Abnormal Detection Method, Device, Electronic Device and Storage Medium for In-Vehicle Messages

By using the MobileNet-GRU teacher model to extract the features of on-board messages and train the lightweight student model, the problem of difficult to take into account in the existing technology of detection accuracy and computing efficiency is solved, and efficient and accurate message abnormality detection in the on-board environment is achieved.

CN119854395BActive Publication Date: 2025-06-13CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510323151.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing on-board message detection methods are difficult to take into account the accuracy of the detection and the computing efficiency of the model, especially in on-board environments where resources are limited.

Method used

The MobileNet-GRU teacher model is used to extract the spatial and temporal characteristics of on-board messages, and train a lightweight student model through knowledge distillation technology, which not only improves detection accuracy but also reduces computational complexity.

Benefits of technology

It realizes efficient deployment of packet abnormality detection in resource-constrained vehicle environments, taking into account the accuracy and computing efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854395B_ABST
    Figure CN119854395B_ABST
Patent Text Reader

Abstract

The present application relates to an abnormal detection method, device, electronic device, and storage medium for in-vehicle messages. The method includes: obtaining training samples, where the training samples include sample in-vehicle messages and hard labels indicating whether the sample in-vehicle messages are abnormal; constructing and training a teacher model using the training samples, where the MobileNet layer in the teacher model is used to extract the spatial features of the sample in-vehicle messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample in-vehicle messages; performing knowledge distillation learning on an initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, where the soft labels refer to the class probability distribution of the sample in-vehicle messages; and detecting whether a target in-vehicle message is abnormal using the trained student model. The present application achieves both the accuracy of message detection and high computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data detection technology, and particularly relates to an abnormal detection method, device, electronic device, and storage medium for vehicle-mounted messages. Background Art

[0002] Existing vehicle-mounted message detection methods mainly rely on rule-based, statistical, or traditional machine learning models. Although these methods are simple to implement, they have obvious limitations. Rule-based methods have poor flexibility and are difficult to adapt to new abnormal patterns; statistical methods are insensitive to sudden anomalies and require a large amount of historical data; while complex deep learning models can improve detection accuracy, but due to their high computational complexity and resource consumption, they are difficult to be efficiently deployed in resource-constrained vehicle-mounted environments. Therefore, the existing technology cannot balance both the accuracy of message detection and the computational efficiency of the model. Summary of the Invention

[0003] This application provides an abnormal detection method, device, electronic device, and storage medium for vehicle-mounted messages to solve the problem of being unable to balance both the accuracy of message detection and the computational efficiency of the model.

[0004] In a first aspect, this application provides an abnormal detection method for vehicle-mounted messages, and the method includes:

[0005] Obtain training samples, where the training samples include sample vehicle-mounted messages and hard labels indicating whether the sample vehicle-mounted messages are abnormal;

[0006] Construct and train a teacher model using the training samples, where the MobileNet layer in the teacher model is used to extract the spatial features of the sample vehicle-mounted messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample vehicle-mounted messages;

[0007] Perform knowledge distillation learning on an initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, where the soft labels refer to the class probability distribution of the sample vehicle-mounted messages;

[0008] Use the trained student model to detect whether the target vehicle-mounted message is abnormal;

[0009] Among them, constructing and training the teacher model using the training samples includes:

[0010] Map the one-dimensional time series features in the sample vehicle-mounted messages to a two-dimensional space to form an image;

[0011] Use the MobileNet layer of the teacher model to extract the overall features and local features of the image, and convert the image into a MobileNet output vector;

[0012] Capture the long-term dependencies of the MobileNet output vector through the update gate and the reset gate in the gated recurrent unit to obtain a gated output vector;

[0013] Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label to obtain a trained teacher model;

[0014] Among them, knowledge distillation learning for the initial student model according to the hard label of the training sample and the soft label output by the teacher model includes:

[0015] Determine an initial MobileNet-GRU student model with a single-layer convolutional kernel and a single-layer gated recurrent unit;

[0016] By adjusting the temperature coefficient, use the softmax function to convert the output of the teacher model into a soft label to obtain a distillation loss function, where the temperature coefficient greater than the first threshold indicates that the distribution area of the soft label is uniform, and the temperature coefficient less than the first threshold indicates that the distribution of the soft label is close to the distribution of the hard label;

[0017] Train the initial student model according to the hard label of the training sample to obtain a hard label loss function;

[0018] Obtain a total loss function according to the weighted sum of the distillation loss function and the hard label loss function, and determine a trained student model according to the loss value of the total loss function.

[0019] Optionally, use the MobileNet layer of the teacher model to extract the global features and local features of the image, and convert the image into a MobileNet output vector, including:

[0020] Extract the local features of the image through the depth convolutional kernel of the MobileNet layer to obtain a first feature map;

[0021] Perform a pointwise convolution operation on the first feature map to obtain a second feature map, where the feature expression ability of the second feature map is greater than that of the first feature map;

[0022] Perform a global max pooling operation on the second feature map to obtain a third feature map containing important features;

[0023] Use the depth convolutional kernel to perform depth convolution on the third feature map until a fourth feature map with a set number of channels is obtained;

[0024] Perform dropout operation and global average pooling operation on the fourth feature map to obtain a MobileNet output vector.

[0025] Optionally, capturing long-range dependencies of the MobileNet output vector through the update gate and reset gate in the gated recurrent unit to obtain a gated output vector includes:

[0026] Process the MobileNet output vector and the hidden state at the previous moment through the update gate in the gated recurrent unit to obtain the first weight value at the current moment, where the first weight value is used to determine how to mix the MobileNet output vector and the hidden state;

[0027] Process the MobileNet output vector and the hidden state at the previous moment through the reset gate in the gated recurrent unit to obtain the second weight value at the current moment, where the second weight value is used to determine the unimportant information that can be ignored in the hidden state;

[0028] Adjust the hidden state at the previous moment through the second weight value, and determine a new candidate hidden state according to the MobileNet output vector;

[0029] Perform a weighted combination of the hidden state at the previous moment and the candidate hidden state through the first weight value to determine the hidden state at the current moment;

[0030] Repeat the above steps in sequence to obtain the hidden state at each moment, and concatenate the hidden states at each moment to obtain a gated output vector.

[0031] Optionally, using the trained student model to detect whether the target vehicle-mounted message is abnormal includes:

[0032] Input the target vehicle-mounted message into the trained student model, where the trained student model is deployed in the vehicle gateway;

[0033] Obtain the message category output by the trained student model, where the message category is used to indicate whether the target vehicle-mounted message is normal or abnormal.

[0034] Optionally, obtaining training samples includes:

[0035] Align the initial vehicle-mounted messages of different lengths, and convert the character-type labels of the aligned messages into numerical-type labels through encoding to obtain sample vehicle-mounted messages;

[0036] Perform radix conversion and normalization processing on the sample vehicle-mounted messages to obtain hard labels.

[0037] In a second aspect, the present application provides an abnormal detection device for in-vehicle messages, and the device includes:

[0038] An acquisition module, configured to acquire training samples, where the training samples include sample in-vehicle messages and hard labels indicating whether the sample in-vehicle messages are abnormal;

[0039] A construction and training module, configured to construct and train a teacher model by using the training samples, where the MobileNet layer in the teacher model is used to extract the spatial features of the sample in-vehicle messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample in-vehicle messages;

[0040] A learning module, configured to perform knowledge distillation learning on an initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, where the soft labels refer to the class probability distribution of the sample in-vehicle messages;

[0041] A detection module, configured to detect whether a target in-vehicle message is abnormal by using the trained student model;

[0042] Wherein, the construction and training module is configured to:

[0043] Map the one-dimensional time series features in the sample in-vehicle messages to a two-dimensional space to form an image;

[0044] Extract the overall features and local features of the image by using the MobileNet layer of the teacher model, and convert the image into a MobileNet output vector;

[0045] Capture the long-range dependence of the MobileNet output vector through the update gate and the reset gate in the gated recurrent unit to obtain a gated output vector;

[0046] Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label to obtain a trained teacher model;

[0047] Wherein, the learning module is configured to:

[0048] Determine an initial MobileNet-GRU student model with a single-layer convolutional kernel and a single-layer gated recurrent unit;

[0049] By adjusting the temperature coefficient, use the softmax function to convert the output of the teacher model into soft labels to obtain a distillation loss function, where the temperature coefficient being greater than a first threshold indicates that the distribution area of the soft labels is uniform, and the temperature coefficient being less than the first threshold indicates that the distribution of the soft labels is close to the distribution of the hard labels;

[0050] Train the initial student model according to the hard labels of the training samples to obtain a hard label loss function;

[0051] Obtain a total loss function based on the weighted sum of the distillation loss function and the hard label loss function, and determine the trained student model according to the loss value of the total loss function.

[0052] In a third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.

[0053] In a fourth aspect, the present application further provides a computer storage medium storing computer-executable instructions for executing the abnormal detection method of vehicle-mounted messages described in any one of the above of the present application.

[0054] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the present application, first, a complex MobileNet-GRU teacher model is constructed. This teacher model can extract both the spatial features and temporal features of vehicle-mounted messages to ensure the accuracy of message anomaly detection. Then, through knowledge distillation technology, a lightweight student model is jointly trained using the soft labels generated by the complex teacher model and the original hard labels. The student model can learn the true category information of the hard labels while obtaining richer feature information from the soft labels, better understanding the similarity between categories, thereby improving the classification accuracy. The lightweight student model can not only improve the calculation efficiency with less computational effort but also ensure the accuracy of message detection, achieving both message detection accuracy and high computational efficiency. Description of the Drawings

[0055] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0056] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0058] Figure 1 Flow chart of a method for detecting anomalies in vehicle-mounted messages provided by an embodiment of the present application;

[0059] Figure 2 Schematic diagram of the preprocessing of training samples provided by an embodiment of the present application;

[0060] Figure 3 Schematic diagram of the construction of a student model provided by an embodiment of the present application;

[0061] Figure 4 Schematic diagram of a process for detecting anomalies in vehicle-mounted messages provided by an embodiment of the present application;

[0062] Figure 5 Schematic diagram of the structure of an anomaly detection device for vehicle-mounted messages provided by an embodiment of the present application;

[0063] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0065] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0066] To solve the problem in the prior art that it is difficult to balance the accuracy of message detection and the high computational efficiency of the model, the embodiments of the present application use a lightweight MobileNet-GRU model for message anomaly detection. The lightweight model can improve the computational efficiency while also improving the message detection accuracy according to the MobileNet-GRU model.

[0067] The method for detecting anomalies in vehicle-mounted messages in the embodiments of the present application can be executed by a processor or a controller of a vehicle.

[0068] Some terms used in the embodiments of the present application are explained below, including the following content:

[0069] In-vehicle message: A data packet transmitted through a CAN (Controller Area Network) bus or a LIN (Local Interconnect Network) bus, containing communication information between various vehicle sensors and controllers, such as speed, temperature, throttle position, etc.

[0070] Hard label: A binary label indicating whether a sample in-vehicle message is abnormal, usually 0 (normal) or 1 (abnormal).

[0071] Soft label: The class probability distribution output by the teacher model, reflecting the possibility of a sample in-vehicle message belonging to different classes, providing richer information on inter-class similarity.

[0072] Knowledge distillation learning: A model compression technique that enables a student model to learn the knowledge of a teacher model to achieve similar performance at a lower computational cost.

[0073] Teacher model: A complex high-performance model used to generate high-quality soft labels to guide the learning of the student model.

[0074] Student model: A simplified model that learns the knowledge of the teacher model through knowledge distillation, has a lower computational complexity, and is suitable for resource-constrained environments.

[0075] MobileNet layer: An efficient convolutional neural network architecture used to extract spatial features of images or data, reducing computational resource consumption.

[0076] Gated Recurrent Unit (GRU): A variant of the recurrent neural network that controls information flow by introducing a gating mechanism, effectively capturing long-range dependencies in time series data.

[0077] Next, in combination with specific implementation manners, a method for detecting anomalies in in-vehicle messages provided by an embodiment of the present application will be described in detail. Taking an application to a processor as an example, as Figure 1 shown, the specific steps are as follows:

[0078] Step 101: Obtain training samples, where the training samples include sample in-vehicle messages and hard labels indicating whether the sample in-vehicle messages are abnormal;

[0079] Step 102: Construct and train a teacher model using the training samples. In the teacher model, the MobileNet layer is used to extract the spatial features of the sample in-vehicle messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample in-vehicle messages;

[0080] Step 103: Perform knowledge distillation learning on the initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, where the soft label refers to the class probability distribution of the sample vehicle-mounted message.

[0081] Step 104: Use the trained student model to detect whether the target vehicle-mounted message is abnormal.

[0082] In step 101, the processor collects a large number of messages from the vehicle communication system as training samples. These messages contain real-time communication information between various vehicle sensors and controllers, such as speed, temperature, throttle position, etc. Each message has a hard label indicating whether the message is abnormal. The hard label is usually a binary value, where 0 represents a normal message and 1 represents an abnormal message, and these labels can be annotated by experts or automatically generated through historical data analysis.

[0083] In step 102, the processor constructs a complex teacher model, which includes two parts: the MobileNet layer and the gated recurrent unit (GRU). The MobileNet layer is used to extract the spatial features of the vehicle-mounted message and capture the static patterns in the message, while the GRU is used to capture the time series features and process the temporal dependencies of the message. The processor uses the cross-entropy loss function to train the teacher model and optimize its classification performance for normal and abnormal messages.

[0084] The processor uses the trained teacher model to predict the training samples and generate soft labels. The soft label refers to the probability distribution of the sample vehicle-mounted message belonging to different classes, reflecting the similarities and differences between classes. The trained teacher model has strong feature extraction and sequence modeling capabilities. On the one hand, it can provide high-quality soft labels to help the student model better learn the subtle differences between classes. On the other hand, it can make full use of rich features to improve the detection accuracy.

[0085] In step 103, the student model is a simplified MobileNet-GRU model with fewer parameters and lower computational complexity. In this embodiment of the application, the soft label is combined with the original hard label to jointly train the initial student model. Through knowledge distillation, the student model not only learns the true class information of the hard label but also learns the soft label provided by the teacher model, enhancing the robustness to noisy data and classification accuracy. This application ensures that the student model can not only inherit the strong feature extraction ability of the teacher model to accurately classify messages but also simplify the calculation through a lightweight model to improve the computational efficiency.

[0086] In step 104, the trained student model is converted into the TensorFlow Lite format (a lightweight model format for running TensorFlow models on specific devices). This format is designed specifically for running models on mobile and embedded devices, supports multiple hardware accelerators, and is suitable for use on automotive gateway devices with limited resources. In this way, the trained student model is deployed to devices such as in-vehicle gateways to receive and process new target vehicle messages in real time. Based on the input target vehicle messages, the student model outputs a prediction result indicating whether they are abnormal, providing timely feedback to the vehicle's safety monitoring system. If an abnormal message is detected, the system will trigger an alarm or take corresponding measures to ensure the stability and safety of the vehicle system.

[0087] The trained student model can operate efficiently on embedded devices and quickly respond to potential security threats. The lightweight student model reduces computational resource consumption, enabling high-precision and low-latency anomaly detection in resource-constrained environments.

[0088] Exemplarily, there is a set of vehicle message data, and each message contains an 8-byte data field. Some of these messages are normal, while others are abnormal. The processor first uses a teacher model (a complex MobileNet-GRU model) to classify these messages and generate soft labels. For example, for a message, the teacher model may output [0.9, 0.1], indicating that the message has a 90% probability of being normal and a 10% probability of being abnormal. Next, a simplified student model is trained using these soft labels and the original hard labels (0 or 1). By learning the soft labels of the teacher model, the student model can not only learn how to distinguish between normal and abnormal messages but also understand the subtle differences between the two. Finally, this lightweight student model can operate efficiently on the in-vehicle gateway to detect in real time whether the newly received messages are abnormal, ensuring the safety and stability of the vehicle system.

[0089] In this application, first, a complex MobileNet-GRU teacher model is constructed. This teacher model can extract both spatial and temporal features of vehicle messages to ensure the accuracy of message anomaly detection. Then, through knowledge distillation technology, a lightweight student model is jointly trained using the soft labels generated by the complex teacher model and the original hard labels, enabling the student model to better understand the similarities between classes and thus improving the classification accuracy. The trained lightweight student model can not only improve computational efficiency with less computational effort but also ensure the accuracy of message detection, achieving both message detection accuracy and high computational efficiency.

[0090] In addition, the lightweight student model occupies less storage space and memory, requires fewer computing resources during inference, and is suitable for deployment on devices with limited storage and memory, such as mobile devices and embedded systems. The reduced computational load and memory access times mean lower energy consumption, which is also applicable to battery-powered devices.

[0091] As an alternative implementation, in step 101, obtaining the training samples includes the following: aligning the initial in-vehicle messages of different lengths, and converting the character-based labels of the aligned messages into numerical labels through encoding to obtain sample in-vehicle messages; performing base conversion and normalization on the sample in-vehicle messages to obtain hard labels.

[0092] Figure 2 This is the preprocessing process of the training samples. The messages support multiple different data lengths. For example, the DLC (Data Length Code) of CAN messages can be up to 8 bytes at most. To facilitate the training of the neural network model, it is necessary to standardize the initial in-vehicle messages of different lengths. First, judge whether the length of the initial in-vehicle message is equal to 8. If it is equal to 8, directly perform One-Hot encoding. If it is less than 8, pad zeros at the end of its payload to ensure that all message lengths are the same.

[0093] The label field of the message dataset is identified by character-based data. For example, normal represents a normal message, and anomaly represents an abnormal message. To facilitate model processing, it is necessary to convert these character-based labels into numerical labels to obtain sample in-vehicle messages. The encoding conversion rule is any one of the following rules. One rule is: normal is encoded as 0, representing a normal message; anomaly is encoded as 1, representing an abnormal message. Another One-Hot encoding rule is: normal is encoded as [1,0], representing a normal message; anomaly is encoded as [0,1], representing an abnormal message.

[0094] Perform base conversion on the sample in-vehicle messages to ensure the consistency of the data format; to scale the message data to the [0,1] interval and improve the training efficiency and stability of the model, the maximum-minimum normalization method is adopted. The specific conversion formula is as follows:

[0095] , where, represents the i-th feature, represents the specific value of the i-th feature in the k-th sample, represents the normalized value. The normalized value is the hard label.

[0096] As an alternative implementation, in step 102, constructing and training the teacher model using the training samples includes the following:

[0097] Step S11: Map the one-dimensional time series features in the sample vehicle-mounted message to a two-dimensional space to form an image.

[0098] Step S12: Use the MobileNet layer of the teacher model to extract the overall features and local features of the image, and convert the image into a MobileNet output vector.

[0099] Step S13: Capture the long-range dependencies of the MobileNet output vector through the update gate and reset gate in the gated recurrent unit to obtain a gated output vector.

[0100] Step S14: Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label to obtain a trained teacher model.

[0101] Each vehicle-mounted message can reach 8 bytes. For example, [0x1A, 0x2B, 0x3C, 0x4D, 0x5E, 0x6F, 0x70, 0x81]. These 8 bytes are the one-dimensional time series features to be processed. The processor maps the one-dimensional time series features of each vehicle-mounted message to a two-dimensional space to form a grayscale or color image. For example, the data of each byte can be converted into pixel values and arranged in a matrix form according to specific rules, so as to ensure that all generated images have the same size (such as 224*224), and perform normalization processing so that the pixel values are in the range of [0, 1], which is convenient for subsequent convolutional neural network processing.

[0102] Exemplarily, there is the following 8-byte vehicle-mounted message data: [0x1A, 0x2B, 0x3C, 0x4D, 0x5E, 0x6F, 0x70, 0x81]. Each byte can be expanded into a 16*16 small block, and all pixel values in each block are set to the value of that byte. Then a total of 16*16*8 small blocks are obtained. Arrange these blocks into a 16*128 matrix. This matrix has 16*128 = 2048 pixels. Use the interpolation method to expand the 16*128 matrix to 224*224, and perform normalization processing on the image so that the pixel values are in the range of [0, 1].

[0103] In the embodiment of the present application, by converting one-dimensional time series data into two-dimensional images, mature image processing technologies and convolutional neural networks (such as MobileNet) can be used to extract complex spatio-temporal features, improve the feature expression ability. In addition, the data in the form of images can more intuitively display the patterns and structures in the message, which helps to capture potential abnormal patterns.

[0104] The processor uses the MobileNet layer in the teacher model to perform multi-step convolution operations on the generated images, extract their global and local features, and finally convert the extracted feature maps into one-dimensional feature vectors (MobileNet output vectors) for subsequent processing. The MobileNet layer can efficiently extract rich spatio-temporal features, significantly improving the model's ability to understand complex patterns. By converting the images into feature vectors, it provides strong support for subsequent time series modeling.

[0105] Next, the processor inputs the MobileNet output vector into the gated recurrent unit (GRU) in the teacher model. Through the update gate and reset gate mechanisms, the GRU can capture the dynamic changes in the time series, ensuring that the model has a good understanding of long-term dependencies. By introducing the gating mechanism, the model can flexibly select which historical information needs to be retained and which information can be ignored, thereby improving the classification accuracy.

[0106] Finally, the processor converts the gated output vector into a prediction result representing the message category, and evaluates the gap between the model prediction result and the true label through a loss function (such as binary cross-entropy loss). The loss function gives a numerical value based on the difference between the predicted value and the true label. The smaller this value is, the more accurate the model's prediction is. By continuously adjusting the model parameters, the performance of the model on the training data gets better and better, and finally converges to a better state. Rigorous evaluation and verification ensure that the model not only performs well on the training data but also maintains high accuracy on unseen data.

[0107] This application converts one-dimensional time series features into two-dimensional images, which is convenient for the MobileNet network to extract features. Then, it extracts the global and local features of the images through a convolutional neural network, and captures the dynamic changes in the time series through a gating mechanism. This application constructs an efficient hybrid model through the combination of the MobileNet layer and the gated recurrent unit (GRU). This combination enables the model to utilize both the spatial and temporal features of in-vehicle messages simultaneously, improving the accuracy and robustness of intrusion detection.

[0108] As an alternative implementation, in step 103, knowledge distillation learning of the initial student model based on the hard labels of the training samples and the soft labels output by the teacher model includes the following content:

[0109] Step S21: Determine the initial MobileNet-GRU student model with a single-layer convolutional kernel and a single-layer gated recurrent unit;

[0110] Step S22: By adjusting the temperature coefficient, the output of the teacher model is converted into soft labels using the softmax function to obtain the distillation loss function. Here, a temperature coefficient greater than the first threshold indicates that the distribution area of the soft labels is uniform, and a temperature coefficient less than the first threshold indicates that the distribution of the soft labels is close to the distribution of the hard labels;

[0111] Step S23: Train the initial student model according to the hard labels of the training samples to obtain the hard label loss function;

[0112] Step S24: Obtain the total loss function based on the weighted sum of the distillation loss function and the hard label loss function, and determine the trained student model according to the loss value of the total loss function.

[0113] Figure 3 As shown in the process schematic diagram constructed for the student model, after the teacher model outputs soft labels, the initial student model is trained using hard labels and soft labels. Based on the distillation loss function and the hard label loss function, the total loss function of the student model is obtained, and thus the student model is trained. The process of training the student model is as follows.

[0114] The processor constructs a lightweight student model, which includes a single-layer convolutional kernel (for preliminary feature extraction) and a single-layer gated recurrent unit (for capturing time series features).

[0115] Then, a temperature coefficient T is introduced to adjust the output distribution of the teacher model. The temperature coefficient T controls the hardness of the probability distribution output by the Softmax function (soft maximum function). When T is large, the Softmax output is smoother, and the category probability distribution area is uniform, indicating that the soft label is softer, such as [0.9, 0.1]. When T is small, the Softmax output is sharper, close to the distribution of the hard labels, indicating that the soft label is closer to the true label, such as 0 or 1. Using the adjusted temperature coefficient T, the output of the teacher model is converted into soft labels, and the distillation loss function is calculated based on these soft labels. The distillation loss function measures the difference between the output of the student model and the soft labels of the teacher model. Through temperature coefficient adjustment, the soft labels can provide more inter-class similarity information, helping the student model better understand the subtle differences between categories and improving classification accuracy.

[0116] Among them, the calculation formula for the soft label is:

[0117] , where is the soft label, T is the temperature coefficient, represents the unnormalized predicted score that the input sample belongs to the i-th category, represents the unnormalized predicted score that the input sample belongs to the j-th category, and k is the total number of samples.

[0118] The formula for the distillation loss function is as follows:

[0119]

[0120] Among them, represents the distillation loss function, is the i-th output of the teacher model at temperature T, is the i-th output of the student model at temperature coefficient T, and k is the total number of samples.

[0121] The processor uses the true labels (i.e., hard labels) of the training samples to supervise the training of the student model. The hard label is a binary value indicating whether the packet is abnormal (0 means normal, 1 means abnormal). An appropriate loss function (such as binary cross-entropy loss) can be defined to evaluate the gap between the prediction results of the student model and the hard labels. The loss function gives a numerical value based on the difference between the predicted value and the true label. The smaller this numerical value is, the more accurate the model's prediction is.

[0122] The formula for the hard label loss function is as follows:

[0123]

[0124] Among them, is the hard label loss function, is a hard label vector, represents the i-th class output of the unsoftened student model, and k is the total number of samples.

[0125] The distillation loss function and the hard label loss function are weighted and combined to form the total loss function. The choice of weights can be adjusted according to the specific application scenario to balance the influence of soft labels and hard labels. By combining the distillation loss and the hard label loss, the total loss function can comprehensively evaluate the performance of the student model, ensuring that it not only inherits the knowledge of the teacher model but also accurately classifies the actual samples.

[0126] The formula for the total loss function is as follows:

[0127] Among them, is the total loss function, is a hyperparameter that can be adjusted to a reference or dynamically adjusted according to experience.

[0128] In this application, the processor first determines an initial student model with a single-layer convolutional kernel and a single-layer gated recurrent unit to ensure its efficiency and adaptability. Then, by adjusting the temperature coefficient T, the output of the teacher model is converted into soft labels using the Softmax function to generate a distillation loss function, which flexibly controls the hardness of the soft labels. Next, the student model is trained according to the hard labels of the training samples to obtain a hard label loss function, ensuring that the model can accurately classify actual samples. Finally, the distillation loss function and the hard label loss function are combined by weighting to form a total loss function, and the trained student model is determined according to the loss value of the total loss function.

[0129] As an alternative implementation, in step S12, determining the MobileNet output vector includes the following:

[0130] Step S121: Extract the local features of the image through the depthwise convolutional kernel of the MobileNet layer to obtain a first feature map;

[0131] Step S122: Perform a pointwise convolution operation on the first feature map to obtain a second feature map, where the feature expression ability of the second feature map is greater than that of the first feature map;

[0132] Step S123: Perform a global max pooling operation on the second feature map to obtain a third feature map containing important features;

[0133] Step S124: Use the depthwise convolutional kernel to perform depthwise convolution on the third feature map until a fourth feature map with a set number of channels is obtained;

[0134] Step S125: Perform a dropout operation and a global average pooling operation on the fourth feature map to obtain the MobileNet output vector.

[0135] The processor converts the sample packet data into an image as the input to ensure that the input format is unified and suitable for subsequent processing. The depthwise convolutional kernel is used to preliminarily extract the local features of the image. The depthwise convolution processes each pixel point channel by channel, capturing the local texture and edge information in the image and generating a first feature map. The depthwise convolution can preliminarily extract the local features in the image while maintaining the computational efficiency, laying a foundation for subsequent more complex feature extraction. This stage mainly focuses on capturing the basic structure and pattern of the image.

[0136] The processor applies a pointwise convolution operation to expand the number of channels of the first feature map, obtaining a second feature map. Pointwise convolution not only improves the feature representation ability but also maintains the computational efficiency. By increasing the number of channels, the model can capture more complex and rich feature information, enhancing the feature representation ability and enabling the model to better understand different types of local features. Pointwise convolution significantly enhances the feature representation ability, enabling the model to capture more diverse local features and providing richer information for subsequent feature fusion and classification tasks.

[0137] The processor performs a global max pooling operation on the second feature map to extract the most important features, obtaining a third feature map. Max pooling selects the maximum value in each local region, reducing the spatial dimension of the feature map while retaining key information. Through the max pooling operation, the model can effectively reduce the spatial dimension of the feature map while retaining the most significant features, improving the computational efficiency and enhancing the robustness of the features. Global max pooling reduces the spatial dimension of the feature map while retaining the most important features, improving the computational efficiency and robustness of the model and enabling the model to focus on the most representative features.

[0138] The processor applies a depth convolution operation again to convolve the third feature map, gradually extracting deeper-level features. Finally, a fourth feature map is obtained. Further extract more complex features and increase the number of channels to capture more details. Through multiple steps of depth convolution operations, the model can gradually extract deeper-level features, enhancing the ability to understand complex patterns, ensuring that the model can capture more detailed information, and thus improving the classification accuracy.

[0139] The processor performs a dropout operation (dropping operation) on the fourth feature map, randomly setting some neurons to zero to prevent overfitting and improve the generalization ability of the model. Dropout, as a regularization technique, can effectively prevent the model from relying too much on certain specific features during training and enhance the robustness of the model. The fourth feature map after the dropout operation is converted into a one-dimensional feature vector through global average pooling, which not only reduces the spatial dimension of the feature map but also provides a smoother and more stable feature representation, ensuring the quality of the feature vector.

[0140] Exemplarily, the processing flow of the MobileNet layer includes the following:

[0141] 1. Initial input: Convert the CAN message data into an image with a dimension of 225*225*3 as the input for feature extraction.

[0142] 2. Depth convolution: Use a 3*3*3 depth convolution kernel to extract local features, obtaining a feature map of 222*222*3. This step initially extracts the local features in the image.

[0143] 3. Pointwise Convolution: Apply a pointwise convolution operation of 1*1*32 to expand the number of channels from 3 to 32, obtaining a feature map of 222*222*32. The pointwise convolution enhances the feature representation ability while maintaining computational efficiency.

[0144] 4. Global Max Pooling: Perform global max pooling on the 222*222*32 feature map to extract the most important features, obtaining a feature map of 111*111*32. Max pooling reduces the spatial dimension while retaining key information.

[0145] 5. Repeated Convolution: Apply the depth convolution operation again to convolve the 111*111*32 feature map, finally obtaining a feature map of 108*108*64. This step further extracts more complex features and increases the number of channels to capture more details.

[0146] 6. Dropout Operation: Perform the dropout operation on the 108*108*64 feature map, randomly setting some neurons to zero to prevent overfitting and improve the generalization ability of the model.

[0147] 7. Global Average Pooling and Reshaping: Convert the feature map into a one-dimensional feature vector through global average pooling and adjust it to a format suitable for input to the GRU network (1*64 tensor) through the reshape operation. This step provides a smoother and more stable feature representation.

[0148] MobileNet is an efficient convolutional neural network architecture that uses depthwise separable convolutions to reduce the amount of computation and the number of parameters. In the embodiments of this application, the MobileNet layer can efficiently and accurately extract the spatial features of in-vehicle CAN messages.

[0149] As an optional implementation manner, in step S13, determining the gated output vector includes:

[0150] Step S131: Process the MobileNet output vector and the hidden state at the previous moment through the update gate in the gated recurrent unit to obtain the first weight value at the current moment, where the first weight value is used to determine how to mix the MobileNet output vector and the hidden state;

[0151] Step S132: Process the MobileNet output vector and the hidden state at the previous moment through the reset gate in the gated recurrent unit to obtain the second weight value at the current moment, where the second weight value is used to determine the unimportant information that can be ignored in the hidden state;

[0152] Step S133: Adjust the hidden state at the previous moment through the second weight value and determine a new candidate hidden state according to the MobileNet output vector;

[0153] Step S134: Determine the hidden state at the current moment by weighted combination of the hidden state and the candidate hidden state at the previous moment using the first weight value;

[0154] Step S135: Repeat the above steps sequentially to obtain the hidden state at each moment, and concatenate the hidden states at each moment to obtain the gated output vector.

[0155] In the GRU layer, first, an Update Gate is introduced. It is responsible for determining how to combine the output vector at the current moment from the MobileNet layer and the hidden state at the previous moment. This process involves performing a linear transformation on both and applying an activation function (usually the Sigmoid function) to generate the first weight value zt. The first weight value zt represents the degree to which the model should retain the old state information (hidden state) and adopt the new input information (i.e., the MobileNet output vector) at the current time point. In other words, zt determines the mixing ratio between the new information and the old information. Through a carefully tuned information fusion mechanism, it is ensured that the model can learn long-term dependencies from the past hidden states and appropriately introduce new feature information, avoiding forgetting important historical data or prematurely accepting new information, thereby improving the model's expressive ability and accuracy.

[0156] Next is the role of the Reset Gate. It also calculates the second weight value rt based on the MobileNet output vector at the current moment and the hidden state at the previous moment. This weight value is used to evaluate which parts of the old state can be ignored as it indicates the degree of unimportant information in the old state. By applying rt to the hidden state at the previous moment, the GRU can selectively forget those less relevant past information to better focus on the novelty and changes brought by the current input. The Reset Gate enables the GRU to have a more refined memory management ability, effectively filtering out unnecessary information, reducing the model burden while enhancing its memory and understanding of key information, and improving the model's understanding depth of sequential data.

[0157] Using the hidden state adjusted by the Reset Gate and the MobileNet output vector at the current moment, the GRU calculates a new candidate hidden state . This step usually includes the application of a tanh activation function to ensure that the candidate state falls within a reasonable numerical range.

[0158] Finally, through the first weight value zt of the Update Gate, the GRU will use the hidden state at the previous moment and the just-calculated candidate hidden state Perform a weighted combination to determine the final hidden state \(h_t\) at the current moment.

[0159] Through the above steps, the GRU can flexibly adjust its internal state at each time point, maintaining effective modeling of long time series while being able to quickly respond to short-term changes. This dynamic characteristic makes the GRU very suitable for processing datasets with complex spatio-temporal structures, such as video frame sequences or time series prediction tasks.

[0160] Exemplarily, the calculation process of the gated recurrent unit is as follows:

[0161] 1. Update the gate to calculate the gate state information at time \(t\).

[0162] Among them, represents the first weight value, is the sigmod function, is the parameter matrix of the update gate, is the hidden state at time \(t - 1\), is the \(1\times64\) tensor finally obtained in the MobileNet network layer, is the bias vector of the update gate.

[0163] 2. Reset the gate to calculate the gate state information at time \(t\).

[0164] Among them, represents the second weight value, is the sigmod function, is the parameter matrix of the reset gate, is the hidden state at time \(t - 1\), is the \(1\times64\) tensor finally obtained in the MobileNet network layer, is the bias vector of the reset gate.

[0165] 3. The reset gate calculates the candidate hidden state.

[0166]

[0167] Among them, represents the candidate hidden state, \(W\) is the parameter matrix for linearly transforming the reset gate, represents the second weight value, is the hidden state at time \(t - 1\), is the \(1\times64\) tensor finally obtained in the MobileNet network layer, \(b\) is the bias vector for adjusting the result after linear transformation.

[0168] 4. The update gate calculates the current hidden state.

[0169] Among them, represents the current hidden state, represents the first weight value, is the hidden state at time t-1, represents the candidate hidden state.

[0170] 5. Loss function of the MobileNet-GRU model.

[0171]

[0172] Among them, loss is the loss function, n is the total number of samples, is the hard label value (0 or 1) of the i-th sample, is the predicted value of the model for the i-th sample, that is, the probability that the model predicts the label value of the i-th sample as 1. In the hard label, 0 represents normal CAN message information, and 1 represents abnormal message information.

[0173] In this application, the output vector of MobileNet is processed through the update gate and reset gate in GRU, which can efficiently capture the long-range dependence relationship of the image features evolving over time.

[0174] As an optional implementation manner, in step 104, using the trained student model to detect whether the target vehicle-mounted message is abnormal includes: inputting the target vehicle-mounted message into the trained student model, where the trained student model is deployed in the vehicle gateway; obtaining the message category output by the trained student model, where the message category is used to indicate whether the target vehicle-mounted message is normal or abnormal.

[0175] The trained student model after knowledge distillation training is deployed in the vehicle gateway, and the preprocessed target vehicle-mounted message is directly input into the student model. Since the student model is optimized and designed, it can operate efficiently in a resource-constrained environment, so real-time processing can be achieved. The student model outputs a probability distribution according to the input target vehicle-mounted message, indicating the possibility that the message belongs to the normal or abnormal category, and then feeds back the determination result to the vehicle safety monitoring system or other relevant modules for further measures (such as triggering an alarm, recording a log, automatic repair, etc.).

[0176] This application provides a schematic diagram of the abnormal detection process of vehicle-mounted messages, as Figure 4 shown, including the following steps.

[0177] Step 401: Preprocess the initial vehicle-mounted message to obtain training samples.

[0178] Step 402: Convert the training samples into images.

[0179] Step 403: Construct and train a teacher model of MobileNet-GRU using images.

[0180] Step 404: Determine whether the model accuracy meets the requirements. If not, return to Step 403 to retrain the model. If it meets the requirements, execute Step 405.

[0181] Step 405: Save the teacher model.

[0182] Step 406: Perform knowledge distillation using the soft labels output by the teacher model and the hard labels of the training samples to train the student model.

[0183] Step 407: Deploy the trained student model to the vehicle gateway.

[0184] Step 408: Use the deployed student model to determine whether the message data is abnormal.

[0185] Based on the same technical concept, the present application provides an abnormal detection device for vehicle-mounted messages, as Figure 5 shown. The device includes:

[0186] An acquisition module 501, configured to acquire training samples, where the training samples include sample vehicle-mounted messages and hard labels indicating whether the sample vehicle-mounted messages are abnormal;

[0187] A construction and training module 502, configured to construct and train a teacher model using the training samples, where the MobileNet layer in the teacher model is used to extract the spatial features of the sample vehicle-mounted messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample vehicle-mounted messages;

[0188] A learning module 503, configured to perform knowledge distillation learning on the initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, where the soft labels refer to the class probability distribution of the sample vehicle-mounted messages;

[0189] A detection module 504, configured to detect whether the target vehicle-mounted message is abnormal using the trained student model;

[0190] Optionally, the construction and training module 502 is configured to:

[0191] Map the one-dimensional time series features in the sample vehicle-mounted messages to a two-dimensional space to form an image;

[0192] Extract the global features and local features of the image using the MobileNet layer of the teacher model, and convert the image into a MobileNet output vector;

[0193] Capture the long-range dependencies of the MobileNet output vector through the update gate and reset gate in the gated recurrent unit to obtain a gated output vector;

[0194] Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label, and obtain the trained teacher model;

[0195] Optionally, the learning module 503 is used to:

[0196] Determine the initial MobileNet-GRU student model with a single-layer convolutional kernel and a single-layer gated recurrent unit;

[0197] By adjusting the temperature coefficient, use the softmax function to convert the output of the teacher model into a soft label, and obtain the distillation loss function. Among them, when the temperature coefficient is greater than the first threshold, it means that the distribution area of the soft label is uniform, and when the temperature coefficient is less than the first threshold, it means that the distribution of the soft label is close to the distribution of the hard label;

[0198] Train the initial student model according to the hard label of the training sample to obtain the hard label loss function;

[0199] Obtain the total loss function according to the weighted sum of the distillation loss function and the hard label loss function, and determine the trained student model according to the loss value of the total loss function.

[0200] Optionally, the construction and training module 502 is used to:

[0201] Extract the local features of the image through the depth convolutional kernel of the MobileNet layer to obtain the first feature map;

[0202] Perform a pointwise convolution operation on the first feature map to obtain the second feature map, where the feature expression ability of the second feature map is greater than that of the first feature map;

[0203] Perform a global max pooling operation on the second feature map to obtain the third feature map containing important features;

[0204] Use the depth convolutional kernel to perform depth convolution on the third feature map until the fourth feature map with a set number of channels is obtained;

[0205] Perform a dropout operation and a global average pooling operation on the fourth feature map to obtain the MobileNet output vector.

[0206] Optionally, the construction and training module 502 is used to:

[0207] Process the MobileNet output vector and the hidden state at the previous moment through the update gate in the gated recurrent unit to obtain the first weight value at the current moment, where the first weight value is used to determine how to mix the MobileNet output vector and the hidden state;

[0208] Process the MobileNet output vector and the hidden state at the previous moment through the reset gate in the gated recurrent unit to obtain the second weight value at the current moment, where the second weight value is used to determine the non-important information that can be ignored in the hidden state;

[0209] Adjust the hidden state at the previous moment through the second weight value, and determine a new candidate hidden state according to the MobileNet output vector;

[0210] Perform a weighted combination of the hidden state at the previous moment and the candidate hidden state through the first weight value to determine the hidden state at the current moment;

[0211] Repeat the above steps in sequence to obtain the hidden state at each moment, and concatenate the hidden states at each moment to obtain the gated output vector.

[0212] Optionally, the detection module 504 is used for:

[0213] Input the target vehicle-mounted message into the trained student model, where the trained student model is deployed in the vehicle gateway;

[0214] Obtain the message category output by the trained student model, where the message category is used to indicate whether the target vehicle-mounted message is normal or abnormal.

[0215] Optionally, the acquisition module 501 is used for:

[0216] Align the initial vehicle-mounted messages of different lengths, and convert the character labels of the aligned messages into numerical labels through encoding to obtain sample vehicle-mounted messages;

[0217] Perform radix conversion and normalization processing on the sample vehicle-mounted messages to obtain hard labels.

[0218] As Figure 6 shown, an embodiment of the present application provides an electronic device, including a processor 601, a communication interface 602, a memory 603, and a communication bus 604, where the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0219] The memory 603 is used to store computer programs.

[0220] In an embodiment of the present application, when the processor 601 is used to execute the program stored on the memory 603, it implements the vehicle-mounted message anomaly detection method provided by any one of the foregoing method embodiments.

[0221] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the abnormal detection method for in-vehicle messages provided in any of the foregoing method embodiments are implemented.

[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0224] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be executed in the particular order described or illustrated, unless the execution order is explicitly stated. It should also be understood that alternative or additional steps may be used.

[0225] The above description is only the specific implementation manners of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for detecting anomalies of vehicle-borne messages, characterized in that: The method comprises: Acquire a training sample, wherein the training sample includes a sample vehicle message and a hard label indicating whether the sample vehicle message is abnormal; The training samples are used to construct and train a teacher model, wherein the MobileNet layer in the teacher model is used to extract the spatial features of the sample vehicle messages, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample vehicle messages; Performing knowledge distillation learning on the initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, wherein the soft labels refer to the category probability distribution of the sample vehicle-borne messages; Use the trained student model to detect whether the target vehicle message is abnormal; Wherein, using the training samples to construct and train the teacher model includes: Mapping the one-dimensional time series features in the sample vehicle message into a two-dimensional space to form an image; Extracting global and local features of the image using the MobileNet layer of the teacher model, and converting the image into a MobileNet output vector; Capturing the long-range dependency of the MobileNet output vector through an update gate and a reset gate in the gated recurrent unit to obtain a gated output vector; Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label to obtain a trained teacher model; Wherein, performing knowledge distillation learning on the initial student model according to the hard labels of the training samples and the soft labels output by the teacher model includes: Determine the initial MobileNet-GRU student model with a single layer of convolutional kernels and a single layer of gated recurrent units; The output of the teacher model is converted into a soft label by adjusting the temperature coefficient and using a softmax function to obtain a distillation loss function, wherein the temperature coefficient being greater than a first threshold value indicates that the distribution area of ​​the soft label is uniform, and the temperature coefficient being less than the first threshold value indicates that the distribution of the soft label is close to the distribution of the hard label; Training the initial student model according to the hard labels of the training samples to obtain a hard label loss function; A total loss function is obtained according to the weighted sum of the distillation loss function and the hard label loss function, and a trained student model is determined according to the loss value of the total loss function; The step of mapping the one-dimensional time series features in the sample vehicle message into a two-dimensional space to form an image includes: Expand each byte in each on-board message into a block, wherein all pixel values ​​in each block are set to the value of the byte; Arrange multiple blocks into a matrix according to specific rules; Expanding the matrix by interpolation to obtain an image; By normalizing the image, an image with pixel values ​​in the interval [0, 1] is obtained.

2. The method according to claim 1, characterized in that Using the MobileNet layer of the teacher model to extract the overall features and local features of the image, and converting the image into a MobileNet output vector includes: Extracting local features of the image through the deep convolution kernel of the MobileNet layer to obtain a first feature map; Performing a point-by-point convolution operation on the first feature map to obtain a second feature map, wherein the feature expression capability of the second feature map is greater than the feature expression capability of the first feature map; Performing a global maximum pooling operation on the second feature map to obtain a third feature map containing important features; Performing depth convolution on the third feature map using the depth convolution kernel until a fourth feature map having a set number of channels is obtained; A dropout operation and a global average pooling operation are performed on the fourth feature map to obtain a MobileNet output vector.

3. The method according to claim 1, characterized in that The long-range dependency of the MobileNet output vector is captured by the update gate and the reset gate in the gated recurrent unit, and the gated output vector is obtained, including: Processing the MobileNet output vector and the hidden state at the previous moment through an update gate in the gated recurrent unit to obtain a first weight value at the current moment, wherein the first weight value is used to determine how to mix the MobileNet output vector and the hidden state; Processing the MobileNet output vector and the hidden state at the previous moment through a reset gate in the gated recurrent unit to obtain a second weight value at the current moment, wherein the second weight value is used to determine non-important information that can be ignored in the hidden state; Adjusting the hidden state at the previous moment by using the second weight value, and determining a new candidate hidden state according to the MobileNet output vector; Performing a weighted combination of the hidden state at the previous moment and the candidate hidden state by using the first weight value to determine the hidden state at the current moment; Repeat the above steps in sequence to obtain the hidden state at each moment, and concatenate the hidden states at each moment to obtain the gated output vector.

4. The method according to claim 1, characterized in that: The trained student model is used to detect whether the target vehicle message is abnormal, including: Inputting the target vehicle-borne message into the trained student model, wherein the trained student model is deployed in a vehicle gateway; Obtain a message category output by the trained student model, wherein the message category is used to indicate whether the target vehicle-mounted message is normal or abnormal.

5. The method according to claim 1, characterized in that Obtaining training samples includes: Aligning initial vehicle-borne messages of different lengths, and converting character labels of the aligned messages into numerical labels through encoding to obtain sample vehicle-borne messages; The sample vehicle-borne message is subjected to base-to-base conversion and normalization processing to obtain a hard label.

6. A vehicle-borne message anomaly detection device, characterized in that: The device comprises: An acquisition module, configured to acquire a training sample, wherein the training sample includes a sample vehicle message and a hard label indicating whether the sample vehicle message is abnormal; A construction and training module, used to construct and train a teacher model using the training samples, wherein the MobileNet layer in the teacher model is used to extract the spatial features of the sample vehicle message, and the gated recurrent unit in the teacher model is used to capture the time series features of the sample vehicle message; A learning module, configured to perform knowledge distillation learning on an initial student model according to the hard labels of the training samples and the soft labels output by the teacher model, wherein the soft labels refer to the category probability distribution of the sample vehicle-borne messages; A detection module, used to detect whether the target vehicle message is abnormal using the trained student model; Wherein, the construction and training module is used for: Mapping the one-dimensional time series features in the sample vehicle message into a two-dimensional space to form an image; Extracting global and local features of the image using the MobileNet layer of the teacher model, and converting the image into a MobileNet output vector; Capturing the long-range dependency of the MobileNet output vector through an update gate and a reset gate in the gated recurrent unit to obtain a gated output vector; Adjust the loss value based on the difference between the detection result corresponding to the gated output vector and the true label to obtain a trained teacher model; Wherein, the learning module is used for: Determine the initial MobileNet-GRU student model with a single layer of convolutional kernels and a single layer of gated recurrent units; The output of the teacher model is converted into a soft label by adjusting the temperature coefficient and using a softmax function to obtain a distillation loss function, wherein the temperature coefficient being greater than a first threshold value indicates that the distribution area of ​​the soft label is uniform, and the temperature coefficient being less than the first threshold value indicates that the distribution of the soft label is close to the distribution of the hard label; Training the initial student model according to the hard labels of the training samples to obtain a hard label loss function; A total loss function is obtained according to the weighted sum of the distillation loss function and the hard label loss function, and a trained student model is determined according to the loss value of the total loss function; The construction and training modules are specifically used for: Expand each byte in each on-board message into a block, wherein all pixel values ​​in each block are set to the value of the byte; Arrange multiple blocks into a matrix according to specific rules; Expanding the matrix by interpolation to obtain an image; By normalizing the image, an image with pixel values ​​in the interval [0, 1] is obtained.

7. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-5 when executing a program stored in a memory.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Lightweight action recognition model training method and system

    CN117877111A

  • Vehicle-mounted network anomaly detection method and system based on semi-supervised reinforcement learning

    CN118337517A

  • Internet-of-things hostile attack traffic detection method and device fusing knowledge distillation and space-time diagram neural network, and storage medium

    CN118631554A