Intelligent equipment identification method and device based on mask auto-encoder
Through self-supervised pre-training and downstream task fine-tuning of masked autoencoder, the problems of feature redundancy and expert dependence in intelligent device recognition are solved, and efficient intelligent device type recognition is achieved.
Patent Information
- Application Number
- CN202410172772.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, intelligent device identification methods rely on expert experience to extract features, resulting in feature redundancy and noise, affecting the generalization ability and identification accuracy of the model.
Mask autoencoder is used to preprocess and feature extraction of intelligent device traffic data. Through self-supervised pre-training and downstream task fine-tuning, an intelligent device recognition model is built to realize automated feature extraction and reduce feature redundancy.
It improves the accuracy and robustness of smart device recognition, reduces dependence on expert experience, and improves the generalization ability of the model.
Smart Images

Figure CN120455383A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer networks and relates to a method and device for identifying intelligent devices based on traffic analysis. Background Art
[0002] In the current intelligent era, with the rapid development of sensing technology, communications technology, big data and artificial intelligence, cloud computing capabilities, and contextual integration and operational capabilities, coupled with people's demand for energy conservation and resource management, convenience, improved efficiency, risk and error reduction, and optimized management, the number of smart devices continues to increase, and their types are becoming increasingly diverse. These devices play a significant role in smart homes, healthcare, smart agriculture, and smart cities, profoundly impacting every aspect of people's production and daily lives. The widespread use of smart devices has led to increasing concerns about their security. The limited resources of smart devices limit their security capabilities. At the same time, manufacturers may overlook security design in their pursuit of market share, resulting in many devices being vulnerable to vulnerabilities and attacks.
[0003] Intelligent device identification technology based on traffic analysis is a crucial prerequisite for secure network management and security policy implementation. For network administrators, identifying all intelligent devices on the network allows them to provide security protection based on publicly available vulnerability information. Leveraging analysis, malicious network activity can be prevented by promptly detecting and isolating suspected compromised devices for further investigation.
[0004] In recent years, numerous research projects have used deep learning techniques to identify smart devices based on traffic analysis, further improving accuracy. However, most studies rely on manual extraction of targeted features by professionals, while a smaller number use genetic algorithms and methods that leverage metric entropy to automatically extract protocol header features. Manual feature extraction relies heavily on expert experience, while automated feature selection suffers from redundancy and noise. Therefore, this reliance on experts and feature redundancy can, to a certain extent, affect the model's generalization and recognition capabilities.
[0005] Therefore, it is necessary to conduct research on smart device identification technology based on masked autoencoders, and to study a method that can automatically extract features and reduce feature redundancy, and ultimately accurately identify the type of smart devices. Summary of the Invention
[0006] In response to the problems existing in the prior art, the present invention aims to provide a method and apparatus for identifying smart devices based on a masked autoencoder. The present invention collects the raw traffic generated by smart devices, labels it into a data set, pre-processes it to form a general representation of the traffic data, and then pre-trains and fine-tunes the masked autoencoder to obtain a smart device identification model. The raw traffic generated by actual smart devices is then pre-processed to form a general representation of the traffic data, which is then input into the smart device identification model to identify its specific smart device type. The present invention implements smart device traffic collection and labeling, data pre-processing, data set partitioning, model structure design, smart device identification model construction, and device type inference functions to address the expert dependency and feature redundancy issues existing in the prior art.
[0007] In a first aspect, an embodiment of the present invention provides a method for identifying a smart device based on masked self-encoding, comprising:
[0008] S1. Collect and annotate smart device traffic data. Use packet capture tools such as TCPDump to collect and annotate smart device traffic. The annotations include the smart device type, manufacturer, and model corresponding to the MAC address in the traffic, forming smart device traffic data (i.e., pcap files) and their corresponding annotations.
[0009] S2. Preprocess the labeled smart device traffic data. Each bit stream in the smart device traffic data (a bit stream contains all the bits of a data packet sent by the device) is processed to form a general representation of the traffic data. At the same time, the formed general representation of the traffic data is converted into a two-dimensional traffic data graph, forming a two-dimensional traffic data graph and its corresponding annotation content. Finally, according to the 80 / 20 principle, a random function is used to divide it into a training data set and a test data set, where the training data set accounts for 80% and the test data set accounts for 20%.
[0010] S3. Construct a masked autoencoder, consisting of a masker, encoder, and decoder. Its main features include a high-masking rate random masking strategy and an asymmetric encoder and decoder design. The masker masks the two-dimensional traffic data graph by 75% to form a partially visible data graph. The encoder, acting as a feature extractor, maps the partially visible data graph into a low-dimensional, compact latent space to form an embedding vector. The decoder, acting as a feature reconstructor, reconstructs the embedding vector to form a reconstructed two-dimensional traffic data graph.
[0011] S4. Perform self-supervised pre-training. Pre-process the two-dimensional traffic data graph from the training dataset and input it into the masked autoencoder to obtain a reconstructed two-dimensional traffic data graph. The masked autoencoder is optimized by minimizing the mean squared error loss calculated for the original and reconstructed pixels of the masked portion, resulting in a pre-trained masked autoencoder model. This part of the training is self-supervised and does not require annotation.
[0012] S5. Fine-tune the downstream tasks. Input the two-dimensional traffic data graph of sample i in the training data set into the encoder of the pre-trained optimized masked autoencoder to obtain an embedding vector and input it into the softmax layer to obtain the annotation content with the maximum probability predicted for the sample i; optimize the encoder according to the loss of the annotation content with the maximum probability predicted for the sample i and the annotation content of the sample i; then use the test data set to verify the optimized encoder, and use it as the smart device recognition model after verification. The present invention sends the divided training data set with annotated content into the encoder of the pre-trained masked autoencoder model, and then connects a softmax layer to obtain the annotation content with the maximum probability predicted, calculates the loss through the cross-entropy loss function, and obtains the smart device recognition model by fine-tuning the set training strategy and verifying with the test data set. The training strategy is to fix some parameters in the encoder (the encoder contains 12 Block layers) in the pre-trained masked autoencoder, and only adjust the parameters of the last Block layer.
[0013] S6. Perform smart device identification. For the original traffic generated by the smart device in practice, a two-dimensional traffic data graph is obtained through step S2. This data graph is input into the smart device identification model. Based on the prediction results of the encoder in the model, the specific smart device type is identified.
[0014] In a second aspect, an embodiment of the present invention provides a smart device identification device based on masked self-encoding, characterized by comprising:
[0015] Data collection module, used to collect and annotate smart device traffic;
[0016] The data preprocessing module is used to process each bit stream of the traffic, form a general representation of the traffic data, and convert it into a two-dimensional traffic data graph, and finally divide it into training data sets and test data sets;
[0017] Model construction module, used to build the network structure of the masked autoencoder, including the masker, encoder, decoder, etc.
[0018] Model pre-training module, used to build a masked autoencoder pre-training model based on self-supervision;
[0019] The model fine-tuning module is used to apply the pre-trained model to actual downstream tasks and fine-tune the smart device recognition model;
[0020] The model recognition module is used to identify the type of smart devices in actual smart device traffic.
[0021] The advantages of the present invention are as follows:
[0022] 1. A general flow representation is constructed based on a single piece of traffic data, and a specific transformation is used to convert it into a two-dimensional flow data graph.
[0023] 2. Self-supervised pre-training was performed to train the masked autoencoder without the need for labeled content.
[0024] 3. Fine-tune downstream tasks. The training strategy used only adjusts some parameters to speed up training and better fit downstream tasks.
[0025] 4. Use it in the field of smart device identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flow chart of a smart device identification method based on masked self-encoding provided by an embodiment of the present invention.
[0027] Figure 2 It is a general representation diagram of the flow data formed in step S2 provided by an embodiment of the present invention.
[0028] Figure 3 This is a two-dimensional flow data graph formed in step S2 provided in an embodiment of the present invention.
[0029] Figure 4 4 is a structural diagram of the masked autoencoder formed in step S3 provided in an embodiment of the present invention.
[0030] Figure 5 This is a functional module diagram of an intelligent device identification apparatus based on masked self-encoding provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The present invention will be described in further detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0032] The Masked AutoEncoder (MAE) is a scalable self-supervised learning method for computer vision. Published at the 2022 Conference on Computer Vision and Pattern Recognition, it is primarily used for image classification. MAE randomly masks patches in the input image and then reconstructs the missing pixels. Using a high-ratio mask and an asymmetric encoder-decoder architecture, it improves model training speed and accuracy.
[0033] like Figure 1 The present invention provides a flowchart of a method for identifying smart devices based on mask self-encoding. Figure 1 , the method mainly includes:
[0034] S1. Intelligent device traffic data collection and labeling;
[0035] Specifically, an intelligent device traffic collection platform including a variety of intelligent devices is built in an intranet environment, wherein the traffic collection device needs to be installed with a TCPDump packet capture tool to capture traffic. In addition, the traffic collection device needs to be connected to a gateway and a network switch. For Ethernet-connected intelligent devices, they need to be connected to the network switch. For wirelessly connected intelligent devices, they need to be connected to an intelligent switch (supporting WIFI, Zigbee, and Bluetooth protocols) and then connected to the network switch. In this way, the platform can capture the traffic of wired and wireless devices at the same time in front of the gateway. At the same time, since it is in an intranet environment, the MAC address can be used to mark the intelligent device. The marking content includes the type, manufacturer, and model of the intelligent device corresponding to the MAC address in the traffic, forming the intelligent device traffic data (i.e., pcap file) and its corresponding marking content. Optionally, in this embodiment, an authoritative open source data set is used to expand the traffic data to improve the multi-source and diversity of the data set.
[0036] S2. Preprocessing the labeled smart device traffic data;
[0037] Specifically, the bit stream of each data item is processed to form a universal representation of the traffic data. This representation is then converted into a two-dimensional traffic data graph, forming a two-dimensional traffic data graph and its corresponding annotations. Finally, based on the 80 / 20 principle, a random function is used to divide the data into training and test datasets, with the training dataset accounting for 80% and the test dataset accounting for 20%.
[0038] The first step is to process each piece of data in the smart device traffic by parsing the IP, TCP, UDP, ICMP, etc. layer by layer. The header field and payload content bits (containing only 0s and 1s) of each layer protocol are spliced together. The bits corresponding to the non-existent protocols are filled with -1 to obtain 4096 bits of data, and finally form a universal representation of the traffic data. Figure 2 The second step is to convert the generated flow data into a 64*64 matrix. A mapping is established for each value of the matrix: -1 becomes 255, 0 becomes 127, and 1 becomes 0. The mapping is then transformed into a grayscale matrix. The grayscale matrix is saved as an image, which is a two-dimensional flow data graph. Figure 3, forming a two-dimensional traffic data graph and its corresponding annotations. Since smart device traffic can vary somewhat in images, mapping traffic data to the image domain allows computer vision algorithms to analyze and process the images, extracting useful feature information, thereby enabling smart device classification and identification, improving classification accuracy and robustness. Finally, according to the 80 / 20 principle, a random function is used to divide the dataset into training and test datasets, with the training dataset accounting for 80% and the test dataset accounting for 20%.
[0039] S3. Construct a masked autoencoder;
[0040] For details, see Figure 4 , including a masker, encoder, and decoder. Its main features include: random masking strategy with high masking rate, and asymmetric encoder and decoder design.
[0041] The masker accepts a two-dimensional traffic data graph, divides the data graph into regular non-overlapping small patches, and then uses a sampling strategy that follows a uniform distribution to randomly sample a portion of these patches while masking the remaining portion (generally 75% of the two-dimensional traffic data graph is masked). It then outputs a mask matrix and visible patches.
[0042] The encoder processes the visible patches after masking, adds their corresponding position encoding vectors, and feeds them into the Transformer Encoder structure to output an embedding vector.
[0043] The decoder processes the masked tokens formed by the embedding vector and the mask matrix, inputs them into the Transformer Decoder structure, and outputs the reconstructed two-dimensional traffic data graph.
[0044] S4. Perform self-supervised pre-training;
[0045] Specifically, the original two-dimensional traffic data graph in the dataset is passed through a masked autoencoder to form a reconstructed two-dimensional traffic data graph. The goal of self-supervised pre-training is to optimize the two two-dimensional traffic data graphs to minimize the loss.
[0046] For two 2D flow data images, only the pixel values of the masked part are used to calculate the loss. The error of the pixel values can be calculated using the mean square error (MSE). For two n×n single-channel images A and B, their mean square error loss can be defined as:
[0047]
[0048] Where A(i, j) and B(i, j) are the pixel values at position (i, j) of images A and B. The mean square error loss is used to self-supervise the pre-trained model, and the pre-trained model is finally obtained.
[0049] S5. Fine-tune the downstream task. Specifically, a softmax layer is added after the encoder of the masked autoencoder. The softmax layer consists of a fully connected layer and a softmax function. The fully connected layer is used to reduce the dimension of the embedding vector output by the encoder to the dimension of the annotation content (i.e., the number of annotations). The softmax function converts the reduced dimension vector into a probability distribution, so that each value ranges between (0, 1) and the sum of all values is equal to 1. After the softmax layer, the annotation content with the highest predicted probability is obtained.
[0050] At the same time, since the masked autoencoder model has been pre-trained, only fine-tuning the downstream tasks is required to obtain better results and better speed. Therefore, the training strategy used is to fix some parameters of the encoder in the pre-trained masked autoencoder (the encoder contains 12 block layers) and only adjust the parameters of the last block layer.
[0051] The cross entropy loss function is used to calculate the loss for the original annotation content and the predicted annotation content, and some parameters of the encoder are fine-tuned. The cross entropy function is:
[0052]
[0053]
[0054] Where N represents the sample, C represents the number of smart device types, l represents the embedding vector output by the encoder, and p(c j |x i ) indicates whether the i-th sample belongs to the j-th type, q(c j |x i ) represents the probability that the model predicts that sample i belongs to type j.
[0055] The divided training and test datasets with annotated content are fed into the encoder and softmax layer of the pre-trained masked autoencoder model. The model parameters are adjusted through the training strategy to minimize the cross entropy loss on the training and validation sets, thereby obtaining the smart device recognition model.
[0056] S6. Perform smart device identification;
[0057] Specifically, for the original traffic generated by smart devices in practice, a two-dimensional traffic data graph is obtained through step S2 and input into the smart device identification model. The model will output a probability vector containing the probability of each smart device type. Based on this probability vector, the maximum probability and its corresponding type c are obtained. j , and the maximum probability is greater than the threshold, then the smart device type can be identified as c j If the maximum probability is less than or equal to the threshold, the smart device type can be identified as an unknown device type.
[0058] The embodiment of the present invention provides a smart device identification device based on mask self-encoding, which is applicable to the methods provided in the above embodiments. Figure 5 The functional module diagram of the masked self-encoding-based intelligent device identification device shown in the figure includes a data acquisition module, a data preprocessing model, a model construction module, a model pre-training module, a model fine-tuning module, and a model identification module. The device executes the embodiment based on the method described in S1 to S6. The specific uses of the modules are:
[0059] Data collection module, used to collect and annotate smart device traffic;
[0060] The data preprocessing module is used to process each bit stream of the traffic, form a general representation of the traffic data, and convert it into a two-dimensional traffic data graph, and finally divide it into training data sets and test data sets;
[0061] Model construction module, used to build the network structure of the masked autoencoder, including mask, encoder, decoder, etc.
[0062] Model pre-training module, used to build a masked autoencoder pre-training model based on self-supervision;
[0063] The model fine-tuning module is used to apply the pre-trained model to actual downstream tasks and fine-tune the smart device recognition model;
[0064] The model recognition module is used to identify the type of smart devices in actual smart device traffic.
[0065] While specific embodiments of the present invention have been disclosed for illustrative purposes, intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the disclosure of the preferred embodiments, and the scope of protection claimed in the present invention shall be determined by the scope of the claims.
Claims
1. A method for identifying smart devices based on a masked autoencoder, comprising the following steps: 1) Collect and annotate the traffic of smart devices to obtain the traffic data of smart devices and their corresponding annotation content; 2) converting each bit stream in the smart device traffic data into a one- or two-dimensional traffic data graph, taking each two-dimensional traffic data graph and its corresponding annotation content as a sample; and dividing the obtained samples into a training data set and a test data set; 3) Constructing a masked autoencoder, comprising a masker, an encoder, and a decoder; wherein the masker is used to mask the input two-dimensional traffic data graph to form a partially visible data graph; the encoder serves as a feature extractor, used to map the partially visible data graph to generate an embedding vector; and the decoder serves as a feature reconstructor, used to reconstruct the two-dimensional traffic data graph based on the embedding vector; 4) inputting the two-dimensional flow data graph in the training data set into a masked autoencoder to obtain a reconstructed two-dimensional flow data graph, and optimizing the masked autoencoder based on the loss of the original pixels and the corresponding reconstructed pixels in the masked part of the two-dimensional flow data graph; 5) Inputting the two-dimensional traffic data graph of sample i in the training dataset into the encoder of the masked autoencoder pre-trained and optimized in step 4) to obtain an embedding vector and inputting it into the softmax layer to obtain the annotation content with the maximum prediction probability for sample i; optimizing the encoder based on the loss between the annotation content with the maximum prediction probability for sample i and the annotation content of sample i; then verifying the optimized encoder using the test dataset, and using it as the smart device recognition model after passing the verification; 6) The intelligent identification traffic to be identified is converted into a two-dimensional traffic data graph and then input into the intelligent device identification model to identify the type of intelligent device corresponding to the intelligent identification traffic to be identified.
2. The method according to claim 1, characterized in that The method for obtaining the two-dimensional traffic data graph is as follows: performing layered protocol analysis on each bit stream, splicing the header field and payload content of each layer protocol together, filling the bits corresponding to non-existent protocols with -1, and forming a general representation of traffic data; then converting the general representation of traffic data into a matrix, establishing a mapping for each value of the matrix, converting the matrix into a grayscale matrix through mapping conversion, and saving the grayscale matrix as an image, which is the two-dimensional traffic data graph.
3. The method according to claim 1, characterized in that The encoder includes 12 Block layers, and the parameters of the last Block layer of the encoder are optimized according to the loss of the annotation content of the sample i and the maximum probability of the annotation content of the sample i.
4. The method according to claim 1, 2 or 3, characterized in that: In step 4), a mean square error loss is calculated based on the original pixels of the masked part and the corresponding reconstructed pixels in the two-dimensional flow data image, and the mask autoencoder is pre-trained and optimized by optimizing the mean square error loss.
5. The method according to claim 1, 2 or 3, characterized in that: In step 5), the encoder is optimized by calculating the loss using a cross entropy loss function based on the annotation content with the maximum probability predicted for the sample i and the annotation content of the sample i.
6. The method according to claim 1, 2 or 3, characterized in that: The masker masks the two-dimensional flow data graph at a ratio of 75% to form a partially visible data graph.
7. The method according to claim 1, 2 or 3, characterized in that: The annotation content includes: the smart device type, manufacturer and model corresponding to the MAC address in the traffic.
8. An intelligent device identification device based on a masked autoencoder, characterized in that: The data collection module is used to collect and annotate the traffic of smart devices to obtain the traffic data of smart devices and their corresponding annotation content; a data preprocessing module, configured to convert each bit stream in the smart device traffic data into a one- or two-dimensional traffic data graph, treating each two-dimensional traffic data graph and its corresponding annotation content as a sample; and dividing the obtained samples into a training data set and a test data set; A model construction module is used to construct a masked autoencoder, which includes a masker, an encoder, and a decoder. The masker is used to mask the input two-dimensional traffic data graph to form a partially visible data graph. The encoder acts as a feature extractor to map the partially visible data graph to generate an embedding vector. The decoder acts as a feature reconstructor to reconstruct the two-dimensional traffic data graph based on the embedding vector. A model pre-training module is used to input the two-dimensional flow data map in the training data set into a masked autoencoder to obtain a reconstructed two-dimensional flow data map, and optimize the masked autoencoder based on the loss of original pixels and corresponding reconstructed pixels in the masked part of the two-dimensional flow data map; A model fine-tuning module is configured to input the two-dimensional flow data graph of sample i in the training dataset into the encoder of the masked autoencoder pre-trained and optimized in step 4), obtain an embedding vector, and input the embedding vector into the softmax layer to obtain the annotation content with the maximum prediction probability for sample i; optimize the encoder based on the loss between the annotation content with the maximum prediction probability for sample i and the annotation content of sample i; then verify the optimized encoder using the test dataset, and use it as the smart device recognition model after passing the verification; The model identification module is used to convert the intelligent identification traffic to be identified into a two-dimensional traffic data graph and input it into the intelligent device identification model to identify the intelligent device type corresponding to the intelligent identification traffic to be identified.
9. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
X-ray image quality intelligent evaluation fusion model and use method
CN120876272A