Power Internet of Things equipment identification method and system based on multi-feature fusion

By collecting and fusion image and traffic data of power IoT devices, and using deep learning models to extract and fusion features, the problem of insufficient identification accuracy and robustness in the prior art is solved, and high accuracy and flexibility of device recognition is achieved.

CN119939337APending Publication Date: 2025-05-06GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411985874.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

Smart Images

  • Figure CN119939337A_ABST
    Figure CN119939337A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power Internet of Things equipment identification method and system based on multi-feature fusion. The method comprises the following steps: acquiring image and flow data of Internet of Things equipment in a power grid environment, marking the type of the equipment, marking the image and flow data acquired by the same equipment as the same power equipment type, and preprocessing the image and flow data; traffic time sequence features and traffic space features are extracted based on the preprocessed traffic data; extracting equipment image features based on the preprocessed image data; integrating the extracted features into a fusion feature vector by using a multi-modal fusion network to form an equipment fingerprint; a power Internet of Things equipment identification model is obtained based on fusion feature vector training; and using the electric power Internet of Things equipment identification model to identify the electric power Internet of Things equipment providing the equipment image and the flow data. According to the invention, through integrating various characteristics of the power equipment image and the traffic data, the accuracy and robustness of identification of the power Internet of Things equipment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power Internet of Things, and in particular to an electric power Internet of Things device identification method and system based on multi-feature fusion. Background Art

[0002] The power industry is promoting the construction of a new power system with new energy as the main body, and actively carrying out the construction of a global Internet of Things (IoT) platform, and proposing an overall solution for cloud-edge-end collaboration. With the gradual development of the full-link verification and implementation of cloud-edge-end collaboration, the IoT platforms of some companies have initially realized the unified access and management of professional terminals such as transmission, transformation and distribution, and have successfully connected more than 1 million terminals. With the rapid development of IoT technology, more and more smart terminal devices will be integrated into the power system, and these devices will exchange and communicate data through wireless networks. The diversity and complexity of power IoT devices require the system to accurately identify and distinguish them, which is crucial for network security, device management and personalized services.

[0003] Traditionally, IoT device identification relies mainly on the physical identifiers of the device, such as the MAC address. This method is simple and direct, but it is easy to be forged or altered, and it cannot provide in-depth information about the device's behavior. Another mainstream method is to identify the device by analyzing the network traffic characteristics of the device. This includes the analysis of the statistical characteristics of the traffic, the protocol type, and the traffic pattern. However, this method can usually only capture some of the behavioral characteristics of the device and is difficult to cope with the diversity and dynamic changes of device behavior. With the development of machine learning technology, some studies have begun to use machine learning algorithms, especially deep learning, to extract more complex features from network traffic to improve the accuracy of identification. These methods have improved the recognition performance to a certain extent, but most of them are limited to a single data source and fail to fully utilize the complementary advantages of multimodal data.

[0004] Whether it is a method based on physical identifiers or single network traffic features, it is difficult to fully capture the multi-dimensional characteristics of the device, resulting in insufficient recognition accuracy and robustness. In an adversarial environment, attackers can evade detection by forging or disguising identifiers, making recognition methods based on static features vulnerable. Although multimodal learning methods show potential, how to effectively fuse data from different modalities so that they can complement and enhance each other rather than simply pile them up remains a technical challenge. Summary of the invention

[0005] In view of this, the present invention provides a method for identifying power Internet of Things devices based on multi-feature fusion. The method collects and annotates the image data and flow data of devices in the power grid environment, uses multiple deep learning models to extract features from each modal data, and gradually fuses multimodal features in a hierarchical manner, ultimately achieving accurate identification of device types. The method of the present invention can effectively improve the accuracy of power Internet of Things device identification, and can adapt to device identification tasks in a variety of scenarios.

[0006] The present invention also provides an electric power Internet of Things device identification system based on multi-feature fusion, a computer device and a computer-readable storage medium.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] In a first aspect, a method for identifying power Internet of Things devices based on multi-feature fusion comprises the following steps:

[0009] Collect images and flow data of IoT devices in the power grid environment, mark the device type, mark the images and flow data collected from the same device as the same device type, and pre-process the images and flow data;

[0010] Based on the preprocessed traffic data, the traffic data is converted into decimal data, and the traffic time series features are extracted using a time series feature extraction network. Based on the preprocessed traffic data, a sliding window is used to extract the features within the window, and the traffic data is structured into a two-dimensional feature matrix. A convolutional neural network is used to perform a convolution operation on the two-dimensional feature matrix to obtain the local features of the traffic space.

[0011] Based on the preprocessed images, a convolutional neural network is used to extract device image features;

[0012] After aligning the traffic time series features, traffic spatial local features, and device image features, a multimodal fusion network is used to integrate them into a fused feature vector, and the fused feature vector is used to train an IoT device recognition model composed of a classifier;

[0013] Use the trained IoT device recognition model to identify IoT devices that provide device images and traffic data.

[0014] Furthermore, the image and flow data of IoT devices in the power grid environment are collected, including:

[0015] Acquire images of the device taken in a controlled environment, the images including images taken from different angles and distances under consistent lighting conditions and including images of different states of the device;

[0016] The traffic on the router accessed by the IoT device is obtained through the traffic capture tool, and the traffic is stored as a pcap file. When the traffic is collected, the IoT device is in an idle or active state; the traffic of each device is extracted from the captured traffic based on the device MAC address.

[0017] Further, the image and traffic data are preprocessed, including:

[0018] For image data, crop the image to a suitable size, convert the image pixel values ​​to floating point numbers and normalize them to [0, 1], and adjust the dimension order according to network requirements;

[0019] For traffic data, it is first cleaned to remove data packets without payload and empty packets and invalid traffic data. Then the cleaned traffic samples are split into five-tuples to generate multiple small-scale traffic files, and the transport layer payload of all data packets is extracted from each session.

[0020] Furthermore, based on the preprocessed traffic data, a sliding window is used to extract features within the window, and the traffic data is structured into a two-dimensional feature matrix, including:

[0021] Create time windows. The window size should balance the required data within the window and the responsiveness to changes in traffic patterns.

[0022] Slide the time window onto the traffic data and extract the features in each window. For each window, extract the packet size, arrival time interval, protocol distribution, IP address and port usage. Continue sliding the window at a specified step size to extract the features in each traffic data.

[0023] The traffic data is structured into a two-dimensional feature matrix, where the rows represent time windows, each window corresponds to a given moment or time interval, and each column corresponds to a feature extracted from the traffic. The size of the two-dimensional feature matrix is ​​(number of windows, number of features in each window).

[0024] Furthermore, after aligning the traffic time series features, traffic spatial local features, and device image features, a multimodal fusion network is used to integrate them into a fused feature vector, including:

[0025] Use the zero-fill method to align device image features, traffic space local features, and traffic time series features;

[0026] The aligned device image features, traffic spatial local features, and traffic temporal feature vectors are defined as single-modal nodes in the graph fusion network. A modal weight network is defined to process each node and calculate the weight of each node. The single-modal layer output vector is obtained by weighted summation based on the weights.

[0027] The unimodal node vectors are integrated in pairs through vector connection and input into the neural network to obtain the bimodal node vector. The weights of the bimodal nodes are determined according to the sum of the weights of the unimodal node vectors and the similarity between them. The bimodal layer output vector is obtained after weighted summation according to the weights.

[0028] Connect each bimodal node vector in the bimodal layer with all other unconnected single and bimodal node vectors, input them into the neural network to obtain trimodal nodes, determine the weight of the trimodal node according to the sum of the weights of the bimodal node vectors and the similarity between them, and obtain the output vector of the trimodal layer after weighted summation according to the weights;

[0029] The output vectors of the unimodal layer, bimodal layer, and trimodal layer are connected by vector connection to obtain the final fused feature vector.

[0030] Furthermore, in the unimodal layer, the weight calculation formula of each node is expressed as:

[0031] α i =σ(W mwn X i +b mwn )

[0032] Where σ is the activation function, W mwn and b mwn is the parameter of the modal weight network, X i is the eigenvector of mode i, a i is the weight of the feature vector of a single modal node, i∈{p1,p2,t}, where p1,p2,t correspond to device image features, traffic local spatial features, and traffic temporal features, respectively;

[0033] The calculation method of the bimodal node vector is expressed as:

[0034]

[0035] In the formula, represents vector connection, ΘNN is the parameter of the neural network, NN∈R 2d →R d It is a neural network that converts from 2D dimension to d dimension;

[0036] The calculation method of the trimodal node vector is expressed as:

[0037]

[0038] In the formula, i,j,m,n∈{p1,p2,t}, and When Y is empty, mn =X m .

[0039] Furthermore, the weight of the bimodal node is determined according to the sum of the weights of the two-way node vectors of the unimodal state and the similarity between them, which is expressed as:

[0040]

[0041] In the formula, a ij Represents a bimodal node Y ij The weight is based on The result after softmax normalization, S ij Represents a unimodal node vector X i With X j The similarity between them; C is a constant used to control the similarity and the relative importance of the weight of the unimodal node to the bimodal node;

[0042] The weight of the trimodal node is determined according to the sum of the weights of the bimodal node vectors and the similarity between them, which is expressed as:

[0043]

[0044] In the formula, a ijmn Represents a trimodal node O ijmn The weight is based on The result after softmax normalization, S ij_mn Represents the bimodal node vector Y ij With Y mn The similarities between.

[0045] In the second aspect, a power Internet of Things device identification system based on multi-feature fusion includes:

[0046] The data collection and preprocessing module is used to collect images and flow data of IoT devices in the power grid environment, mark the device type, mark the images and flow data collected from the same device as the same device type, and preprocess the images and flow data;

[0047] The traffic feature extraction module is used to convert the traffic data into decimal data based on the preprocessed traffic data, and use the time series feature extraction network to extract the traffic time series features; based on the preprocessed traffic data, use the sliding window to extract the features within the window, structure the traffic data into a two-dimensional feature matrix, and use the convolutional neural network to perform convolution operations on the two-dimensional feature matrix to obtain the local features of the traffic space;

[0048] An image feature extraction module, used to extract device image features using a convolutional neural network based on the preprocessed image;

[0049] The multi-feature fusion and recognition model training module is used to align the traffic time series features, traffic spatial local features, and device image features, and integrate them into a fused feature vector using a multimodal fusion network, and use the fused feature vector to train the IoT device recognition model composed of a classifier;

[0050] The IoT device identification module is used to use the trained IoT device identification model to identify the power IoT devices that provide device images and flow data.

[0051] In a third aspect, a computer device comprises: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the method for identifying electric power Internet of Things devices based on multi-feature fusion as described above are implemented.

[0052] In a fourth aspect, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for identifying electric power Internet of Things devices based on multi-feature fusion as described above.

[0053] Based on the above technical scheme, the beneficial effect of the present invention compared with the prior art is that by integrating the multimodal characteristics of the image and flow data of the power Internet of Things devices, the accuracy and robustness of the identification of the power Internet of Things devices are significantly improved. Using a hierarchical multimodal fusion network, this method effectively integrates features from different modalities to form a unique "device fingerprint" and enhances the feature expression capability. In addition, the application of diversified deep learning models fully utilizes the characteristics of each modal data. The present invention also has good adaptability and can adapt to device identification tasks in a variety of scenarios. It also has good scalability and is easy to expand to new device types and scenarios. In short, the present invention provides an innovative, flexible and accurate solution for IoT device identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of the power Internet of Things device identification method based on multi-feature fusion;

[0055] Figure 2 A schematic diagram of extracting features in each window in a sliding window based on traffic data;

[0056] Figure 3 It is the network structure of the multimodal fusion network. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings.

[0058] Reference Figure 1 , an embodiment of the present invention provides a method for identifying power Internet of Things devices based on multi-feature fusion, comprising the following steps:

[0059] Step S1: Collect images and flow data of IoT devices in the power grid environment.

[0060] For the power Internet of Things, the shape of the terminal equipment connected to the system is fixed in some cases, especially when it is necessary to comply with industry standards or installation specifications, which provides inspiration for image recognition through the shape of the equipment. In addition, the communication data generated during the working process of the equipment is a deep description of the intrinsic characteristics of the equipment. Therefore, the present invention constructs fusion features by combining images and traffic. Image feature extraction can capture the appearance attributes of the device, such as shape, color, texture, and size. These features help to distinguish different types of Internet of Things devices, especially when the devices have obvious differences in physical form. In the multi-feature fusion method, image features are combined with traffic features to provide a more comprehensive device description for training Internet of Things device recognition models. Since image features have a certain invariance to changes in illumination, changes in viewing angle, and occlusion, this makes the recognition model more robust.

[0061] According to an embodiment of the present invention, the collection of image data includes: obtaining device images taken in a controlled environment to ensure consistent lighting conditions, simple backgrounds, and reduced variables. Including device images taken from different angles and distances to capture the diversity of the device. If the device has multiple states (such as on / off), ensure that it is photographed in different states.

[0062] Traffic data collection, including:

[0063] (1) In a real power grid environment, use the router's tcpdump to capture the traffic passing through the router, and store the traffic as a pcap file on the local server for further processing. When collecting traffic, keep the IoT device in an idle or active state. Note that both states need to be collected. Traffic data in different states can fully understand the behavior of the device in the actual environment.

[0064] (2) Use the device's MAC address and the tcpreplay tool to extract the traffic of each device from the captured traffic. That is, based on the MAC address, extract the corresponding traffic of each MAC address.

[0065] The collected data is labeled with the device type. Note that the images and traffic data collected by the same device are labeled with the same device type. The purpose of labeling is to build a training data set for the following model learning.

[0066] Step S2: pre-process the collected images and flow data.

[0067] The preprocessing of image data includes image cropping, pixel value conversion and normalization, and adjustment of dimension requirements according to network requirements. In an embodiment of the present invention, the short side of the image is scaled to 256 pixels, the aspect ratio is kept unchanged, and a central area of ​​224×224 pixels is cropped from the scaled image for subsequent feature extraction. The image pixel values ​​are then converted to floating point numbers and normalized to [0,1], and the dimension order is adjusted to (C, H, W), where C represents the number of channels, representing the number of color channels in the image; H represents the height of the image, representing the number of pixel rows in the image; W represents the width of the image, representing the number of pixel columns in the image. The order of dimensions is adjusted to match the expected input dimension order of the convolutional neural network used later.

[0068] The preprocessing of traffic data includes:

[0069] (1) Clean the original traffic data in pcap format and remove invalid traffic data such as data packets without payload and empty packets to reduce the interference of noise data and improve the efficiency of traffic feature extraction.

[0070] (2) Use the SplitCap tool to split the cleaned large-scale traffic samples according to the five-tuple (source IP address, destination IP address, source port, destination port, and transport layer protocol) to generate multiple small-scale traffic files. This step helps to generate more fine-grained and information-rich session data.

[0071] (3) IoT devices usually have specialized communication protocols and data formats, and the payload data contains characteristic information carried by specific devices. The Scapy tool is used to extract the transport layer payload of all data packets from each session. This step requires fewer computing resources and does not rely on specific protocols or standards.

[0072] Step S3: extracting device image features based on the preprocessed device image.

[0073] Any convolutional neural network can be used to extract device image features. In the embodiment of the present invention, ResNet50 is used to extract device image features. ResNet50 is a widely used deep learning network. The specific extraction process is not described in detail in the present invention.

[0074] The present invention introduces device image data as a modality to supplement the physical characteristics of the device to address the problem of highly overlapping traffic characteristics when different devices have similar characteristics such as communication protocols, data packet sizes, and transmission frequencies.

[0075] Step S4: Extracting flow characteristics based on the preprocessed flow data.

[0076] According to the implementation mode of the present invention, since the flow data of the electric power Internet of Things device in the power grid environment has long-term dependence (that is, the current data point may be related to the data point in the past period of time), time series characteristics (that is, the flow data is arranged in chronological order), dynamic variability (that is, the flow pattern will change over time) and nonlinear complexity (that is, the flow data contains complex nonlinear relationships), based on the pre-processed flow data, a deep learning network (such as RNN, LSTM, etc., which can extract time series features) is used to extract the flow time series features, and the flow data is extracted through a sliding window to construct a two-dimensional feature matrix and extract the local features of the flow space. In an embodiment of the present invention, the pre-processed flow data is converted into decimal data, and the long short-term memory network LSTM is used to extract the time series features of the network flow. LSTM is a mature technology and will not be repeated here. The flow data is extracted through a sliding window to construct a two-dimensional feature matrix and extract the local features of the flow space, which specifically includes the following steps:

[0077] S401: Create a time window. The window size should balance the required data in the window and the responsiveness to changes in the traffic pattern. Shorter windows capture finer-grained patterns and are suitable for burst traffic. Longer windows are suitable for capturing more stable traffic or periodic patterns. In the embodiment of the present invention, the size of the time window is set to 3.

[0078] S402: Slide the time window onto the traffic data, and extract features from each traffic data using a sliding window with a step size of 1. Since the power Internet of Things terminal devices have specific behavior patterns in dimensions such as data packet size (for example, the size of the data packets sent and received may vary depending on their functions), interaction frequency (for example, many power Internet of Things terminal devices send data according to a predetermined cycle or event trigger), and protocol switching behavior (for example, some devices will switch to a low-power communication protocol when high-speed communication is not required to save energy), for each window, extract features such as data packet size, arrival time interval, protocol distribution, IP address, and port usage: Count the data packet size and create a histogram of the data packet size within the window; Calculate the mean of the data packet arrival time interval within the window; Calculate the frequency of different protocols used within the window (TCP / IP, DL / T645, Modbus, DNP3.0, IEC 60880-5, and wireless communication protocols such as ZigBee, LoRa, WiFi, and 5G); Count the number of occurrences of IP and port. The process of sliding window calculation is as follows: Figure 2 shown.

[0079] S403: The traffic data is structured into a two-dimensional feature matrix, where rows represent time windows, each window corresponds to a specific moment or time interval, and each column corresponds to a feature extracted from the traffic. The size of the two-dimensional feature matrix is ​​(number of windows, number of features in each window).

[0080] S404: Adjust the ResNet18 network to adapt it to one-dimensional convolution to form ResNet18-1D, and use it to perform convolution operation on the two-dimensional feature matrix obtained in S403 to capture the local features of the flow space.

[0081] Spatial local features use a sliding window method to analyze the local changes in traffic data in different time periods and dimensions, revealing the relationship between short-term traffic patterns and dimensions. LSTM is good at capturing the temporal dependencies in traffic data and can effectively identify the long-term behavior patterns of devices, especially with strong sensitivity to periodic or bursty traffic. Combining these two feature extraction methods can improve the recognition accuracy and robustness of the model in complex and changing network environments.

[0082] Step S5: Multimodal feature fusion, through a specially designed multimodal graph fusion network, the key features of different modalities are integrated into a fused feature vector to form a "device fingerprint".

[0083] The feature fusion process specifically includes:

[0084] S501: Use the zero-completion method to align device image features, traffic space local features, and traffic time series features.

[0085] S502: Define the aligned device image features, traffic spatial local features, and traffic temporal feature vectors as single-modal nodes in the graph fusion network, and define the modal weight network MWN∈R d →R 1 Processing each node, that is, the MWN inputs a d-dimensional feature vector and outputs a weight value (scalar, dimension 1) to control the weight of each node. Here, the modal weight network consists of a fully connected layer activated by a parameterized sigmoid activation function, and finally outputs the node weight vector, which can also be a neural network of any other structure. The following is the formula for calculating the weight of a single modal node:

[0086] α i =σ(W mwn X i +b mwn )

[0087] Where σ is the activation function, W mwn and b mwn is the parameter of the modal weight network, X iis the eigenvector of mode i, a i is the weight of a single modal feature vector i∈{p1,p2,t}, where p1,p2,t correspond to device image features, traffic spatial local features, and traffic temporal features, respectively.

[0088] S503: Calculate the output of the unimodal layer, which is the vector U after weighted summation using the weight parameters in step S502:

[0089]

[0090] S504: Integrate the single-mode node vectors by connecting them two by two and input them into NN∈R 2d →R d Network, NN represents neural network, NN∈R 2d →R d Indicates that the NN inputs a 2d-dimensional feature vector and outputs a d-dimensional feature vector. In the embodiment of the present invention, the NN is implemented using a multilayer perceptron, and the output of the network is a bimodal node vector. Other networks can also be used. The links generated between the two unimodal nodes and the corresponding bimodal nodes are shown in the attached figure. Figure 3 As shown in the link between the first and second layers, the calculation formula of the bimodal node vector is as follows:

[0091]

[0092] in represents vector connection, and ΘNN is the parameter of the neural network.

[0093] S505: Calculate the softmax normalized vector of the unimodal node vector, and use the normalized vector to estimate the similarity between each pair. The similarity estimation method can be any method for calculating vector similarity, such as cosine similarity, Jaccard similarity, etc. The calculation formula of the similarity estimation method here is as follows:

[0094]

[0095] in Yes X i The softmax normalized vector of .

[0096] S506: Calculate bimodal node Y based on similarity ij The weight a ij :

[0097]

[0098] The constant C here is used to control the similarity and the relative importance of the weight of the unimodal node to the bimodal node. In the embodiment, C = 0.5, a ij yes The normalized result.

[0099] S507: The output of the bimodal layer is a vector obtained by weighting the bimodal nodes using the weight parameters in step S506, which is calculated by the following formula:

[0100]

[0101] S508: In the bimodal layer, each bimodal node vector is connected to all other unconnected single and bimodal node vectors and input into the neural network NN∈R 2d →R d Get the trimodal node, such as Figure 3 As shown, the first trimodal node of the trimodal layer Bimodal Node Vector and The vectors are connected and input into the NN network to obtain the last node By bimodal node vector and the unimodal node vector The vectors are connected and then input into the NN network to obtain. Here, the NN is implemented using a multi-layer perceptron. Other networks can also be used. The nodes corresponding to the two connected vectors and the links generated by the generated trimodal nodes are shown in the attached figure. Figure 3 The second and third links are shown as follows, where usually i,j,m,n∈{p1,p2,t}, when assuming that When Y is empty, mn =X m :

[0102]

[0103] S509: Calculate the weight of each node in the trimodal layer according to the process of S505, S506, and S507, and perform weighted summation on the node vectors output by S508 to obtain the output vector of the trimodal layer. The calculation formula is as follows: usually i, j, m, n ∈ {p1, p2, t}, when assuming that When Y is empty, mn =X n ,a mn =a n , C is a constant, take 0.5:

[0104]

[0105] S510: The final fusion feature F, i.e., “device fingerprint”, is obtained by vector connection. Figure 3 The complete structure of the image fusion network is shown:

[0106]

[0107] S6: Classification and model training: Use the classifier to classify and identify IoT devices based on the fused feature vector, and train it on the data preprocessed by S2 to obtain the IoT device recognition model.

[0108] Because the neural network parameters for feature extraction and feature fusion also need to be trained, the IoT device recognition model here refers to an end-to-end network model that includes image feature and traffic feature extraction networks, multimodal feature fusion networks, and classifier networks. The training process is not the focus of the invention and will not be repeated here.

[0109] The information of different modalities is gradually fused in a hierarchical manner, which allows the model to capture the interaction between modalities at different levels of abstraction, improves the expressiveness of the fused features, and thus improves the robustness and accuracy of recognition.

[0110] S7: Use the IoT device recognition model trained in step S6 to identify the power IoT device that provides device images and flow data. Before recognition, it is necessary to preprocess the image and flow data with reference to step S2. The model will perform feature extraction and feature fusion on the preprocessed device images and flow data, and finally classify and output the determined device type recognition result.

[0111] The above steps show the specific implementation mode of the present invention. Figure 1 The overall architecture of the present invention is demonstrated. The present invention improves the accuracy and robustness of IoT device identification by integrating the multimodal characteristics of image and traffic data. Using a special graph fusion network technology, the method effectively integrates features from different modalities to form a unique "device fingerprint" and enhances the feature expression capability. In addition, the application of diverse deep learning models fully utilizes the characteristics of each modality of data. The present invention also has good adaptability and can adapt to device identification tasks in a variety of scenarios. It also has good scalability and can be easily expanded to new device types and scenarios. In short, the present invention provides an innovative, flexible and accurate IoT device identification solution.

[0112] It should be noted that the method described in this embodiment is not the only one and can be adjusted and optimized according to actual needs. For example, different deep learning models can be replaced to adapt to specific data characteristics, or feature fusion strategies can be adjusted to improve recognition performance. Those skilled in the art can make various improvements and modifications based on the spirit and scope of the present invention, and these improvements and modifications should also be regarded as the scope of protection of the present invention.

[0113] Based on the same technical concept as the method embodiment, the embodiment of the present invention also provides a power Internet of Things device identification system based on multi-feature fusion, which mainly includes:

[0114] The data collection and preprocessing module is used to collect images and flow data of IoT devices in the power grid environment, mark the device type, mark the images and flow data collected from the same device as the same device type, and preprocess the images and flow data;

[0115] The traffic feature extraction module is used to convert the traffic data into decimal data based on the preprocessed traffic data, and use the time series feature extraction network to extract the traffic time series features; based on the preprocessed traffic data, use the sliding window to extract the features within the window, structure the traffic data into a two-dimensional feature matrix, and use the convolutional neural network to perform convolution operations on the two-dimensional feature matrix to obtain the local features of the traffic space;

[0116] An image feature extraction module, used to extract device image features using a convolutional neural network based on the preprocessed image;

[0117] The multi-feature fusion and recognition model training module is used to align the traffic time series features, traffic spatial local features, and device image features, and integrate them into a fused feature vector using a multimodal fusion network, and use the fused feature vector to train the IoT device recognition model composed of a classifier;

[0118] The IoT device identification module is used to use the trained IoT device identification model to identify the power IoT devices that provide device images and flow data.

[0119] It should be understood that the electric power Internet of Things device identification system based on multi-feature fusion in the embodiment of the present invention can implement all the technical solutions in the above-mentioned method embodiment, and the functions of its various functional modules can be specifically implemented according to the methods in the above-mentioned method embodiments. The specific implementation process can refer to the relevant description in the above-mentioned embodiments, which will not be repeated here.

[0120] The present invention also provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the method for identifying electric power Internet of Things devices based on multi-feature fusion as described above are implemented.

[0121] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for identifying electric power Internet of Things devices based on multi-feature fusion as described above are implemented.

[0122] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, devices (systems), computer equipment or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0123] The present invention is described with reference to a flowchart of a method according to an embodiment of the present invention. It should be understood that each process in the flowchart and a combination of processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A device that specifies functions in a process or multiple processes.

[0124] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A function specified in a process or multiple processes.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.

Claims

1. A method for identifying power Internet of Things devices based on multi-feature fusion, characterized in that: The following steps are involved: Collect images and flow data of IoT devices in the power grid environment, mark the device type, mark the images and flow data collected from the same device as the same device type, and pre-process the images and flow data; Based on the preprocessed traffic data, the traffic data is converted into decimal data, and the traffic time series features are extracted using a time series feature extraction network; Based on the preprocessed traffic data, a sliding window is used to extract the features within the window, and the traffic data is structured into a two-dimensional feature matrix. A convolutional neural network is used to perform a convolution operation on the two-dimensional feature matrix to obtain the local features of the traffic space. Based on the preprocessed images, a convolutional neural network is used to extract device image features; After aligning the traffic time series features, traffic spatial local features, and device image features, a multimodal fusion network is used to integrate them into a fused feature vector, and the fused feature vector is used to train an IoT device recognition model composed of a classifier; Use the trained IoT device recognition model to identify power IoT devices that provide device images and flow data.

2. The method according to claim 1, characterized in that: Collect images and traffic data of IoT devices in the power grid environment, including: Acquire images of the device taken in a controlled environment, the images including images taken from different angles and distances under consistent lighting conditions and including images of different states of the device; The traffic on the router accessed by the IoT device is obtained through the traffic capture tool, and the traffic is stored as a pcap file. When the traffic is collected, the IoT device is in an idle or active state; the traffic of each device is extracted from the captured traffic based on the device MAC address.

3. The method according to claim 1, characterized in that Preprocessing of images and traffic data, including: For image data, crop the image to a suitable size, convert the image pixel values ​​to floating point numbers and normalize them to [0, 1], and adjust the dimension order according to network requirements; For traffic data, it is first cleaned to remove data packets without payload and empty packets and invalid traffic data. Then the cleaned traffic samples are split into five-tuples to generate multiple small-scale traffic files, and the transport layer payload of all data packets is extracted from each session.

4. The method according to claim 1, characterized in that: Based on the preprocessed traffic data, a sliding window is used to extract the features within the window, and the traffic data is structured into a two-dimensional feature matrix, including: Create time windows. The window size should balance the required data within the window and the responsiveness to changes in traffic patterns. Slide the time window onto the traffic data and extract the features in each window. For each window, extract the packet size, arrival time interval, protocol distribution, IP address and port usage. Continue sliding the window at a specified step size to extract the features in each traffic data. The traffic data is structured into a two-dimensional feature matrix, where the rows represent time windows, each window corresponds to a given moment or time interval, and each column corresponds to a feature extracted from the traffic. The size of the two-dimensional feature matrix is ​​(number of windows, number of features in each window).

5. The method according to claim 1, characterized in that After aligning the traffic time series features, traffic spatial local features, and device image features, a multimodal fusion network is used to integrate them into a fused feature vector, including: Use the zero-fill method to align device image features, traffic space local features, and traffic time series features; The aligned device image features, traffic spatial local features, and traffic temporal feature vectors are defined as single-modal nodes in the graph fusion network. A modal weight network is defined to process each node and calculate the weight of each node. The single-modal layer output vector is obtained by weighted summation based on the weights. The unimodal node vectors are integrated in pairs through vector connection and input into the neural network to obtain the bimodal node vector. The weights of the bimodal nodes are determined according to the sum of the weights of the unimodal node vectors and the similarity between them. The bimodal layer output vector is obtained after weighted summation according to the weights. Connect each bimodal node vector in the bimodal layer with all other unconnected single and bimodal node vectors, input them into the neural network to obtain trimodal nodes, determine the weight of the trimodal node according to the sum of the weights of the bimodal node vectors and the similarity between them, and obtain the output vector of the trimodal layer after weighted summation according to the weights; The output vectors of the unimodal layer, bimodal layer, and trimodal layer are connected by vector connection to obtain the final fused feature vector.

6. The method according to claim 5, characterized in that In the unimodal layer, the weight calculation formula for each node is expressed as: a i =σ(W mwn X i +b mwn ) Where σ is the activation function, W mwn and b mwn is the parameter of the modal weight network, X i is the eigenvector of mode i, a i is the weight of the feature vector of a single modal node, i∈{p1,p2,t}, where p1,p2,t correspond to device image features, traffic local spatial features, and traffic temporal features, respectively; The calculation method of the bimodal node vector is expressed as: In the formula, represents vector connection, ΘNN is the parameter of the neural network, NN∈R 2d →R d It is a neural network that converts from 2D dimension to d dimension; The calculation method of the trimodal node vector is expressed as: In the formula, i,j,m,n∈{p1,p2,t}, and When Y is empty, mn =X m .

7. The method according to claim 6, characterized in that The weight of the bimodal node is determined according to the sum of the weights of the two-way node vectors of the unimodal and the similarity between them, which is expressed as: In the formula, a ij Represents a bimodal node Y ij The weight is based on The result after softmax normalization, S ij Represents a unimodal node vector X i With X j The similarity between them; C is a constant used to control the similarity and the relative importance of the weight of the unimodal node to the bimodal node; The weight of the trimodal node is determined according to the sum of the weights of the bimodal node vectors and the similarity between them, which is expressed as: In the formula, a ijmn Represents a trimodal node O ijmn The weight is based on The result after softmax normalization, S ij_mn Represents the bimodal node vector Y ij With Y mn The similarities between.

8. A power Internet of Things equipment identification system based on multi-feature fusion, characterized in that: include: The data collection and preprocessing module is used to collect images and flow data of IoT devices in the power grid environment, mark the device type, mark the images and flow data collected from the same device as the same device type, and preprocess the images and flow data; A traffic feature extraction module is used to convert the traffic data into decimal data based on the preprocessed traffic data, and extract the traffic time series features using a time series feature extraction network; Based on the preprocessed traffic data, a sliding window is used to extract the features within the window, and the traffic data is structured into a two-dimensional feature matrix. A convolutional neural network is used to perform a convolution operation on the two-dimensional feature matrix to obtain the local features of the traffic space. An image feature extraction module, used to extract device image features using a convolutional neural network based on the preprocessed image; The multi-feature fusion and recognition model training module is used to align the traffic time series features, traffic spatial local features, and device image features, and integrate them into a fused feature vector using a multimodal fusion network, and use the fused feature vector to train the IoT device recognition model composed of a classifier; The IoT device identification module is used to use the trained IoT device identification model to identify the power IoT devices that provide device images and flow data.

9. A computer device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the power Internet of Things device identification method based on multi-feature fusion as described in any one of claims 1-7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for identifying electric power Internet of Things devices based on multi-feature fusion as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Hidden scanning behavior recognition method, device, equipment, medium and program

    CN120415909A

  • Concealed scanning behavior identification method, device, equipment, medium and program

    CN120415909B

  • Communication equipment identification method and electronic equipment

    CN120915691A