Multimedia Internet of Things anomaly detection method and device based on tensor multi-feature fusion

By adopting a tensor multi-feature fusion method in IoMT anomaly detection, using graph attention network and tensor product technology, the problem of sensor connection and data association being not fully considered is solved, and the accuracy of anomaly detection and the robustness of the model are improved.

CN119939319APending Publication Date: 2025-05-06HUBEI CHUTIAN SMART COMM CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510117628.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing IoMT time series anomaly detection methods fail to fully consider the relationship between sensors and the association between data, resulting in a low accuracy of abnormal detection.

Method used

Using a method based on tensor multi-feature fusion, multiple sensor features are extracted through graph attention network and tensor product technology, and graph tensors are constructed for low-tube rank approximation to identify abnormal time points in the multimedia Internet of Things.

Benefits of technology

By fully digging out the dependencies between sensors, the accuracy of anomaly detection is improved, and the robustness and adaptability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939319A_ABST
    Figure CN119939319A_ABST
Patent Text Reader

Abstract

The invention discloses a multimedia Internet of Things anomaly detection method and device based on tensor multi-feature fusion. The method comprises the following steps: acquiring time sequence historical data of the multimedia Internet of Things through a sensor; inputting the time sequence historical data into the graph attention network to obtain sensor features; wherein the step of inputting the time sequence historical data into the graph attention network to obtain the sensor features comprises the following steps: respectively inputting the time sequence historical data into a plurality of feature extractors to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating low tube rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through a tensor product, and outputting the sensor features of the current time sequence based on the tensor attention network; and identifying the abnormal time point of the multimedia Internet of Things based on the sensor characteristics by using a graph deviation scoring method. According to the invention, the multimedia Internet of Things anomaly detection accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multimedia Internet of Things data detection, and in particular to a multimedia Internet of Things anomaly detection method, device, storage medium and electronic device based on tensor multi-feature fusion. Background Art

[0002] With the rapid development of the Internet of Things (IoT), our cities are equipped with a large number of devices, each of which can collect information related to the status of the IoT in real time and generate time series to accurately reflect these statuses. In addition, with the digitization of multimedia data, various types of multimedia data such as images, videos, and audio can also be converted into time series for better transmission, analysis, and storage. More importantly, multimedia data can be integrated with the Internet of Things through the "bridge" of time series to form the Multimedia Internet of Things (IoMT).

[0003] The time series of all multimedia data generated from the IoMT system can be integrated into a multivariate time series (MTS). However, anomalies may occur in the MTS due to the damage of some IoMT facilities or attacks during data transmission. We call those data that are significantly different from normal MTS data anomalies. In order to ensure the security and reliability of the IoMT system, MTS anomaly detection becomes a crucial research task.

[0004] In recent years, many AI-based MTS anomaly detection works have emerged, which can be divided into two categories: prediction-based and reconstruction-based. Prediction-based models mainly use neural networks such as CNN and RNN to learn the inherent characteristics of normal data from historical data and predict future MTS data values. If the predicted value is far from the observed value at a certain point, we regard it as an anomaly. Reconstruction-based models mainly use generative AI such as Transformer and GAN to reconstruct the entire MTS, and detect anomalies by comparing the observed MTS with the reconstructed MTS. If the reconstructed value at a certain point deviates greatly from the observed value, it is regarded as an anomaly.

[0005] There are many sensors in IoMT, and the connections between them are intricate. The current IoMT time series anomaly detection methods do not consider the connections between sensors and the associations between data, resulting in low anomaly detection accuracy. Summary of the invention

[0006] The embodiments of the present application provide a multimedia Internet of Things anomaly detection method, device, storage medium and electronic device based on tensor multi-feature fusion, which can fully explore the dependency relationship between sensors and improve the accuracy of anomaly detection.

[0007] The embodiment of the present application provides a multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, including: Obtaining time series historical data of multimedia IoT through sensors; Inputting the time series historical data into a graph attention network to obtain sensor features; wherein, inputting the time series historical data into a graph attention network to obtain sensor features includes: Inputting the time series historical data into multiple feature extractors respectively to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating a low-rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through tensor product, and outputting the sensor features of the current time series based on the tensor attention network; The graph deviation scoring method is used to identify the time points of multimedia IoT anomalies based on the sensor features.

[0008] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the multiple feature extractors include sensor embedding vectors, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The time series historical data is converted into a sensor embedding vector, and the sensor embedding vector is:

[0009] in, Represents time series historical data, represents the sensor embedding vector, Represents the dimension of the sensor embedding vector.

[0010] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the multiple feature extractors include a local feature extractor, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted through the linear layer to obtain the local features of the sensor:

[0011] in, is the input time series historical data, is the local feature of the sensor, Indicates mapping the input data to dimensional space, are its parameters.

[0012] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the multiple feature extractors include a time feature extractor, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted through a recurrent neural network to obtain a time feature graph:

[0013]

[0014] in, is a linear layer used to obtain the initial state of the recurrent neural network. are the parameters of the recurrent neural network, Represents the initial state of the recurrent neural network.

[0015] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the multiple feature extractors include a multi-scale feature extractor, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted by dilated convolution to obtain multi-scale features: For sensor i, define -Dilated convolution operation for:

[0016] in, is the time series generated by sensor i, is a 1D convolution filter kernel of size 1×k, d is the dilation factor; Stacking There are dilated convolution layers, and the dilated convolution used is:

[0017]

[0018]

[0019]

[0020] Where i=1, 2, …, N, j=1,2,…, represents the input time series historical data with sliding window size w generated by sensor i, and are two parallel dilated convolution operations, is the multi-scale feature of sensor i in the j-th dilated convolutional layer, are c 1D convolution kernels of different sizes, Indicates that the feature is truncated to dimensions, and then connect them together. is the input to the dilated convolutional layer of sensor i in layer j.

[0021] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the step of constructing a graph structure based on the initial features and converting the graph structure into a graph tensor comprises: A graph structure is constructed based on the initial features, and the graph structure is defined as: ,i=1,2,…, ;in, Represents a node, represents the edge, and the sensor is regarded as a node. Indicates that there is an edge from node i to node j; Calculating similarities between sensors based on the initial features, and when the similarities exceed a preset similarity threshold, establishing edges between corresponding sensors; The weighted adjacency matrices corresponding to the graph structure are concatenated into a graph tensor.

[0022] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the low-rank r approximation of the graph tensor is calculated based on the graph tensor, including: Use the tensor singular value decomposition method to decompose the graph tensor Perform a low-rank r approximation:

[0023] Among them, r represents The management order of " is the tensor product operation, as well as Respectively represent and The first r transverse slices of It only contains The diagonal tensor of the first r tensor tubes of , is a tensor after the low-rank r approximation based on T-SVD.

[0024] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the graph attention network is expanded into a tensor attention network through tensor product, and the sensor features of the current time series are output based on the tensor attention network, including: Use matrix product to represent a two-layer graph attention network:

[0025]

[0026] in, is a parameter set, is the weighted adjacency matrix, and is a weight matrix that can be learned along with the training process, represents a nonlinear activation function; The graph attention network is extended to the tensor attention network using the tensor product operation:

[0027]

[0028]

[0029] in, represents the tensor product operation, represents the number of neural network layers of the tensor attention network, is the feature tensor, Represents the weight tensor of the cth layer ( , ,), H is the feature dimension of the hidden layer of the tensor attention network, represents a nonlinear activation function, which is a weight matrix used to map a three-dimensional tensor to two-dimensional sensor features. is the parameter set containing all parameters of layer c, The final sensor features representing the output of the model.

[0030] Furthermore, in the above-mentioned multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, the time point of identifying the multimedia Internet of Things anomaly based on the sensor features using the graph deviation scoring method includes: Predicting a time series at a current moment based on the sensor feature and the sensor embedding vector; The anomaly score at the current moment is calculated, and if the anomaly score at the current moment is greater than the anomaly threshold, it is determined that the multimedia Internet of Things state at the current moment is abnormal.

[0031] The embodiment of the present application also provides a multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion, including: An acquisition module is used to acquire the time series historical data of the multimedia Internet of Things through sensors; A feature extraction module, used for inputting the time series historical data into a graph attention network to obtain sensor features; wherein the step of inputting the time series historical data into a graph attention network to obtain sensor features includes: Inputting the time series historical data into multiple feature extractors respectively to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating a low-rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through tensor product, and outputting the sensor features of the current time series based on the tensor attention network; The anomaly identification module is used to identify the time point of multimedia Internet of Things anomalies based on the sensor features using a graph deviation scoring method.

[0032] An embodiment of the present application also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are suitable for being loaded by a processor to execute any of the above-mentioned multimedia Internet of Things anomaly detection methods based on tensor multi-feature fusion.

[0033] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion as described in any of the above items.

[0034] The present application provides a multimedia Internet of Things anomaly detection method, device, storage medium and electronic device based on tensor multi-feature fusion. The present application extracts sensor features in a time series from multiple data analysis perspectives through multiple feature extractors, and thereby mines the dependency between multiple sensors. These feature extractors cover almost all data analysis perspectives of the time series, and can fully mine the dependency between sensors and improve the accuracy of model anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The technical solution and other beneficial effects of the present application will be made apparent by describing in detail the specific implementation methods of the present application in conjunction with the accompanying drawings.

[0036] Figure 1 A flowchart of a multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in an embodiment of the present application.

[0037] Figure 2 Another flowchart of the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in an embodiment of the present application.

[0038] Figure 3 A schematic diagram of values ​​generated within a certain time period provided in an embodiment of the present application.

[0039] Figure 4 This is a flow chart of feature extraction of a multi-scale feature extractor provided in an embodiment of the present application, where DCL represents a dilated convolutional layer.

[0040] Figure 5 A schematic diagram of the structure of the dilated convolutional layer provided in an embodiment of the present application.

[0041] Figure 6 t-SNE schematic diagram of the sensor embedding vector provided in an embodiment of the present application.

[0042] Figure 7 t-SNE schematic diagram of the local features of the sensor provided in the embodiment of the present application.

[0043] Figure 8 A t-SNE schematic diagram of the sensor time characteristics provided in an embodiment of the present application.

[0044] Fig. 9 t-SNE schematic diagram of the multi-scale features of the sensor provided in the embodiment of the present application.

[0045] Fig.10 A t-SNE diagram of the final sensor features output by the tensor attention network provided in an embodiment of the present application.

[0046] Fig.11 A schematic diagram of the model performance of different feature extractors provided in the embodiments of the present application when constructing tensors.

[0047] Fig.12 A schematic diagram of sensitivity analysis of parameter r on 6 data sets provided in an embodiment of the present application.

[0048] Fig.13 A schematic diagram of sensitivity analysis of parameter k on 6 data sets provided in an embodiment of the present application.

[0049] Fig.14 A schematic diagram of the model performance on 6 data sets when using TGAT with different numbers of layers provided in an embodiment of the present application.

[0050] Fig.15 Detailed information statistics of the six data sets provided in the embodiments of the present application.

[0051] Fig.16 A schematic diagram of comparative experimental results of anomaly detection on 6 data sets using Pre (Precision, %), Rec (Recall rate, %), and F1 (F1-score) indicators provided in an embodiment of the present application.

[0052] Fig.17The ablation experiment results on 6 data sets are provided in the embodiments of the present application.

[0053] Fig.18 A schematic diagram of the structure of a multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion provided in an embodiment of the present application.

[0054] Fig.19 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0056] The current difficulties in detecting anomalies in IoMT time series are mainly reflected in four aspects. First, there are a large number of sensors in IoMT, and the connections between them are intricate. These sensors are distributed in various devices, systems and environments, and their tasks cover a variety of functions such as monitoring environmental parameters, collecting data and communicating with other devices. However, these sensors do not exist in isolation, but are interrelated and influence each other. They may work together to perform specific tasks, or they may rely on each other to provide accurate data or trigger specific operations; second, different types of multimedia data are multimodal data, which leads to the IoMT time series data converted from multimedia data containing a large amount of heterogeneous information. When processing these time series data, the existence of these heterogeneous information must be taken into account; third, the existing graph neural network models are only suitable for processing a single graph, and do not have sufficient ability to perform graph neural network learning in multidimensional tensors converted from multiple graph structures; fourth, the scarcity or unavailability of anomaly labels. In short, we lack sufficient anomaly samples when training the model or anomaly samples are difficult to obtain, which will bring a series of problems to the accuracy and reliability of IoMT anomaly detection.

[0057] To solve the above problems, the embodiments of the present application provide a multimedia Internet of Things anomaly detection method, device, storage medium and electronic device based on tensor multi-feature fusion. The multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion provided in the embodiments of the present application can be integrated in an electronic device, which can be a terminal, a server and other devices, wherein the terminal can include a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a micro processing box, or other devices.

[0058] The following first explains and illustrates the technical terms of the present invention: Slide Window (SW): Sliding window is a common algorithmic technique. The basic idea is to maintain a fixed-size window that moves over the sequence and, at each step, perform specific operations on the elements within the window, and then adjust the position of the window according to the requirements of the problem.

[0059] Sensor Embedding (SE): Sensor embedding refers to converting sensor data into an embedded representation in a high-dimensional vector space. This embedded representation can capture the relationships and features between sensor data and facilitate higher-level data analysis, pattern recognition, and prediction in the model.

[0060] Unsupervised deep learning methods: Unsupervised deep learning methods are a class of machine learning techniques that aim to learn feature representations of data from unlabeled data without the need for manually annotated labels or category information. These methods are often used for tasks such as data exploration, feature learning, data compression, and generating new data.

[0061] Tensor: It is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. It is a generalization of the concept of vector. Vector is a first-order tensor that can be used to represent multilinear functions of linear relationships between some vectors, scalars and other tensors.

[0062] View: In this invention, we regard each graph containing independent sensor dependencies as a view.

[0063] Tensor Singular Value Decomposition (T-SVD): Tensor singular value decomposition is a method of decomposing a tensor. It is an extension of the traditional singular value decomposition (SVD) on multidimensional tensors. SVD is a powerful tool in linear algebra for decomposing a matrix into three other matrices that can represent the original matrix in a simplified form, capturing its most important features.

[0064] Tensor Low-Rank Approximation: Tensor low-rank approximation is a method to reduce the dimensionality and complexity of a high-dimensional tensor by approximating it as a low-rank tensor.

[0065] Tubular rank: Tubular rank is a special rank of a tensor, which is particularly suitable for describing tensors in high-dimensional space. Specifically, for a three-dimensional tensor (such as a video sequence), tubal rank can be regarded as the rank of the tensor in the time dimension.

[0066] Tensor product (T-product): Tensor product (T-product) is an operation between two tensors, which is similar to the operation of matrix multiplication in two-dimensional matrices. However, tensor product is different from matrix multiplication because it is applicable to high-dimensional tensors and can handle more complex data structures.

[0067] GRU (Gated Recurrent Unit): GRU is a variant of recurrent neural network (RNN) that aims to solve the long-term dependency problem while reducing the gradient vanishing problem in traditional RNN.

[0068] Graph Attention Network (GAT): Graph Attention Network is a graph neural network based on the attention mechanism. Compared with traditional graph neural networks, GAT introduces an attention mechanism that enables each node to dynamically weight its importance according to its related neighbor nodes. This attention mechanism enables GAT to flexibly learn the relationship between each node and its neighbor nodes, and perform information transfer and aggregation in graph data.

[0069] Interquartile range (IQR): Interquartile range is a statistic that describes a set of data. It is a measurement method in statistics that is used to measure the degree of dispersion of data. IQR represents the gap between the median and the upper and lower quartiles of the data. It can help us understand the distribution range and dispersion of the data.

[0070] Mean Squared Error (MSE): It is used to measure and evaluate the reconstruction and estimation quality of the unused spatial environment variable values, that is, the error measure between the estimated value and the reference point value.

[0071] See also Figure 1 and Figure 2 , Figure 1 A flowchart of a multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in an embodiment of the present application, Figure 1 Another flowchart of a multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in an embodiment of the present application, which is applied to an electronic device, the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion includes the following steps: S1, obtains the time series historical data of multimedia IoT through sensors.

[0072] Among them, the number of sensors in the time series historical data is N, the total time of the time series is represented by T, and the entire time series can be expressed as ,in represents the data generated by N sensors at time t.

[0073] At time t, obtain the data of a sliding window of length w:

[0074] Our goal is to predict the data at time t using historical data of length w .

[0075] S2, inputting the time series historical data into the graph attention network to obtain sensor features; wherein, inputting the time series historical data into the graph attention network to obtain sensor features includes: S21, inputting the time series historical data into multiple feature extractors respectively to obtain different initial features.

[0076] In order to fully mine sensor features from different data analysis angles, the present application embodiment designs 4 different feature extractors, which almost cover all data analysis angles of time series. Step S21 includes the following steps: S211 uses the sensor embedding vector that can be trained together with the model as a data analysis perspective to describe sensor features, thereby extracting the flexible and changeable device dependencies in the IoMT scenario and converting the time series historical data into a sensor embedding vector. The sensor embedding vector is:

[0077] in, Represents time series historical data, represents the sensor embedding vector, Represents the dimension of the sensor embedding vector.

[0078] S212, extract the features of the time series historical data through the linear layer to obtain the local features of the sensor:

[0079] in, is a time series historical data with a sliding window size of w. is the local feature of the sensor, Indicates mapping the input data to dimensional space, are its parameters.

[0080] Since the data values ​​of the original time series already contain a lot of available information, the data values ​​generated by a sensor in the SWAT dataset over a period of time are visualized as an example. Figure 3 A schematic diagram of the values ​​generated within a certain period of time provided in the embodiment of the present application, such as Figure 3 As shown in the figure, it can be seen that there is a significant difference in data values ​​between abnormal data and normal data, and the value of abnormal data will suddenly increase or decrease at this moment. Based on this phenomenon, by directly analyzing the local data values ​​generated by the device, we can intuitively obtain some dependencies between sensors. In order to capture these clear dependencies, we use a linear layer to describe the local features of each sensor and use it as a feature generated by one of the data analysis perspectives.

[0081] S213, extract the features of the time series historical data through the recurrent neural network GRU to obtain the time feature graph:

[0082]

[0083] in, is a linear layer used to obtain the initial state of the recurrent neural network GRU. are the parameters of the recurrent neural network, Represents the initial state of the recurrent neural network GRU.

[0084] S214, extract the features of the time series historical data through dilated convolution to obtain multi-scale features: Figure 4 This is a flow chart of feature extraction of a multi-scale feature extractor provided in an embodiment of the present application, where DCL represents a dilated convolutional layer. Figure 5 For a schematic diagram of the structure of the dilated convolutional layer provided in the embodiment of the present application, please refer to Figure 4 and Figure 5 , step S214 specifically includes: For sensor i, define -Dilated convolution operation for:

[0085] in, is the time series generated by sensor i, is a 1D convolution filter kernel of size 1×k, and d is the dilation factor.

[0086] Stacking There are dilated convolution layers, and the dilated convolution used can be formulated as:

[0087]

[0088]

[0089]

[0090] Where i=1, 2, …, N, j=1,2,…, represents the input time series historical data with sliding window size w generated by sensor i, and are two parallel dilated convolution operations, is the multi-scale feature of sensor i in the j-th dilated convolutional layer, are c 1D convolution kernels of different sizes, Indicates that the feature is truncated to dimensions, and then connect them together. is the input of the dilated convolutional layer of sensor i in the jth layer. In each dilated convolutional layer, we organize the multi-scale features of all sensors into a matrix, represented as .

[0091] Figure 6 A t-SNE schematic diagram of a sensor embedding vector provided in an embodiment of the present application, Figure 7 A t-SNE schematic diagram of the local features of the sensor provided in the embodiment of the present application, Figure 8 This is a t-SNE schematic diagram of the sensor time characteristics provided in the embodiment of the present application. Fig. 9 This is a t-SNE schematic diagram of the multi-scale features of the sensor provided in the embodiment of the present application. Fig.10 A t-SNE diagram of the final sensor features output by the tensor attention network provided in an embodiment of the present application. Figure 6-Figure 10 It is shown that the present invention can effectively extract multiple sensor dependencies of time series, and also shows that the present invention can transmit heterogeneous information between different views and multiple dependencies between sensors in a graph tensor.

[0092] S22, construct a graph structure based on the initial features and convert the graph structure into a graph tensor.

[0093] Fig.11 A schematic diagram of the model performance of different feature extractors provided in the embodiment of the present application when constructing a tensor. In one embodiment, step S22 includes the following steps: S221, build a graph structure based on the initial features, and the graph structure is defined as: ,i=1,2,…, ;in, Represents a node, represents the edge, and the sensor is regarded as a node. It means that there is an edge from node i to node j, and it also means that node j is affected by or depends on node i.

[0094] S222, calculating the similarity between sensors based on the initial features, and when the similarity exceeds a preset similarity threshold, establishing edges between corresponding sensors.

[0095] Specifically, the features represented by F∈{V,M,H,D1…v} are used to calculate the similarity between sensors, and the TopK algorithm is used to establish edges between sensors with high similarity:

[0096]

[0097] Where i=1,2,…,N, is the weighted adjacency matrix (g=1,2,…, ), which means that in the figure There is a path from node j to node i with weight The edge, Represents the attention mechanism used to calculate the weights between two adjacent nodes.

[0098] The attention mechanism and similarity calculation formula used to calculate edge weights are as follows:

[0099]

[0100] where Fi represents the features F∈{V,M,H,D1…v} representing node / sensor i.

[0101] S223, connect the weighted adjacency matrices corresponding to the graph structure into a graph tensor.

[0102] Multi-view diagram (g=1,2,…, ) is organized as a rank-3 tensor:

[0103] in, These weighted adjacency matrices are Connect along the third dimension to form a graph tensor .

[0104] S23, calculates a low-rank r approximation of the graph tensor based on the graph tensor.

[0105] Specifically, tensor singular value decomposition (T-SVD) is introduced to obtain a low tubular rank-r approximation of the graph tensor. This is done to enable information to be transferred between different graphs and to fully consider the factors that affect the sensor dependency of the time series. Step S23 includes the following steps: Use the tensor singular value decomposition method to decompose the graph tensor Perform a low-rank r approximation:

[0106] Among them, r represents tubal rank, " is the tensor product (T-product) operation, as well as Respectively represent and The first r lateral slices of It only contains The diagonal tensor of the first r tensor tubes, is a tensor after the low-rank r approximation based on T-SVD.

[0107] S24, expands the graph attention network into a tensor attention network through tensor product, and outputs the sensor features of the current time series based on the tensor attention network.

[0108] Specifically, step S24 includes: S241, use matrix product to represent a two-layer graph attention network GAT:

[0109]

[0110] in, is a parameter set, is the weighted adjacency matrix, and is a weight matrix that can be learned along with the training process, represents a non-linear activation function.

[0111] S242, using the tensor product (T-product) operation to extend the graph attention network to the tensor attention network (TGAT):

[0112]

[0113]

[0114] in, represents the tensor product operation, represents the number of neural network layers of the tensor attention network, is the feature tensor, Represents the weight tensor of the cth layer ( , , ), H is the feature dimension of the hidden layer of the tensor attention network, represents a nonlinear activation function, which is a weight matrix used to map a three-dimensional tensor to two-dimensional sensor features. is the parameter set containing all parameters of layer c, The final sensor features representing the output of the model.

[0115] S3, uses a graph deviation scoring method to identify the time points of multimedia IoT anomalies based on sensor features.

[0116] In one embodiment, step S3 includes: S31, predicting the time series at the current moment based on the sensor features and the sensor embedding vector.

[0117] Specifically, the final features of all sensors at the current time t are expressed as:

[0118] in, .

[0119] Then the time series value at the current time t is predicted by element-wise multiplication of the sensor features and the sensor embedding vector:

[0120] in, Indicates that the parameter is The fully connected layer.

[0121] S32, calculating the anomaly score at the current moment. If the anomaly score at the current moment is greater than the anomaly threshold, it is determined that the multimedia Internet of Things state at the current moment is abnormal.

[0122] Specifically, step S32 includes the following steps: S321, first calculate the predicted value and the original data The error between:

[0123] Among them, the tensor attention network model is trained using a loss function based on Mean Squared Error:

[0124] in, is the predicted data at time t, is the original data at time t.

[0125] Among them, the time series in the training set used to train the tensor attention network have no abnormal labels, and the embodiment of the present application adopts an unsupervised deep learning method to train the model.

[0126] Fig.12 A schematic diagram of sensitivity analysis of parameter r on 6 data sets provided in the embodiment of the present application, Fig.13 A schematic diagram of sensitivity analysis of parameter k on 6 data sets provided in an embodiment of the present application, Fig.14 A schematic diagram of the model performance on 6 data sets when using TGAT with different numbers of layers provided in an embodiment of the present application. Fig.15 This is a statistical chart of detailed information of the six data sets provided in the embodiments of the present application. For all data sets, only the test set has abnormal data. Fig.16 A schematic diagram of comparative experimental results of anomaly detection on six data sets using Pre (Precision, %), Rec (Recall rate, %), and F1 (F1-score) indicators provided in an embodiment of the present application, wherein the bold data represent the highest performance on the corresponding data set indicators. Fig.17 The ablation experiment results on 6 data sets are provided in the embodiments of the present application. Figure 12-Figure 17 The dataset and training results for the Tensor Attention Network model are shown.

[0127] S322, use median and interquartile range To normalize the error:

[0128] S323, select standardized error The maximum value on N sensors is taken as the anomaly score at time t, and the maximum anomaly score along the time dimension of the validation dataset is used as the threshold. If the anomaly score at time t is greater than the threshold, we regard it as "abnormal": .

[0129] The multimedia IoT anomaly detection method used in the above steps S1-S3 can be applied to an anomaly detection model. The anomaly detection model is constructed based on the sensor coverage scenario of the target area. The specific structure of the anomaly detection model can be found in Figure 2 .

[0130] The number of sensors in the time series is represented as N, and the total time of the time series is represented as T. The entire time series can be represented as ,in Represents the data generated by N sensors at time t. At time t, the input of the anomaly detection model is the data of a sliding window of length w: , the goal of the anomaly detection model is to predict the data at time t through historical data of length w Among them, the abnormal label of the predicted data is expressed as , It means that the data at that moment is abnormal, otherwise it is normal data.

[0131] The anomaly detection model consists of five main modules: (1) Graph tensor construction. We use multi-scale feature extractors, temporal feature extractors, local feature extractors, and sensor embedding to extract sensor features from different perspectives. We use these features to construct a graph tensor.

[0132] (2) Low-rank-r approximation based on T-SVD. It is used to preserve the most closely related content between views in the graph tensor. This process can be regarded as passing messages between different views of the graph tensor.

[0133] (3) Graph learning based on tensor GAT. GAT is extended to a tensor version, called TensorGAT (TGAT).

[0134] (4) Output layer. After this layer, the sensor features learned by TGAT are decoded into predicted time series data at time t.

[0135] (5) Graph Deviation Scoring: A special anomaly scoring calculation method will be used to determine the abnormal time points.

[0136] The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in this application includes the following beneficial effects: (1) High anomaly detection accuracy. The present invention designs multiple feature extractors to extract sensor features in time series from multiple data analysis perspectives, and thereby mines the dependencies between multiple sensors. These feature extractors cover almost all data analysis perspectives of time series, and can fully mine the dependencies between sensors, thereby improving the accuracy of model anomaly detection. (2) The model is highly robust. The present invention utilizes low-rank approximation based on tensor product (T-product) and tensor GAT to transfer information between multiple sensor dependencies. These tensor operations can also effectively transfer and integrate heterogeneous information from multimedia, making the model adaptable to different multimedia application scenarios; (3) Comprehensive evaluation indicators. The present invention comprehensively considers the impact of factors such as sensor embedding vectors, sensor local features, sensor time dimension features, and sensor multi-observation scale features on the intrinsic dependencies of time series. Tensors are used to uniformly represent multiple sensor correlation factors in a high-dimensional manner, and each factor is integrated into a framework. Tensor neural networks are used to predict the time series value at the current moment, which can comprehensively mine and integrate the complex attributes and features of time series. (4) Strong versatility. The present invention is based on an unsupervised deep learning method. We do not need a time series with anomaly labels. In addition, the present invention is a time series anomaly detection model based on prediction. That is, we predict whether the data generated at the current time is abnormal data from the historical data of the time series, rather than waiting until the anomaly occurs before detecting the anomaly. This is more in line with the needs of practical applications.

[0137] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion. The multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion can be implemented as an independent entity or integrated in an electronic device, which can be a terminal, server and other devices. The terminal can include a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a micro processing box, or other devices.

[0138] See also Fig.18 , Fig.18 The multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion provided in the embodiment of the present application is specifically described and applied to electronic devices. The multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion may include: An acquisition module is used to acquire the time series historical data of the multimedia Internet of Things through sensors; The feature extraction module is used to input the time series historical data into the graph attention network to obtain the sensor features; wherein the time series historical data is input into the graph attention network to obtain the sensor features, including: Input the time series historical data into multiple feature extractors respectively to obtain different initial features; construct a graph structure based on the initial features and convert the graph structure into a graph tensor; calculate the low-rank r approximation of the graph tensor based on the graph tensor; expand the graph attention network into a tensor attention network through tensor product, and output the sensor features of the current time series based on the tensor attention network; Anomaly identification module is used to identify the time points of multimedia IoT anomalies based on sensor features using a graph deviation scoring method.

[0139] During specific implementation, the above modules and / or units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above modules and / or units can refer to the previous method embodiments. The specific beneficial effects that can be achieved can also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.

[0140] In addition, the embodiment of the present application further provides an electronic device, which may be a computer, a tablet computer, or other device. The electronic device may implement the steps in any embodiment of the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in the embodiment of the present application, and therefore, may achieve the beneficial effects that can be achieved by any multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in the embodiment of the present application, as detailed in the previous embodiment, which will not be repeated here.

[0141] Fig.19 The specific structural block diagram of the electronic device provided in the embodiment of the present invention is shown, and the electronic device can be used to implement the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in the above embodiment. The electronic device 500 can be a terminal, a server and other devices, wherein the terminal can include a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a micro processing box, or other devices.

[0142] The RF circuit 510 is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, user identity module (SIM) cards, memories, etc. The RF circuit 510 can communicate with various networks such as the Internet, corporate intranets, wireless networks, or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The above-mentioned wireless networks may use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE802.11g and / or IEEE802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, even those that have not yet been developed.

[0143] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizing the functions of taking pictures with the front camera, processing the captured images, and switching the display color of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 520 may further include a memory remotely arranged relative to the processor 580, and these remote memories may be connected to the electronic device 500 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0144] The input unit 530 may be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function control. The display unit 540 may be used to display information input by the user or information provided to the user and various graphical user interfaces, which may be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541, which may be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like.

[0145] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, which is converted into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and converted into audio data, and then the audio data is output to the processor 580 for processing, and then sent to another terminal through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earplug jack to provide communication between an external headset and the electronic device 500.

[0146] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), which provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it is not a necessary component of the electronic device 500 and can be omitted as needed without changing the essence of the invention.

[0147] The processor 580 is the control center of the electronic device 500. It uses various interfaces and lines to connect various parts of the entire mobile phone. By running or executing software programs and / or modules stored in the memory 520, and calling data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580.

[0148] The electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to various components. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so that the power management system can manage charging, discharging, and power consumption management. The power supply 590 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0149] Although not shown, the electronic device 500 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be described in detail here. Specifically in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, wherein one or more programs are stored in the memory, and are configured to be executed by one or more processors. One or more programs include instructions for performing the following operations: Obtaining time series historical data of multimedia IoT through sensors; Inputting the time series historical data into a graph attention network to obtain sensor features; wherein, inputting the time series historical data into a graph attention network to obtain sensor features includes: Inputting the time series historical data into multiple feature extractors respectively to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating a low-rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through tensor product, and outputting the sensor features of the current time series based on the tensor attention network; The graph deviation scoring method is used to identify the time points of multimedia IoT anomalies based on the sensor features.

[0150] In specific implementation, the above modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. The specific implementation of the above modules can be found in the previous method embodiments, which will not be repeated here.

[0151] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions, and the instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium, which stores multiple instructions, and the instructions can be loaded by a processor to execute the steps of any embodiment of the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided by the embodiment of the present invention.

[0152] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0153] Since the instructions stored in the storage medium can execute the steps in any embodiment of the multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in the embodiments of the present invention, the beneficial effects that can be achieved by any multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion provided in the embodiments of the present invention can be achieved. For details, please refer to the previous embodiments and will not be repeated here.

[0154] The above is a detailed introduction to a multimedia Internet of Things anomaly detection method, device, storage medium and electronic device based on tensor multi-feature fusion provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion, characterized in that: The method comprises: Obtaining time series historical data of multimedia IoT through sensors; Inputting the time series historical data into a graph attention network to obtain sensor features; wherein, inputting the time series historical data into a graph attention network to obtain sensor features includes: Inputting the time series historical data into multiple feature extractors respectively to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating a low-rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through tensor product, and outputting the sensor features of the current time series based on the tensor attention network; The graph deviation scoring method is used to identify the time points of multimedia IoT anomalies based on the sensor features.

2. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The multiple feature extractors include sensor embedding vectors, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The time series historical data is converted into a sensor embedding vector, and the sensor embedding vector is: in, Represents time series historical data, represents the sensor embedding vector, Represents the dimension of the sensor embedding vector.

3. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The multiple feature extractors include a local feature extractor, and the time series historical data is respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted through the linear layer to obtain the local features of the sensor: in, is the input time series historical data, is the local feature of the sensor, Indicates mapping the input data to dimensional space, are its parameters.

4. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The multiple feature extractors include a time feature extractor, and the time series historical data are respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted through a recurrent neural network to obtain a time feature graph: in, is a linear layer used to obtain the initial state of the recurrent neural network. are the parameters of the recurrent neural network, Represents the initial state of the recurrent neural network.

5. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The multiple feature extractors include a multi-scale feature extractor, and the time series historical data is respectively input into the multiple feature extractors to obtain different initial features, including: The features of the time series historical data are extracted by dilated convolution to obtain multi-scale features: For sensor i, define -Dilated convolution operation for: in, is the time series generated by sensor i, is a 1D convolution filter kernel of size 1×k, d is the dilation factor; Stacking There are dilated convolution layers, and the dilated convolution used is: Where i=1, 2, …, N, j=1,2,…, represents the input time series historical data with sliding window size w generated by sensor i, and are two parallel dilated convolution operations, is the multi-scale feature of sensor i in the j-th dilated convolutional layer, are c 1D convolution kernels of different sizes, Indicates that the feature is truncated to dimensions, and then connect them together. is the input to the dilated convolutional layer of sensor i in layer j.

6. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The step of constructing a graph structure based on the initial features and converting the graph structure into a graph tensor includes: A graph structure is constructed based on the initial features, and the graph structure is defined as: ,i=1,2,…, ;in, Represents a node, represents the edge, and the sensor is regarded as a node. Indicates that there is an edge from node i to node j; Calculating similarities between sensors based on the initial features, and when the similarities exceed a preset similarity threshold, establishing edges between corresponding sensors; The weighted adjacency matrices corresponding to the graph structure are concatenated into a graph tensor.

7. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The calculating a low-rank r approximation of a graph tensor based on the graph tensor comprises: Use the tensor singular value decomposition method to decompose the graph tensor Perform a low-rank r approximation: Among them, r represents The management rank of " is the tensor product operation, as well as Represented and The first r transverse slices of It only contains The diagonal tensor of the first r tensor tubes of , is a tensor after the low-rank r approximation based on T-SVD.

8. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 1 is characterized in that: The method of expanding the graph attention network into a tensor attention network by tensor product, and outputting sensor features of the current time series based on the tensor attention network, includes: Use matrix product to represent a two-layer graph attention network: in, is a parameter set, is the weighted adjacency matrix, and is a weight matrix that can be learned along with the training process, represents a nonlinear activation function; The graph attention network is extended to the tensor attention network using the tensor product operation: in, represents the tensor product operation, represents the number of neural network layers of the tensor attention network, is the feature tensor, Represents the weight tensor of the cth layer ( , ,), H is the feature dimension of the hidden layer of the tensor attention network, represents a nonlinear activation function, which is a weight matrix used to map a three-dimensional tensor to two-dimensional sensor features. is the parameter set containing all parameters of layer c, The final sensor features representing the output of the model.

9. The multimedia Internet of Things anomaly detection method based on tensor multi-feature fusion according to claim 2 is characterized in that: The method for using graph deviation scoring to identify the time point of multimedia Internet of Things anomalies based on the sensor features includes: Predicting a time series at a current moment based on the sensor feature and the sensor embedding vector; The anomaly score at the current moment is calculated, and if the anomaly score at the current moment is greater than the anomaly threshold, it is determined that the multimedia Internet of Things state at the current moment is abnormal.

10. A multimedia Internet of Things anomaly detection device based on tensor multi-feature fusion, characterized in that: include: An acquisition module is used to acquire the time series historical data of the multimedia Internet of Things through sensors; A feature extraction module, used for inputting the time series historical data into a graph attention network to obtain sensor features; wherein the step of inputting the time series historical data into a graph attention network to obtain sensor features includes: Inputting the time series historical data into multiple feature extractors respectively to obtain different initial features; constructing a graph structure based on the initial features, and converting the graph structure into a graph tensor; calculating a low-rank r approximation of the graph tensor based on the graph tensor; expanding the graph attention network into a tensor attention network through tensor product, and outputting the sensor features of the current time series based on the tensor attention network; The anomaly identification module is used to identify the time point of multimedia Internet of Things anomalies based on the sensor features using a graph deviation scoring method.

Citation Information

Cited By

  • Multi-parameter fusion air suspension heat pump AI load prediction energy-saving control method

    CN121635009A