Intelligent data monitoring method and monitoring platform

By constructing a dynamic heterogeneous hypergraph model and a multi-dimensional tensor decomposition algorithm, the problem of the inability to effectively express the high-order synergy relationship between multiple nodes in the existing monitoring technology is solved, and the accurate prediction of the abnormal propagation path and multi-level coordinated fault detection are realized, which improves the abnormal prediction accuracy and fault response speed of the monitoring system.

CN120492213APending Publication Date: 2025-08-15ZHANGJIAGANG BIG DATA CO LTD

Patent Information

Application Number
CN202510978397.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing monitoring technologies cannot effectively express and calculate the high-order synergy relationship between multiple nodes, and it is difficult to accurately distinguish the differentiated patterns and rates of abnormal propagation between different types of nodes. They lack multi-level coordinated fault detection capabilities and cannot accurately predict the abnormal diffusion path.

Method used

A dynamic heterogeneous hypergraph model is constructed, and the embedded representation of nodes and relationships is obtained through multi-dimensional tensor representation and decomposition algorithms are used to capture complex interaction features using the hyper-edge attention mechanism and meta-path random walk, anomaly propagation rate model is established, and multi-level collaborative anomaly detection is implemented.

Benefits of technology

It realizes accurate prediction of abnormal propagation paths, improves the accuracy of abnormal prediction, and can comprehensively monitor from single-node exceptions to global system coordinated faults, accurately distinguishes the abnormal propagation mode and rate between different types of nodes, and shortens the system recovery time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492213A_ABST
    Figure CN120492213A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data monitoring, and discloses an intelligent data monitoring method and a monitoring platform, and the intelligent data monitoring method comprises the steps: constructing a dynamic heterogeneous hypergraph model, and representing the complex relation and time sequence evolution characteristics among multi-source data nodes; based on a dynamic heterogeneous hypergraph model, realizing high-order relationship representation and calculation, and obtaining embedded representation of nodes and relationships; performing heterogeneous graph information transmission and aggregation by utilizing embedded representation of nodes and relationships, and capturing complex interaction characteristics among multiple nodes; according to the complex interaction characteristics, an abnormal propagation rate model is established, and accurate prediction of an abnormal diffusion path is realized; based on the abnormal propagation rate model and the abnormal diffusion path, multi-level collaborative anomaly detection is implemented, and collaborative faults across subsystems are identified; according to the invention, the abnormal propagation path can be predicted in advance, the abnormal prediction accuracy is improved, and the system fault response time is advanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data monitoring, and more particularly to an intelligent data monitoring method and a monitoring platform. Background Art

[0002] With the development of IoT technology and artificial intelligence, the demand for monitoring large-scale, complex systems is growing. These systems often consist of multiple types of device nodes, with complex relationships and collaboration patterns between them. Traditional monitoring methods, which primarily rely on threshold detection or statistical models, have limitations in processing high-dimensional, heterogeneous data, capturing implicit relationships, and predicting the propagation of anomalies.

[0003] Existing monitoring technologies, primarily based on binary graph models, are unable to effectively express and calculate high-level collaborative relationships between multiple nodes. Furthermore, these methods struggle to accurately distinguish the differentiated patterns and rates of anomaly propagation between different types of nodes, and lack the ability to monitor multiple levels of anomalies, from single-node anomalies to global coordinated system failures. Furthermore, traditional methods are unable to accurately predict the anomaly propagation path, making it difficult to precisely locate the anomaly source and assess the impacted area.

[0004] Therefore, there is an urgent need for intelligent data monitoring methods that can handle complex heterogeneous systems and implement high-order relationship representation, anomaly propagation modeling, and multi-level collaborative fault detection. With the development of Internet of Things (IoT) technology and artificial intelligence (AI), the demand for monitoring large-scale, complex systems is increasing. These systems typically consist of multiple types of device nodes with complex inter-node relationships and collaboration patterns. Traditional monitoring methods, which primarily rely on threshold detection or statistical models, have limitations in handling high-dimensional, heterogeneous data, capturing implicit relationships, and predicting anomaly propagation.

[0005] Existing monitoring technologies, primarily based on binary graph models, are unable to effectively express and calculate high-level collaborative relationships between multiple nodes. Furthermore, these methods struggle to accurately distinguish the differentiated patterns and rates of anomaly propagation between different types of nodes, and lack the ability to monitor multiple levels of anomalies, from single-node anomalies to global coordinated system failures. Furthermore, traditional methods are unable to accurately predict the anomaly propagation path, making it difficult to precisely locate the anomaly source and assess the impacted area.

[0006] Therefore, there is an urgent need for an intelligent data monitoring method that can handle complex heterogeneous systems and achieve high-order relationship representation, anomaly propagation modeling, and multi-level collaborative fault detection. Summary of the Invention

[0007] The present invention provides an intelligent data monitoring method and a monitoring platform to solve the technical problem in related technologies that high-order collaborative relationships between multiple nodes cannot be effectively expressed and calculated.

[0008] The present invention provides an intelligent data monitoring method, comprising: Construct a dynamic heterogeneous hypergraph model to represent the complex relationships and temporal evolution characteristics between multi-source data nodes; Based on the dynamic heterogeneous hypergraph model, high-order relationship representation and calculation are realized to obtain the embedded representation of nodes and relationships; Leveraging embedded representations of nodes and relationships, we perform heterogeneous graph information transfer and aggregation, capturing complex interaction features between multiple nodes. Based on the complex interaction characteristics, an abnormal propagation rate model is established to achieve accurate prediction of the abnormal diffusion path; Based on the anomaly propagation rate model and anomaly diffusion path, multi-level collaborative anomaly detection is implemented to identify collaborative failures across subsystems.

[0009] Furthermore, the step of constructing a dynamic heterogeneous hypergraph model includes: Standardize various monitoring data sources in the system, extract features and generate corresponding node representations; Analyze the collaborative relationship patterns between multiple nodes, identify sets of nodes with strong correlations, and construct hyperedges; Through time window sliding and snapshot sampling methods, the temporal evolution characteristics of the hypergraph structure are captured to form a time-varying hypergraph sequence.

[0010] Furthermore, the step of implementing high-order relation representation and calculation includes: Representing high-order relations in heterogeneous hypergraphs as multi-dimensional tensors; Using tensor decomposition algorithm, the high-dimensional relationship tensor is decomposed into a low-dimensional factor matrix to obtain the embedded representation of nodes and relationships; Combining timing information and heterogeneity constraints to optimize embedding representations.

[0011] Furthermore, the step of performing heterogeneous graph information transmission and aggregation includes: Designing attention weights for each hyperedge enables the system to adaptively aggregate information from different nodes; Design a type-aware information aggregation function based on the characteristics of heterogeneous graphs; Generate semantically relevant neighborhoods of nodes through meta-path based random walks.

[0012] Furthermore, the step of establishing the abnormal propagation rate model includes: Define a propagation rate tensor to describe the rate at which anomalies propagate from one type of node to another type of node through a specific relationship; Design a time-varying graph convolutional network to capture the temporal characteristics of anomaly propagation; Based on the trained propagation rate model, the abnormal diffusion path is predicted.

[0013] Furthermore, the steps of implementing multi-level collaborative anomaly detection include: Calculate node-level anomaly scores based on node feature representation and historical state sequences; Calculate subgraph-level anomaly scores based on the set of subgraphs defined by hyperedges; Calculate a system-level anomaly score based on the global hypergraph structure and all detected node-level and subgraph-level anomalies; Based on the anomaly propagation model and multi-level anomaly detection results, the risk of each part of the system being affected by anomalies is assessed.

[0014] Furthermore, the high-order relations are decomposed into a number of low-order relation combinations by using a hierarchical decomposition method; when there are a large number of sparse high-order relations in the system, a sparse tensor storage and calculation method is adopted.

[0015] Furthermore, the attention weights of each hyperedge design adopt a multi-layer stacked hyperedge attention network structure, and through the multi-head attention mechanism, multiple groups of independent attention weights are calculated in parallel to enrich the expressive power of the model.

[0016] Furthermore, the prediction of the anomaly diffusion path adopts a recursive rolling prediction strategy, first predicting the state of all nodes at the next moment, then using the prediction result as a new input to predict the state at the next moment until the entire prediction time window is completed; and by calculating the abnormal state changes of each node at different time points, a heat map of abnormal propagation is drawn.

[0017] The present invention provides an intelligent data monitoring platform for executing the above-mentioned intelligent data monitoring method, comprising: Dynamic heterogeneous hypergraph construction module, used to build a hypergraph model representing the complex relationships between multi-source data nodes in the system; High-order relational computing module, used to implement multi-dimensional tensor representation and decomposition, and obtain embedded representations of nodes and relations; Heterogeneous graph information transfer module, used to capture complex interaction features between multiple nodes through hyperedge attention mechanism and meta-path random walk; Anomaly propagation modeling module, which is used to build a propagation rate tensor and a time-varying graph convolutional network to accurately predict anomaly diffusion paths; The multi-level anomaly detection module is used to perform anomaly analysis at the node level, subgraph level, and hypergraph level to identify collaborative failures across subsystems.

[0018] The beneficial effects of the present invention are: it can predict the abnormal propagation path in advance, improve the accuracy of abnormality prediction, and advance the system fault response time; through a unified hypergraph representation, it can realize comprehensive monitoring from single-node abnormalities to multi-node collaborative faults, and the accuracy of hidden collaborative fault detection is improved compared with traditional methods; it can accurately distinguish the abnormal propagation modes and rates between different types of nodes, adapt to complex heterogeneous systems with multiple node and relationship types, and improve the accuracy of cross-subsystem abnormality propagation prediction; break through the limitations of binary relationships, express and analyze high-order collaborative relationships involving up to 8 nodes, and improve the accuracy of complex collaborative fault identification; provide a global view of system-level abnormalities, accurately locate the source of abnormalities and quantitatively evaluate the scope of impact, shorten the system recovery time, and reduce economic losses and service interruption time. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of an intelligent data monitoring method in the present invention; Figure 2 is a flow chart of step 1 of the present invention; Figure 3 is a flow chart of step 2 of the present invention; Figure 4 is a flow chart of step 3 of the present invention; Figure 5 is a flow chart of step 4 of the present invention; Figure 6 It is a flow chart of step 5 of the present invention. DETAILED DESCRIPTION

[0020] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0021] At least one embodiment of the present invention discloses an intelligent data monitoring method, such as Figures 1 to 6 As shown, the following steps are included: Step 1: Build a dynamic heterogeneous hypergraph model to represent the complex relationships and temporal evolution characteristics between multi-source data nodes; In an embodiment of the present application, this step represents the complex relationships and temporal evolution characteristics between multi-source data nodes in the system by constructing a dynamic heterogeneous hypergraph model, thereby breaking through the limitation that traditional graph models can only express binary relationships.

[0022] Specifically, this application defines a dynamic heterogeneous hypergraph model of the system : ; in Represents a node set, corresponding to various monitoring objects in the system, represents the set of hyperedges, Represents the node type mapping function, represents the hyperedge type mapping function, Represents the time dimension; ; in 、 、 Represents the first, second, and nodes, is the total number of nodes in the system; ; in Represents a set of hyperedges, each hyperedge Can connect two or more nodes to represent a high-order collaborative relationship between multiple nodes. 、 、 Respectively represent Article 1, Article 2, Super edge, is the total number of hyperedges in the system; ; Represents a node type mapping function that transforms the node set Mapping to a collection of node types : ; in Represents a set of node types, 、 、 Respectively represent the first, second, and Node types, is the total number of node types in the system; ; Represents the hyperedge type mapping function, which transforms the hyperedge set Mapping to a collection of relationship types : ; in Represents a set of relationship types, 、 、 Respectively represent the first, second, and Types of relationships, is the total number of relationship types in the system; ; in Represents the time dimension, used to capture the temporal evolution of the system state; 、 、 Represents the first, second, and A point in time, is the length of the time series.

[0023] This step contains the following sub-steps: Step 1.1, data preprocessing and node generation; In this step, various monitoring data sources in the system are standardized, features are extracted and corresponding node representations are generated. , construct its initial feature vector , contains the static attributes and dynamic state information of the node; It should be noted that the specific implementation methods include: first, normalizing various types of monitoring data to make data of different dimensions comparable; second, using principal component analysis to reduce feature dimensions and retain key information; finally, fusing static attributes (such as device type, location information, etc.) and dynamic states (such as current load, temperature, etc.) to form an initial feature vector.

[0024] In some implementations, self-supervised learning methods can be optionally employed to enhance node feature representations. For example, by masking some node features and training a model to predict these masked features, a more robust node representation can be obtained. This self-supervised approach can effectively improve feature quality in scenarios with sparse data or high noise levels.

[0025] In the smart city infrastructure monitoring scenario, for example, for the power system, nodes can represent substations, distribution equipment, sensors, etc. The initial feature vector of each node contains static attributes (equipment type, rated power, installation location, etc.) and dynamic status (current load rate, temperature, operating time, etc.). Specifically, for the substation node , its initial eigenvector can be expressed as: ; in Represents a substation node The first three elements of the initial feature vector are the one-hot encoded device type, and the subsequent elements are the current load rate, temperature normalized value, and runtime normalized value.

[0026] Step 1.2, hyperedge identification and construction; In addition, this sub-step analyzes the collaborative relationship patterns between multiple nodes, identifies sets of nodes with strong correlations, and constructs hyperedges. Hyperedge construction uses a variety of correlation measurement methods, including: Explicit relationship definition based on domain knowledge; implicit relationship discovery based on correlation analysis; causal relationship inference based on temporal consistency In the specific implementation, for explicit relationships based on domain knowledge, this application predefines a set of hyperedge templates based on the system topology and functional associations. For example, in the power system, a substation and all its connected distribution equipment form a hyperedge to represent the power supply relationship.

[0027] For implicit relationships based on correlation analysis, this application uses a sliding time window to calculate the mutual information or Pearson correlation coefficient matrix between nodes. Nodes with high correlation coefficients are then grouped into hyperedges using clustering algorithms (such as hierarchical clustering or community discovery algorithms). In an industrial production environment, such as a chemical production line, by analyzing the correlation of sensor data from each device, a collection of devices with process dependencies can be discovered, forming a hyperedge structure, even if these devices may not be physically adjacent.

[0028] Optionally, in some implementations, this application may also employ a hyperedge generation method based on a graph neural network. Specifically, an initial binary relationship graph is constructed, node representations are learned through a graph neural network, and finally, hyperedges are automatically generated based on the similarity or clustering results of the node representations. This method is particularly suitable for large-scale systems with complex and implicit relationships, such as urban transportation networks or large clusters of IoT devices.

[0029] In addition, when dealing with dynamically changing systems, the present application can optionally implement an adaptive hyperedge update mechanism. This mechanism dynamically adjusts the composition of hyperedges by continuously monitoring changes in node status and relationship strength. For example, when the correlation between certain nodes is detected to be increasing, the system automatically forms a new hyperedge; when the node relationship within a hyperedge weakens below a threshold, the hyperedge is automatically split or deleted. This adaptive mechanism enables the hypergraph structure to reflect the dynamic changes of the system in real time, improving the accuracy of anomaly detection.

[0030] Step 1.3, time-varying relationship modeling; In addition, through time window sliding and snapshot sampling methods, the evolution characteristics of the hypergraph structure over time are captured to form a time-varying hypergraph sequence: ; in 、 、 Respectively indicate time points 、 、 A hypergraph snapshot of is the length of the time series.

[0031] In this application, a fixed-length time window (e.g., 10 minutes) is used to slide and generate a hypergraph sequence. The window sliding step size can be adjusted according to the scenario requirements (e.g., 1 minute). For each time window, node features and hyperedge relationship strengths are recalculated to generate a hypergraph snapshot at the corresponding time point.

[0032] In IoT device cluster monitoring applications, such as smart building systems, time-varying hypergraphs can be used to capture the changing collaborative working patterns of building equipment over time. For example, during working hours (9:00 AM to 5:00 PM), hyperedges strongly connect the air conditioning, lighting, and elevator systems. However, during off-hours, the connections between these systems weaken, while the connections between the security system and other systems strengthen. This time-varying behavior can be effectively represented and analyzed using hypergraph sequences.

[0033] Alternatively, in scenarios with large data volumes or high real-time requirements, this application can employ an incremental time-varying relationship modeling approach. This approach does not require recalculating the entire hypergraph at each time window. Instead, it updates only the changed parts based on the hypergraph structure at the previous moment, thereby reducing computational complexity. For example, for a large urban transportation network, the state of relevant nodes and hyperedges can be updated only when traffic flow changes on certain sections of road, while leaving the rest unchanged.

[0034] In another implementation, the present application can incorporate multi-scale time windows for modeling time-varying relationships. Specifically, hypergraph sequences with different time granularities, such as minute-, hour-, and day-level windows, are simultaneously maintained to capture system evolution characteristics at different time scales. This multi-scale approach is particularly suitable for complex systems with varying temporal patterns, such as energy consumption monitoring scenarios with both short-term fluctuations and long-term trends.

[0035] Therefore, the output of this step is a complete dynamic heterogeneous hypergraph model, which serves as the basic data structure for subsequent high-order relation representation and calculation.

[0036] Step 2: Based on the dynamic heterogeneous hypergraph model, high-order relationship representation and calculation are implemented to obtain embedded representations of nodes and relationships; This step uses multi-dimensional tensor representation and decomposition algorithms to achieve effective representation and calculation of high-order relationships, thereby breaking through the limitations of traditional binary relationship representation and obtaining low-dimensional embedding representations of nodes and relationships.

[0037] Step 2.1, construct a multidimensional relationship tensor; In this substep, high-order relations in the heterogeneous hypergraph are represented as multidimensional tensors: ; in Indicates the The size of the dimensions, represents a multidimensional relational tensor, express dimensional real tensor space, is the number of dimensions, Represents the field of real numbers.

[0038] for High-order relationships between nodes, tensor elements Representation node The strength or probability of the relationship between Representing a tensor In the index The element value at .

[0039] It should be noted that in the specific implementation, this application uses two methods to construct multidimensional relationship tensors: First, for small-scale high-order relationships (usually involving 3-4 nodes), a complete multi-dimensional tensor is directly constructed. For example, in smart grid monitoring, a third-order tensor is constructed: ; in 、 and Respectively represent the node sets of power generation equipment, transmission equipment and distribution equipment, represents a multidimensional relational tensor, Represents the field of real numbers.

[0040] Secondly, for large-scale high-order relationships (involving five or more nodes), a hierarchical decomposition approach is used to decompose the high-order relationship into a combination of several lower-order relationships. In industrial production monitoring scenarios, such as an automated production line, if there is a collaborative relationship formed by eight process nodes, it can be decomposed into two fourth-order relationships, which are then represented and calculated separately to improve computational efficiency.

[0041] Optionally, when there are a large number of sparse high-order relations in the system, the present application may adopt sparse tensor storage and calculation methods. Specifically, only the non-zero elements in the tensor and their indices are stored, which greatly reduces the storage space and computational complexity. For example, in a large-scale IoT monitoring system, the number of nodes can reach tens of thousands, but the actual high-order relations formed usually only account for a very small part of the theoretical total. Using sparse tensor representation can reduce the storage requirement from TB level to GB level.

[0042] In addition, when processing time-varying high-order relations, the present application may optionally adopt a tensor flow representation method to represent the time-varying relations as a tensor sequence to capture the changing pattern of the relation strength over time, namely: ; in 、 、 Respectively indicate time points 、 、 The relation tensor of is the length of the time series.

[0043] In urban traffic flow monitoring, this representation method can effectively capture the changing patterns of traffic flow in different time periods and assist in identifying abnormal traffic events.

[0044] In addition, for different types of nodes and relationships, type-related tensors are constructed , represents a specific type of relationship between specific types of nodes. Indicates the node type and Relationship Type A connected relation tensor. Ellipses indicate that more types may be involved.

[0045] Step 2.2, tensor decomposition and embedding learning; Next, this application uses a tensor decomposition algorithm to decompose the high-dimensional relationship tensor into a low-dimensional factor matrix, thereby obtaining an embedded representation of nodes and relationships. Specifically, the following optimization problem is solved:

[0046] in It is The embedding matrix of dimensions, is the embedding dimension; is the embedding matrix of the relation type, represents the rank of tensor decomposition; represents the outer product operation; is a regularization term used to prevent overfitting; is the regularization coefficient; Represents the Frobenius norm, which is used to measure the difference between tensors; 、 、 Represents the first, second, and Embedding matrix of dimensions 、 、 No. row, corresponding to potential factors, Indicates the number of dimensions; Represents the relation embedding matrix No. row, corresponding to potential factors; Represents the summation symbol.

[0047] It should be noted that in the specific implementation, this application uses an alternating optimization strategy to solve the above optimization problem. First, all embedding matrices are randomly initialized. Then, the other matrices are fixed and a single matrix is optimized, iterating until convergence. To improve computational efficiency, large-scale tensor decomposition uses a distributed computing framework, distributing data and computation tasks across multiple computing nodes.

[0048] In some embodiments, the present application may optionally employ a neural network-based tensor decomposition method, such as neural CP decomposition or neural Tucker decomposition, to enhance the model's expressiveness and capture more complex, high-order relational patterns by introducing nonlinear activation functions and multi-layer structures. For example, in smart grid fault propagation analysis, traditional linear tensor decomposition may not accurately express the nonlinear propagation characteristics of faults between different types of equipment, while neural network-based methods can better model this complex pattern.

[0049] Furthermore, when the system contains relational evolution patterns at different time scales, the present application can optionally implement a multi-scale tensor decomposition method. Specifically, the relational tensors at different time granularities are decomposed separately, and then the embedding representations at different scales are integrated through a cross-scale attention mechanism. This approach is particularly effective in smart city monitoring, as urban systems often contain minute-level (e.g., traffic flow), daily (e.g., energy consumption), and seasonal (e.g., water resource utilization) variations.

[0050] In smart city traffic monitoring applications, such as road network congestion prediction, tensor decomposition can be used to obtain embedded representations of road segment nodes, time nodes, and traffic flow types. These embedded vectors capture the spatial dependencies between road segments, temporal evolution patterns, and mutual influences under different traffic flow conditions, which is helpful for subsequent congestion propagation analysis. Specifically, for a certain intersection , whose embedding representation It may show a similar pattern to the main road intersections it connects to, reflecting the characteristics of traffic flow propagation.

[0051] Step 2.3, embedding fusion and optimization; It should be noted that, in view of the time-varying characteristics, this step combines the timing information to optimize the embedding representation and adopts the timing regularization term: ; in Indicates a time point The embedding matrix of represents the timing regularization term, Represents a time collection The number of elements (cardinality), and Respectively indicate time points and The embedding matrix of represents the Frobenius norm; Represents the summation symbol.

[0052] In its implementation, this application also introduces an adaptive time series smoothing mechanism that dynamically adjusts the time series regularization strength based on node type and time period characteristics. For example, the evolution of the traffic network differs significantly on weekdays and weekends. Therefore, different regularization coefficients are used for these two time periods to improve the model's adaptability to time-varying patterns.

[0053] Optionally, to better capture long-term dependencies, this application can use sequence models such as Recursive Neural Network (RNN), Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU) to enhance the temporal embedding representation. Specifically, the node embedding sequence is input into the sequence model to generate an enhanced embedding containing historical information: ; in 、 、 Represents nodes respectively At the time point 、 、 The embedding representation of Indicates the time window size.

[0054] In smart grid monitoring, this approach can effectively capture long-term trends and periodic patterns in device status and improve the accuracy of anomaly detection; In another embodiment, the present application may implement an embedding enhancement method based on a graph convolutional network. By treating node embeddings as node features on a graph, a multi-layer graph convolutional network is applied to further incorporate neighborhood information, enhancing the expressive power of the embeddings. In multi-level system monitoring, this approach can effectively integrate node information from different levels, while simultaneously considering device-level, subsystem-level, and system-level features.

[0055] In addition, to address heterogeneity, type-related embedding constraints are introduced: ; in Represents type-related embedding constraints, 、 denote the node embedding matrix and the relationship embedding matrix respectively, Represents a set of node types, Represents a set of relationship types, and Represents the node type mapping function and the relationship type mapping function respectively, Representation type The embedding matrix of the node, Representation type The embedding matrix of the relation, Representation type The embedding matrix of the node, Representation matrix The transpose of Representation type Nodes and Types Nodes pass type The relation tensor of relation connections, represents the Frobenius norm; Represents the summation symbol.

[0056] In industrial equipment monitoring scenarios, this application implements type-specific embedding constraints to ensure that the embedded representations of different types of equipment (such as production equipment, testing equipment, and storage devices) retain common characteristics within the type while also reflecting the interaction patterns between different types. For example, for processing equipment and quality inspection equipment on a production line, although the equipment types are different, their close relationship in the process flow will result in similarities in specific dimensions, facilitating the subsequent identification of collaborative anomaly patterns across different equipment types.

[0057] Therefore, the output of this step is the node embedding matrix and the relation embedding matrix , providing a basis for subsequent heterogeneous graph information transmission and aggregation.

[0058] Step 3: Using the embedded representation of nodes and relationships, we perform heterogeneous graph information transfer and aggregation to capture the complex interaction features between multiple nodes. According to another embodiment of the present application, this step realizes information transmission and aggregation in the heterogeneous graph through the hyperedge attention mechanism and meta-path random walk, thereby effectively capturing the complex interaction features and semantic associations between multiple nodes.

[0059] Step 3.1, super edge attention mechanism; In this step, attention weights are designed for each hyperedge so that the system can adaptively aggregate information from different nodes. Specifically, for node , update its representation by the following formula: ; in Contains nodes The set of all hyperedges of ; is a node In the Representation of layers; is a node In the Representation of layers; is the information aggregation function; Represents a hyperedge Remove the node All nodes outside Layer representation set Representation node In the Representation of layers; is a nonlinear activation function; It is a super edge The attention weight is calculated as follows: ; in is the LeakyReLU activation function; is a learnable parameter matrix used to transform the input of the attention calculation; Represents vector concatenation operation; represents the natural exponential function; Indicates that it contains nodes Another superedge of ; Indicates that the node The representation of its hyperedge The neighbor aggregation in represents splicing.

[0060] It should be noted that in the specific implementation, this application uses a multi-layer stacked hyperedge attention network structure, usually consisting of 2-3 layers, to capture relational information at different levels of abstraction. For each layer, the hyperedge attention weight is first calculated, then the neighbor information is aggregated based on the weight, and finally converted through a nonlinear activation function (such as ReLU or LeakyReLU).

[0061] In some embodiments, the present application may optionally implement a multi-head attention mechanism to enrich the expressive power of the model by computing multiple sets of independent attention weights in parallel.

[0062] Specifically, for each attention head , using different parameter matrices Calculating attention weights , and then concatenate or sum the outputs of multiple heads: ; in It is the node after splicing multiple outputs In the The layer representation, is the number of attention heads, usually set to 4-8; It is Hyperedges computed by attention heads The attention weight of It is The information aggregation function used by each attention head; Indicates the sum of the results of all attention heads; Represents a hyperedge Remove the node All nodes outside Layer representation set Representation node In the Layer representation.

[0063] In complex industrial system monitoring, multi-head attention can simultaneously focus on different types of collaborative relationships, such as functional dependencies, physical connections, and data flows between devices, thereby improving the comprehensiveness of anomaly detection.

[0064] Additionally, when processing large-scale hypergraphs, this application can optionally employ a sampling-based sparse attention calculation method to reduce computational complexity. For each node, instead of calculating its attention with all hyperedges, a fixed number of hyperedges are randomly sampled for calculation. For large-scale systems, such as city-level traffic network monitoring, this approach can reduce computational complexity from quadratic to linear while maintaining model performance.

[0065] In smart manufacturing scenarios, such as semiconductor wafer production line monitoring, the hyperedge attention mechanism can adaptively focus on other device nodes that are most relevant to the current device status anomaly. If the parameters are abnormal, the super-edge attention will be automatically assigned to the upstream lithography machine and downstream cleaning machines Higher attention weights, because they form a process flow hyperedge with the etcher, and their state has a direct impact on the etcher. Specifically, in a certain state transfer, it may be calculated (in Include 、 and ),and (in Include and other maintenance equipment), indicating that the system pays more attention to the impact of process flow relationships on abnormal conditions.

[0066] Step 3.2, type-aware information aggregation; In addition, based on the characteristics of heterogeneous graphs, this application designs a type-aware information aggregation function: ; in is the information aggregation function, Representation node Type; Represents a hyperedge Type; Representation node Type; It is node-based 、 and nodes Type Calculated normalization coefficient; is the transformation matrix specific to the node type and edge type; Represents the hyperedge Remove the node Sum all nodes outside the . Represents a hyperedge Remove the node All nodes outside Layer representation set Representation node In the Layer representation.

[0067] In the implementation, this application designs a special transformation matrix for different types of node pairs to adapt to the differences in node features in heterogeneous graphs. Specifically, the attention mechanism is used to calculate the normalization coefficient : ; in It is node-based 、 and nodes Type Calculate the normalization coefficient, is a learnable attention vector; Is specific to the node type The transformation matrix of Is specific to the node type The transformation matrix of Represents a vector The transpose of is the leaky rectified linear unit activation function; represents the natural exponential function; Represents a hyperedge Remove the node Another node outside; Represents the hyperedge Remove the node Sum all nodes outside the . In the smart city water system monitoring application, type-aware information aggregation can handle the heterogeneous relationships between different types of nodes such as water plants, pipe networks, and pumping stations. For example, when monitoring a certain area (node , type is "region") when the water pressure is abnormal, the system will automatically adjust the flow from different types of nodes (such as upstream pump station nodes) based on the type perception mechanism. and trunk nodes ) Aggregate information in a way that applies different transformation matrices to the pump station status information and the pipeline network status information respectively, and then determines the weights of their influence on the water pressure anomaly in the current area through the attention mechanism.

[0068] Step 3.3, meta-path random walk; Next, this application generates semantically relevant neighborhoods of nodes through random walks based on meta-paths Meta Path Defined as: ; in 、 、 Respectively represent the first, second, and Node types, is the length of the meta-path; Indicates the total number of nodes; Indicates the The relationship type between node types; For nodes , perform multiple random walks based on different meta-paths to obtain a node set Then calculate the node and The semantic similarity between: ; in Representation node and nodes In the meta path The semantic similarity under It is a slave node Set out along the metapath The set of nodes visited by random walks; It is a slave node Set out along the metapath The set of nodes visited by random walks; Indicates the size of the intersection of two sets; Indicates the size of the union of two sets; Based on this similarity, this application defines semantically related neighborhoods: ; in Representation node and nodes In the meta path The semantic similarity under is a predefined similarity threshold; is a node The semantically relevant neighborhood of node Semantic similarity is higher than the threshold All nodes of In actual implementation, this application predefines a set of important meta-path patterns based on different application scenarios. Each meta-path reflects a specific type of semantic relationship. For example, in smart grid monitoring, you can define: Meta Path :Substation distribution station User area, which represents the power supply route; Meta Path :Substation Control Center Substation, indicating the equipment relationship of the shared control system.

[0069] Optionally, in complex heterogeneous systems, the present application may adopt a meta-path automatic discovery algorithm instead of relying solely on predefined meta-paths. Specifically, by analyzing the distribution of node types and relationship types in the system, potential important meta-path patterns are identified. The algorithm first generates all possible meta-path candidates, and then scores and screens the candidate paths based on indicators such as path instance frequency, node coverage, and semantic relevance, and finally retains the most representative set of meta-paths. This automatic discovery method is particularly suitable for scenarios with complex and dynamically changing system structures, such as smart city integrated monitoring platforms, where node and relationship types may increase or change over time.

[0070] In addition, during the meta-path random walk process, the present application can optionally implement a biased random walk strategy to dynamically adjust the walk probability based on the attributes of nodes and edges. For example, in power system monitoring, when the current node is in an abnormal state, the walk tends to select neighboring nodes with strong correlations with it, rather than uniform random selection; or in a specific time period, a specific type of relationship is preferentially selected for walks. This biased strategy can improve the semantic relevance of the walk path and more effectively discover potential abnormal propagation paths.

[0071] For each node, a random walk of a fixed length and number of walks is performed (for example, a random walk of length 10 is performed 100 times per node). In multimodal IoT monitoring systems, meta-path random walks can discover semantic associations across device types and subsystems. For example, for a temperature sensor node, a random walk of the meta-path "sensor-monitoring-device-impact-sensor" can identify other sensor nodes that have indirect impact relationships with it but may be physically distant. These nodes may play an important role in the anomaly propagation chain.

[0072] In some implementations, the present application can optionally combine meta-path random walks with knowledge graph technology to construct a semantic knowledge base between system components. By introducing external domain knowledge and historical event records, the semantic understanding capabilities of random walks can be enhanced. For example, in industrial production line monitoring, equipment manuals, maintenance records, and failure cases can be used to construct a knowledge graph to guide the generation and priority setting of meta-paths, making the discovered semantic associations more consistent with actual engineering experience and improving the interpretability of anomaly detection and prediction.

[0073] Therefore, by integrating the hyperedge attention mechanism and meta-path random walk, this step outputs a node representation that integrates the high-order relationship information of multiple nodes. .

[0074] It can be seen that this representation effectively captures the structural and semantic features in the heterogeneous graph, providing input for the subsequent establishment of anomaly propagation rate model.

[0075] Step 4: Based on the complex interaction characteristics, an anomaly propagation rate model is established to achieve accurate prediction of the anomaly diffusion path; According to another embodiment of the present application, this step models the propagation pattern of anomalies between different types of nodes by establishing a propagation rate tensor and a time-varying graph convolutional network, thereby achieving accurate prediction of the anomaly diffusion path and impact range.

[0076] Step 4.1, construct the propagation rate tensor; In this sub-step, the application defines the propagation rate tensor , used to describe the exception type Node Warp Relationship Propagate to Type The rate of the node. It should be noted that for each propagation rate element , learn by: ; in Is the node type Embedding vector of Is the relationship type Embedding vector of Is the node type Embedding vector of is a learnable parameter matrix; is the activation function; Represents vector concatenation operation; Indicates that the vector will be embedded 、 、 Splice into a long vector; represents a three-dimensional tensor space, where is the number of node types, is the number of relationship types; In the specific implementation, this application uses historical abnormal event data to train the propagation rate tensor. First, a sequence of abnormal events that have occurred in the system is collected, including the time, location, and type of the abnormal events. Then, for each pair of abnormal events with a causal relationship (i.e., one abnormality triggers another), the time delay between them is calculated to construct a labeled dataset of abnormal propagation. Based on this data, a gradient descent optimization is used to parameters so that the predicted propagation rate is as close as possible to the actually observed propagation delay.

[0077] Alternatively, when abnormal event data is limited, this application can employ transfer learning and semi-supervised learning methods to enhance the propagation rate model. Specifically, the model is first pre-trained in a similar domain or simulated environment, and then fine-tuned using a small amount of abnormal event data from the target system. Furthermore, the model is further optimized using self-supervised learning tasks (such as predicting future changes in node status) using unlabeled data. In newly established monitoring systems, this approach can effectively alleviate the "cold start" problem and accelerate the model's adaptation to the propagation patterns of a specific system.

[0078] In another implementation, this application can combine physical models and data-driven models to construct a hybrid propagation rate model. For systems with well-defined physical laws (such as power grids or water conservancy systems), domain knowledge can be used to define initial propagation laws, which can then be adjusted and optimized based on actual observational data. This hybrid approach maintains reasonable predictive power even when data is insufficient, while gradually improving accuracy as data accumulates.

[0079] In smart grid monitoring applications, the propagation rate tensor can represent the propagation law of abnormal states between different types of devices. For example, It might represent the propagation rate of an anomaly from type 1 (e.g., transformer) through relationship type 2 (e.g., power supply) to type 3 (e.g., distribution equipment). By analyzing historical data, the system might learn that transformer overheating anomalies typically propagate to downstream distribution equipment within 5-10 minutes, while voltage fluctuation anomalies may propagate within 2-3 minutes. These differentiated propagation patterns are accurately encoded in the propagation rate tensor.

[0080] Step 4.2, time-varying graph convolutional network; In addition, this application designs a time-varying graph convolutional network to capture the temporal characteristics of abnormal propagation. Specifically, for the time step ,node Abnormal status update: ; in Representation node In time abnormal state; Representation node In time abnormal state; is a node The set of neighbors of node In time abnormal state; Indicates the node Sum all neighbor nodes; It is the time-varying propagation coefficient of the time-varying graph convolutional network, which means that Slave nodes To Node The abnormal propagation intensity of ; the calculation formula is: ; in is the time-varying propagation coefficient, which means that Slave nodes To Node Abnormal transmission intensity; Is a slave node type Economic relationship type To Node Type The rate of spread; It is a gating function used to dynamically adjust the propagation coefficient according to the node status and time characteristics; It is a temporal feature representation, encoding time points characteristics; is a learnable parameter matrix; Indicates that the node Representation, node The representation and time features are concatenated into a vector; Representation node Type; Representation node and nodes the type of relationship between them; Representation node Type; In the implementation, this application uses a multi-layer stacked graph convolutional network, usually containing 3-4 layers, to capture the spatial propagation characteristics of anomalies layer by layer. It includes periodic encoding of timestamps (such as periodic features at the hour, day, week, etc.) and sequence position encoding to capture the time-varying pattern of anomaly propagation. It is implemented as a gating unit containing two layers of feedforward neural network, which is used to adaptively adjust the propagation coefficient according to the node status and time characteristics.

[0081] In some embodiments, the present application may optionally use a graph attention network (GAT) instead of a time-varying graph convolutional network to assign different attention weights to different neighbor nodes. Specifically, the time-varying propagation coefficient calculation is changed to: ; in is the time-varying propagation coefficient of the graph attention network, which represents the Slave nodes To Node Abnormal transmission intensity; Is a slave node type Economic relationship type To Node Type The rate of spread; is the attention weight, indicating the node For Node Attention intensity; It is a gating function used to dynamically adjust the propagation coefficient according to the node status and time characteristics; is a learnable parameter matrix; Indicates that the node Representation, node The representation and time features are concatenated into a vector.

[0082] This approach can better distinguish the degree of influence of different neighboring nodes on the abnormal state of the target node, improving prediction accuracy. For example, in urban water supply network monitoring, the weight of the impact of water quality anomalies at an upstream water plant on downstream pipeline nodes will be dynamically adjusted based on current flow rate, pressure, and other conditions, reflecting the heterogeneous propagation characteristics of the actual system.

[0083] Furthermore, for anomaly propagation processes with significant latency characteristics, this application optionally implements a latency modeling method based on a graph diffusion convolutional network (GDCN). This method treats anomaly propagation as a diffusion process over continuous time and accurately models the latency effects between nodes at different distances by solving the heat diffusion equation on the graph. In large-scale industrial production line monitoring, this method can accurately capture the propagation latency of anomalies from the source device to downstream devices at varying distances, enabling more precise time predictions.

[0084] In urban traffic flow monitoring applications, time-varying graph convolutional networks can model how traffic congestion propagates over time in a road network. For example, a node on a certain trunk road segment In time Congestion abnormality occurs ( The system uses a time-varying graph convolutional network to calculate how the congestion status will propagate to connected sections at different time points (e.g., 15 minutes later, 30 minutes later). By introducing time features ,The system can distinguish different patterns of congestion propagation during the morning and evening peaks, for example, morning peak congestion mainly spreads from residential areas to commercial areas, while the opposite is true during the evening peak.

[0085] Step 4.3, anomaly diffusion prediction; Then, based on the trained propagation rate model, this application predicts the abnormal diffusion path. First, using the current hypergraph state and the initial abnormal node set Trigger the simulation, then recursively calculate the future Abnormal status of each node within time: ; in is the anomaly propagation function, which comprehensively considers the current state of the node, the state of the neighbors, and the time-varying propagation coefficient; Representation node In time abnormal state; Representation node In the future The predicted abnormal state; Representation node All neighbor nodes at time The abnormal state set of Indicates time Slave nodes To Node The time-varying propagation coefficient of Indicates time arrive During this period, all neighbor nodes To Node The set of time-varying propagation coefficients of ; Representation node The set of neighbors of Indicates the time span of the forecast; Indicates the time offset, ranging from arrive ; Indicates the value range of the time offset; In actual implementation, this application uses a recursive rolling prediction strategy. First, the state of all nodes at the next moment is predicted. Then, the predicted results are used as new input to predict the state at the next moment, and this recursive process is repeated until the entire prediction time window is completed. To improve prediction stability, a decay mechanism and threshold control are introduced to prevent error accumulation in long sequence predictions.

[0086] In some embodiments, the present application may optionally use a Monte Carlo simulation method to perform probabilistic anomaly diffusion prediction. Through multiple random simulations, a probability distribution is generated for the abnormal state of each node at different time points, rather than a single determined value. This method can reflect the inherent randomness and uncertainty of the system and provide more comprehensive information for risk assessment. For example, in smart city energy network monitoring, for a certain substation, it may be predicted that "there is a 30% probability of a moderate anomaly and a 15% probability of a severe anomaly in 15 minutes", rather than simply predicting "an anomaly in 15 minutes".

[0087] In addition, for large-scale complex systems, the present application can optionally implement a hierarchical anomaly diffusion prediction method. First, a coarse-grained prediction is performed at the subsystem level to determine the main areas that the anomaly may affect; then a fine-grained node-level prediction is performed for these areas to improve computational efficiency. In large-scale smart city monitoring, this hierarchical approach can reduce computational complexity from Reduce to , enabling the system to process city-level monitoring data in real time.

[0088] In industrial IoT equipment failure prediction applications, anomaly diffusion prediction can accurately predict the propagation chain and impact range of equipment failures. For example, if an anomaly is detected in a chemical plant's temperature control system, the system can predict how the anomaly will affect the related pressure control system, flow control system, and downstream process equipment within the next 30 minutes, and calculate the probability and extent of impact on each device. This enables plant managers to take targeted measures in advance, such as reducing the load on critical equipment or preparing backup systems, to minimize production losses.

[0089] In addition, by calculating the abnormal state changes of each node at different time points, this application draws an abnormal propagation heat map , where the elements Representation node In time The probability of being affected by an anomaly.

[0090] Heat maps use color gradients to visually illustrate the propagation of anomalies, providing system managers with a global view. In smart city water supply system monitoring, if a water quality anomaly occurs at a water plant, the heat map can clearly show how the anomaly gradually spreads along the water supply network, as well as the probability of different areas being affected at different times, helping to accurately formulate emergency response plans.

[0091] Therefore, the output of this step is a complete anomaly propagation prediction result, including the abnormal state change trend, propagation path and affected probability of each node in the future time window, providing basic data for multi-level collaborative anomaly detection.

[0092] Step 5: Based on the anomaly propagation rate model and anomaly diffusion path, multi-level collaborative anomaly detection is implemented to identify collaborative failures across subsystems. According to a further embodiment of the present application, this step realizes comprehensive anomaly monitoring of complex systems from local to global levels by simultaneously performing anomaly analysis at three levels: node level, subgraph level, and hypergraph level, thereby effectively identifying collaborative failures across multiple subsystems.

[0093] Step 5.1, node-level anomaly detection; In this substep, for each node , based on its feature representation And the historical state sequence: ; in 、 、 Represents nodes respectively In time ,time ,time The historical state sequence, Indicates the length of the historical time window; Indicates the current time point; Calculate the node anomaly score: ; in, Representation node In time The anomaly score, Representation node go through The feature representation vector obtained after the layer graph neural network, is an anomaly scoring function that combines the following aspects: Time series pattern deviation: the deviation between the current state of a node and the historical pattern; feature space deviation: the distance between the node feature and the normal distribution; relationship consistency deviation: the difference between the node relationship pattern and the expected It should be noted that for different types of nodes, this application adopts specific anomaly scoring criteria and thresholds to improve detection accuracy.

[0094] In actual implementation, this application uses an integrated learning method to implement the anomaly scoring function Specifically, it combines multiple anomaly detection technologies, including: Reconstruction error metric based on autoencoder; density-based local outlier factor calculation; dynamic time warping distance based on time series; GNN anomaly scoring based on graph structure.

[0095] The final anomaly score is obtained by weighted fusion of these scores: ; in represents the final anomaly score of the node; is the weight of each scoring component; Indicates the The score generated by an anomaly detection technique; Indicates that all Sum the score components; represents the total number of anomaly detection techniques used; Optionally, to address uncertainty in anomaly scoring, this application may implement an anomaly scoring framework based on Bayesian deep learning. This approach not only outputs an anomaly score but also provides a confidence estimate, representing the degree of uncertainty in the score. Specifically, a Bayesian neural network (such as a feedforward network with Monte Carlo dropout) is used to model the parameter distribution of the scoring function, rather than a single point estimate. This approach is particularly valuable in monitoring critical infrastructure, such as nuclear power plants, where high-uncertainty anomaly detection results trigger a more cautious manual review process, while high-confidence results can lead to faster automated response measures.

[0096] Furthermore, in scenarios where system operating modes frequently change, this application optionally employs an adaptive anomaly detection strategy. This strategy uses an online learning mechanism to dynamically adjust anomaly detection model parameters and thresholds based on recently observed data. In scenarios with seasonal fluctuations (such as urban energy consumption monitoring), this adaptive approach can effectively address cyclical changes in normal modes and reduce false alarm rates.

[0097] In smart manufacturing scenarios, node-level anomaly detection can accurately identify abnormal conditions in individual devices. For example, in the semiconductor manufacturing process, the system simultaneously monitors multiple parameters, such as temperature, pressure, and light source stability, for key equipment nodes like lithography machines. This system then calculates a comprehensive anomaly score based on the device's historical operating patterns and its relationships with other devices. When a lithography machine's exposure system parameters fluctuate abnormally, even if the fluctuations are within a single parameter threshold, the multidimensional anomaly scoring function can still detect potential anomalies through correlation analysis, providing early warning.

[0098] Step 5.2, subgraph-level collaborative anomaly detection; In addition, the subgraph set defined based on hyperedges: ; in 、 、 Represents the first, second, and sub-graphs, is the total number of subgraphs in the system; This application calculates the subgraph-level anomaly score: ; in is the subgraph collaborative anomaly scoring function; It is a subgraph Description of the structural characteristics; Representing a subgraph In time Anomaly score; Representing a subgraph All nodes in time The set of anomaly scores; Indicates the subgraphs; Representation node Belongs to subgraph ; It should be understood that the subgraph collaborative anomaly score not only considers the combination of node anomaly scores, but also considers whether the interaction pattern between nodes is abnormal, and is calculated using the following formula: ; in is the subgraph collaborative anomaly scoring function; Representing a subgraph All nodes in time The set of anomaly scores; is a node Importance weight of is a measure of node interaction anomaly within a subgraph; is the balance coefficient; Represents a subgraph Sum all nodes in ; In the implementation, this application builds a subgraph-level anomaly detection model based on graph neural network. First, the normal structural pattern of the subgraph is learned using the graph autoencoder, and then the deviation between the current subgraph structure and the learned normal pattern is measured to calculate the interaction anomaly metric. Importance weight Determined by node centrality indicators (such as eigenvector centrality or PageRank value), the impact of key nodes on the overall subgraph anomaly is highlighted.

[0099] Alternatively, in application scenarios with strong domain knowledge, this application can implement a knowledge-guided subgraph anomaly detection method. This method incorporates domain rules and expert experience to provide a priori constraints for subgraph anomaly scoring. For example, in chemical production process monitoring, constraints related to changes in the concentration of specific substances are defined based on chemical reaction principles. When the node state in the subgraph violates these constraints, the anomaly score is increased. This method can combine domain knowledge with data-driven detection, improving detection accuracy and interpretability.

[0100] In another embodiment, for subgraph anomaly detection in large-scale systems, the present application can adopt hierarchical clustering and hierarchical anomaly detection methods. First, nodes are clustered into subgraphs of different levels according to the system topology and functional relationships, and then anomaly detection and propagation analysis are performed step by step from the bottom subgraph upwards. This method can effectively deal with the computational complexity problem in large-scale systems. For example, in a city-level monitoring system containing tens of thousands of nodes, the computational complexity can be reduced from 100 to 1000. Reduce to .

[0101] In smart city power system monitoring, subgraph-level collaborative anomaly detection can identify coordinated failure patterns within a regional power grid. For example, a regional power grid subgraph might contain a substation node and multiple distribution station nodes. Even if the individual anomaly scores for each node are low, if they exhibit coordinated behavior inconsistent with normal load distribution patterns (e.g., stable voltage at a substation but inconsistent voltage fluctuations at downstream distribution stations), subgraph-level detection can identify this coordinated anomaly and prevent regional power failures caused by load imbalance.

[0102] Step 5.3, hypergraph-level system anomaly identification; Next, based on the global hypergraph structure and all detected node-level and subgraph-level anomalies, this application calculates a system-level anomaly score: ; in is the system-level anomaly scoring function; Indicates that all nodes at time The set of anomaly scores; Represents all subgraphs at time The set of anomaly scores; represents the entire hypergraph; Represents the set of all nodes in the hypergraph; Representing a subgraph It is a hypergraph A subset of The entire system at time Anomaly score; System-level anomaly scoring function Comprehensive consideration: Spatial distribution pattern of abnormal nodes; topological correlation characteristics of abnormal subgraphs; and temporal evolution trend of anomalies. It should be emphasized that system-level anomalies not only reflect the aggregation of single-point anomalies, but more importantly, capture the coordinated failure modes across subsystems, which are usually difficult to detect in local monitoring.

[0103] In implementation, this application uses a hierarchical attention mechanism to implement a system-level anomaly scoring function. First, a node-subgraph-system hierarchy is constructed. Then, a cross-level attention network is used to adaptively aggregate anomaly information at different levels. For the spatial distribution of abnormal nodes, a spatial clustering algorithm is used to detect abnormal clustering patterns. For the temporal evolution of anomalies, a recurrent neural network is used to capture the dynamic changes in the system's abnormal state.

[0104] Optionally, to enhance the explainability of system-level anomaly identification, the present application may implement an anomaly identification method enhanced by causal reasoning. This method constructs a causal graph between system components, analyzes the causal relationship between abnormal events, and identifies the root cause and impact path. Specifically, a structural causal model (such as an intervention algorithm or counterfactual analysis) is used to infer the propagation link of the anomaly and distinguish between primary and secondary anomalies. In complex industrial system monitoring, this method can help engineers quickly locate the root cause of the anomaly, rather than just discovering the surface symptoms, greatly improving the efficiency of fault recovery.

[0105] In another implementation, for scenarios requiring real-time response, this application can employ a progressive anomaly identification strategy. This strategy first rapidly assesses the system status and provides a preliminary anomaly determination, then gradually refines the analysis in the background and continuously updates the anomaly assessment results. For example, in an urban traffic management system, a preliminary traffic anomaly alert can be issued within seconds, followed by a progressively more detailed analysis of the anomaly's root cause, impact, and recommended measures over the next few minutes, supporting managers' hierarchical response decisions.

[0106] In large-scale industrial park monitoring applications, hypergraph-level system anomaly recognition can identify complex coordinated failures across multiple subsystems. For example, in a petrochemical industrial park, when minor anomalies simultaneously occur in the energy supply system, raw material processing system, and product production system, while the severity of each individual anomaly is insufficient to trigger an alarm, system-level analysis can identify that this cross-system coordinated anomaly may indicate a significant risk to the overall production process. By calculating a global anomaly score, the system promptly issues an early warning to park managers, preventing potential chain reactions that could lead to production accidents.

[0107] Step 5.4, cascade risk assessment; In addition, based on the anomaly propagation model and multi-level anomaly detection results, this application assesses the risk of each part of the system being affected by anomalies: ; in is the conditional probability function; is the abnormal threshold; Representation node In the future Value at risk; Representation node In the future Anomaly score; Indicates the current time The set of anomaly scores for all nodes; Indicates the time interval for prediction.

[0108] In implementation, this application uses a combination of Bayesian networks and Monte Carlo simulations to implement cascade risk assessment. First, a Bayesian network is constructed based on historical anomaly propagation data to capture the conditional dependencies between nodes. Then, using the currently detected anomaly state as observational evidence, probabilistic reasoning is used to calculate the anomaly probability distribution of each node in the future. Finally, through multiple Monte Carlo simulations, possible anomaly propagation scenarios are generated to assess the risk level of different nodes and subsystems.

[0109] Optionally, for critical infrastructure monitoring requiring high reliability, this application can implement a multi-scenario risk assessment approach. This approach not only evaluates the most likely anomaly propagation path, but also analyzes the risks under various extreme and edge scenarios. By considering multiple possible interventions and changes in external conditions, it provides decision makers with a more comprehensive view of the risks. For example, in urban water supply system monitoring, the system will evaluate the potential propagation paths and impact range of water quality anomalies under various scenarios, including normal, high-load, and emergency operations, supporting the development of robust emergency response plans.

[0110] In another implementation, this application can employ reinforcement learning to optimize risk response strategies. The system not only predicts the risk of anomaly propagation but also uses reinforcement learning algorithms to learn the optimal timing and method of intervention to minimize overall risk. In smart grid monitoring, when a voltage anomaly is detected in a specific area, the system can recommend the optimal load adjustment strategy, striking a balance between ensuring power supply reliability and system safety.

[0111] In intelligent building management systems, cascading risk assessments predict how an initial anomaly will impact various subsystems throughout the building. For example, when an abnormal refrigerant pressure in the air conditioning system is detected, the system not only assesses the risk of the air conditioning subsystem itself but also analyzes potential cascading anomalies in the air supply system, the fresh air system, and the building's energy management system. The risk assessment results are presented as a risk heat map, helping building managers prioritize and prevent potential problems before they occur.

[0112] Therefore, through cascading risk assessment, system managers can obtain the following key information: Accurate location of the anomaly source; the diffusion path and rate of the anomaly impact; the probability and degree of impact on each part of the system; and the risk level of key nodes and subsystems. It can be seen that the output of this step is a complete multi-level anomaly detection report, including information such as anomaly source location, anomaly propagation path, collaborative failure pattern identification, and risk level assessment, providing system managers with a comprehensive view of the anomaly status.

[0113] An intelligent data monitoring platform, used to execute the above-mentioned intelligent data monitoring method, comprising: Dynamic heterogeneous hypergraph construction module, used to build a hypergraph model representing the complex relationships between multi-source data nodes in the system; High-order relational computing module, used to implement multi-dimensional tensor representation and decomposition, and obtain embedded representations of nodes and relations; Heterogeneous graph information transfer module, used to capture complex interaction features between multiple nodes through hyperedge attention mechanism and meta-path random walk; Anomaly propagation modeling module, which is used to build a propagation rate tensor and a time-varying graph convolutional network to accurately predict anomaly diffusion paths; The multi-level anomaly detection module is used to perform anomaly analysis at the node level, subgraph level, and hypergraph level to identify collaborative failures across subsystems.

[0114] Here, the present invention provides an implementation example: In order to verify the effectiveness and practicality of the intelligent data monitoring method proposed in this application, an application example in smart city power system monitoring is given below.

[0115] This application example focuses on a city's smart grid system, which consists of 57 substation nodes (Type A1), 213 distribution station nodes (Type A2), 1,268 user area nodes (Type A3), and 92 monitoring center nodes (Type A4). Multiple types of relationships exist between nodes, including power supply relationships (R1), control relationships (R2), and monitoring relationships (R3). The system needs to monitor the grid's operating status in real time, predict potential abnormality propagation paths, and promptly identify coordinated faults across subsystems.

[0116] First, the system collects various monitoring data, including voltage, current, power factor, equipment temperature and other parameters, and generates the initial node feature vector after standardization.

[0117] Next, based on domain knowledge and historical data analysis, the system identifies various high-order relationships and constructs hyperedges. For example, a main substation, its five connected distribution substations, and the corresponding user areas form a hyperedge, representing a complete power supply chain. Multiple substations under the same monitoring center form another type of hyperedge, representing control domain relationships. Based on a week of historical data, a 10-minute time window and a 1-minute sliding step are used to generate a time-varying hypergraph sequence, capturing the time-varying characteristics of the grid load.

[0118] Then, the system represents the high-order relationship as a multi-dimensional tensor. For the ternary relationship of substation-distribution station-user area, the tensor is constructed For complex structures containing control relationships, a hierarchical decomposition approach is used. A tensor decomposition algorithm is applied, with the embedding dimension set to k = 64. Through alternating optimization, embedding representations for various nodes and relationships are obtained.

[0119] In the information transmission and aggregation stage, the system implements a 3-layer stacked hyperedge attention network, with 6 attention heads in each layer. At the same time, the predefined key meta-paths are as follows P1: Substation → Power supply → Distribution station → Power supply → User area; P2: Substation → Control → Monitoring Center → Control → Substation; We perform 100 random walks of length 10 for each node to generate semantically relevant neighborhoods.

[0120] The system trains an anomaly propagation rate model based on historical anomaly event data (including 156 equipment overheating events, 89 voltage anomaly events, and 72 load imbalance events). For example, it learns that substation overheating anomalies typically propagate to downstream distribution stations within 7.5 minutes, while voltage fluctuation anomalies propagate within an average of 3.2 minutes. A four-layer time-varying graph convolutional network is used to model the temporal characteristics of anomaly propagation, with temporal features including hourly cycles, weekdays / holidays, and load level encoding.

[0121] Finally, the system implements collaborative anomaly detection at three levels: at the node level, an ensemble learning approach is used, combining techniques such as autoencoder reconstruction error and local outlier factor calculation; at the subgraph level, graph autoencoders are used to learn normal structural patterns; and at the system level, a hierarchical attention mechanism is used to integrate anomaly information from different levels. Furthermore, the system implements cascaded risk assessment based on Bayesian networks and Monte Carlo simulations to predict the anomaly propagation path and impact range.

[0122] During the three months of deployment, the system detected 22 potential coordination failures, 17 of which were confirmed by system administrators, for a detection accuracy rate of 77.3%. More importantly, the system successfully predicted the propagation paths of all high-risk failures, issuing warnings an average of 18.7 minutes in advance, providing system administrators with ample response time.

[0123] The data on the improvement of abnormal prediction ability is shown in Table 1: Table 1: Abnormal prediction ability improvement data

[0124] The system response effect improvement data is shown in Table 2: Table 2: System response effect improvement data

[0125] Data shows that this system achieved a nearly threefold improvement in anomaly prediction lead time, significantly enhancing system administrators' response capabilities. Furthermore, the collaborative fault identification rate increased from 29.5% to 77.3%, demonstrating the method's strength in detecting complex cross-subsystem fault patterns. Furthermore, the system's average recovery time was shortened by 63.5%, directly reducing the impact of service disruptions caused by anomalies.

[0126] In a typical case, the system detected a slight anomaly in a temperature parameter at a main substation. While this did not trigger a traditional threshold alarm, high-level relationship analysis and propagation prediction enabled the system to identify this as a potential load imbalance within the connected distribution substation network, potentially impacting power supply stability in the user area within 15 minutes. The system-generated anomaly propagation heat map accurately displayed the potentially affected areas, enabling managers to proactively adjust load distribution and completely avoid potential widespread power supply fluctuations.

[0127] Practical application results show that the intelligent data monitoring method based on dynamic heterogeneous hypergraph and high-order relationship representation proposed in this application can effectively capture multi-node collaborative anomaly patterns in complex systems, achieve accurate prediction of anomaly diffusion paths, and provide system managers with a comprehensive abnormal status view and decision support.

[0128] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. An intelligent data monitoring method, characterized in that: include: Construct a dynamic heterogeneous hypergraph model to represent the complex relationships and temporal evolution characteristics between multi-source data nodes; Based on the dynamic heterogeneous hypergraph model, high-order relationship representation and calculation are realized to obtain the embedded representation of nodes and relationships; Leveraging embedded representations of nodes and relationships, we perform heterogeneous graph information transfer and aggregation, capturing complex interaction features between multiple nodes. Based on the complex interaction characteristics, an abnormal propagation rate model is established to achieve accurate prediction of the abnormal diffusion path; Based on the anomaly propagation rate model and anomaly diffusion path, multi-level collaborative anomaly detection is implemented to identify collaborative failures across subsystems.

2. The intelligent data monitoring method according to claim 1, characterized in that: The steps of constructing a dynamic heterogeneous hypergraph model include: Standardize various monitoring data sources in the system, extract features and generate corresponding node representations; Analyze the collaborative relationship patterns between multiple nodes, identify sets of nodes with strong correlations, and construct hyperedges; Through time window sliding and snapshot sampling methods, the temporal evolution characteristics of the hypergraph structure are captured to form a time-varying hypergraph sequence.

3. The intelligent data monitoring method according to claim 1, characterized in that: The steps of implementing high-order relation representation and calculation include: Representing high-order relations in heterogeneous hypergraphs as multi-dimensional tensors; Using tensor decomposition algorithm, the high-dimensional relationship tensor is decomposed into a low-dimensional factor matrix to obtain the embedded representation of nodes and relationships; Combining timing information and heterogeneity constraints to optimize embedding representations.

4. The intelligent data monitoring method according to claim 1, characterized in that: The steps of performing heterogeneous graph information transmission and aggregation include: Designing attention weights for each hyperedge enables the system to adaptively aggregate information from different nodes; Design a type-aware information aggregation function based on the characteristics of heterogeneous graphs; Generate semantically relevant neighborhoods of nodes through meta-path based random walks.

5. The intelligent data monitoring method according to claim 1, characterized in that: The steps of establishing the abnormal propagation rate model include: Define a propagation rate tensor to describe the rate at which anomalies propagate from one type of node to another type of node through a specific relationship; Design a time-varying graph convolutional network to capture the temporal characteristics of anomaly propagation; Based on the trained propagation rate model, the abnormal diffusion path is predicted.

6. The intelligent data monitoring method according to claim 1, characterized in that: The steps of implementing multi-level collaborative anomaly detection include: Calculate node-level anomaly scores based on node feature representation and historical state sequences; Calculate subgraph-level anomaly scores based on the set of subgraphs defined by hyperedges; Calculate a system-level anomaly score based on the global hypergraph structure and all detected node-level and subgraph-level anomalies; Based on the anomaly propagation model and multi-level anomaly detection results, the risk of each part of the system being affected by anomalies is assessed.

7. The intelligent data monitoring method according to claim 3, characterized in that: The high-order relations are decomposed into a number of low-order relation combinations in a hierarchical decomposition manner; when there are a large number of sparse high-order relations in the system, a sparse tensor storage and calculation method is adopted.

8. The intelligent data monitoring method according to claim 4, characterized in that: The attention weights of each hyperedge design adopt a multi-layer stacked hyperedge attention network structure, and through the multi-head attention mechanism, multiple groups of independent attention weights are calculated in parallel to enrich the expression ability of the model.

9. The intelligent data monitoring method according to claim 5, characterized in that: The prediction of the abnormal diffusion path adopts a recursive rolling prediction strategy, which first predicts the state of all nodes at the next moment, and then uses the prediction result as a new input to predict the state at the next moment until the entire prediction time window is completed; and by calculating the abnormal state changes of each node at different time points, a heat map of abnormal propagation is drawn.

10. An intelligent data monitoring platform, characterized in that: An intelligent data monitoring method for executing any one of claims 1 to 9, comprising: Dynamic heterogeneous hypergraph construction module, used to build a hypergraph model representing the complex relationships between multi-source data nodes in the system; High-order relational computing module, used to implement multi-dimensional tensor representation and decomposition, and obtain embedded representations of nodes and relations; Heterogeneous graph information transfer module, used to capture complex interaction features between multiple nodes through hyperedge attention mechanism and meta-path random walk; Anomaly propagation modeling module, which is used to build a propagation rate tensor and a time-varying graph convolutional network to accurately predict anomaly diffusion paths; The multi-level anomaly detection module is used to perform anomaly analysis at the node level, subgraph level, and hypergraph level to identify collaborative failures across subsystems.

Citation Information

Patent Citations

  • Dynamic graph anomaly detection method based on hypergraph contrast learning

    CN120067950A

  • Construction method of multidimensional time sequence anomaly detection system

    CN120277571A

  • Abnormality prediction system, method, and computer program

    WO2023210518A1

Cited By

  • Industrial equipment state sensing method based on PageRank graph calculation algorithm

    CN120974383A

  • An industrial equipment state perception method based on a PageRank graph calculation algorithm

    CN120974383B

  • Road maintenance construction monitoring system and method based on vehicle-road cooperation

    CN121303558A

  • A highway maintenance construction monitoring system and method based on vehicle-road cooperation

    CN121303558B

  • Data transmission security method based on graph nerve detection

    CN121462302A