Abnormal event detection method and system based on multi-source operation and maintenance data fusion
By fusing multi-source operation and maintenance data and using graph neural network technology, a dynamic abnormal event map is constructed, which solves the problem of inaccurate multi-source data correlation analysis, realizes accurate location and root cause analysis of abnormal events, and improves operation and maintenance efficiency and accuracy.
Patent Information
- Application Number
- CN202511328618.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing technologies neglect the inherent connections between data in multi-source operation and maintenance data analysis, resulting in inaccurate correlation analysis of abnormal events and an inability to comprehensively and accurately reflect the actual operating status of the system.
An anomaly detection method based on multi-source operation and maintenance data fusion is adopted. By collecting and preprocessing multi-source operation and maintenance data, a dynamically updated anomaly map is constructed. Graph neural network technology is used to perform cross-type feature aggregation and differentiated causal inference. The map is then corrected by combining user feedback information to achieve accurate location and root cause analysis of anomalies.
It improves the accuracy and real-time performance of abnormal event detection, reduces the false alarm rate, enhances the system's ability to detect and handle abnormal events, and meets the real-time response requirements of operation and maintenance.
Smart Images

Figure CN120803804B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing and abnormal event detection, specifically to an abnormal event detection method and system based on multi-source operation and maintenance data fusion. Background Technology
[0002] With the advancement of informatization, the volume of operational data has increased dramatically, making the fusion and analysis of multi-source heterogeneous data crucial for improving operational efficiency. In today's complex operational environment, the development of information technology has resulted in a wide range of operational data sources and diverse formats. This massive amount of data contains vital information about the system's operational status. Effective analysis and utilization of this data can promptly identify potential problems in the system, allowing for proactive measures to prevent failures, thereby significantly improving operational efficiency and ensuring the stable operation of the system.
[0003] Traditional anomaly correlation analysis primarily relies on two methods. One is rule matching, which involves pre-setting a series of rules; when data matches these rules, it's considered an anomaly. This method is simple, direct, and easy to implement, and can be effective in some simple operational scenarios. The other is statistical analysis of a single data source, which involves performing statistical calculations on data from a single data source, such as mean and variance, to determine if the data is anomaly. This method is more commonly used when the data volume is small and the data relationships are relatively simple.
[0004] However, existing technologies have significant drawbacks. They often overlook the inherent relationships between multi-source data, analyzing only single-source data in isolation or performing simple cross-source data comparisons. This leads to inaccurate correlation analysis of abnormal events and an inability to comprehensively and accurately reflect the actual operating status of the system. Summary of the Invention
[0005] To achieve accurate detection of abnormal events, this application provides an abnormal event detection method and system based on the fusion of multi-source operation and maintenance data.
[0006] Firstly, this application provides an anomaly event detection method based on multi-source operation and maintenance data fusion, including:
[0007] Collect multi-source operation and maintenance data and perform preprocessing; the preprocessing includes noise reduction, data alignment and completion.
[0008] Based on the pre-built mapping relationship between multi-source operation and maintenance data and abnormal event types, the abnormal event types mapped by the pre-processed multi-source operation and maintenance data are determined, and the differential processing of multi-source operation and maintenance data driven by abnormal event types is completed to obtain a dynamically updated abnormal event map. The differential processing process includes: feature extraction based on the abnormal event type matching feature engine, relationship modeling based on the correlation relationship of operation and maintenance data matched by abnormal event type, generating map update instructions, and completing the dynamic update of the abnormal event map based on the map update strategy matched by abnormal event type.
[0009] Based on a dynamically updated abnormal event graph, graph neural network technology is used to aggregate cross-type features of all nodes in the graph, outputting the initial values of the abnormal probability of all nodes and the initial values of the probability distribution of each type of abnormal event. This allows for a preliminary determination of whether an abnormal event exists, its type, and its level. Differential causal reasoning driven by the abnormal event type is then completed to obtain the final judgment of the abnormal event and the root cause location.
[0010] Statistical analysis is performed on the finally identified abnormal events, and the statistical content of abnormal events within the recent preset time period is automatically displayed or can be queried and retrieved, including: the proportion of various types of abnormal events and the proportion of abnormality level of each type of abnormal event.
[0011] By adopting the above scheme, multi-source operation and maintenance data can be differentiated based on mapping relationships to obtain dynamically updated abnormal event maps, thereby uncovering potential correlations between multi-source data. Cross-type feature aggregation can be performed using graph neural network technology, and differentiated causal reasoning based on abnormal event type can be applied to accurately locate abnormal events and their root causes. Statistical analysis of the finally identified abnormal events can be performed to intuitively present the proportion of various abnormal events and the proportion of abnormal levels, ensuring the accuracy and real-time nature of abnormal event correlation analysis and improving operation and maintenance efficiency.
[0012] Preferred options also include:
[0013] Based on the statistical data of abnormal events within a preset time period, determine the types of abnormal events that are high-frequency abnormalities;
[0014] In the process of differentiated processing of multi-source operation and maintenance data driven by high-frequency abnormal event types, a first acceleration strategy is initiated to obtain a sub-graph of the abnormal event graph that is dynamically updated at an accelerated pace. The first acceleration strategy includes: extracting key features based on the feature engine matching the abnormal event types that are high-frequency abnormal, performing relationship modeling based on the relationship between the operation and maintenance data matching the abnormal event types that are high-frequency abnormal, removing relationship edges with weights lower than a preset threshold, generating a graph update instruction, and completing the accelerated dynamic update of the sub-graph of the abnormal event graph based on the graph update strategy matching the abnormal event types that are high-frequency abnormal.
[0015] Based on the sub-graph of the dynamically updated anomaly event graph, a second acceleration strategy is initiated during the process of completing differentiated causal reasoning driven by anomaly event type to obtain the final identified anomaly event and root cause localization. A confidence evaluation function is designed based on the root cause localization of anomaly event type to complete the confidence evaluation of root cause localization and output the root cause localization with a confidence level higher than the preset confidence level. The second acceleration strategy adopts a pruning reasoning method to accelerate causal reasoning, including retaining only the nodes with the top-K correlation with the anomaly node to participate in the root cause calculation.
[0016] By adopting the above scheme, for high-frequency abnormal event types, the processing speed of abnormal event map update and causal reasoning is improved by using the first and second acceleration strategies. At the same time, the confidence assessment is combined to output high-confidence root cause localization, thereby enhancing the efficiency and accuracy of abnormal event detection and root cause localization in high-frequency abnormal scenarios.
[0017] Preferred options also include:
[0018] Design a structured root cause feedback template and receive structured root cause feedback information for different types of abnormal events uploaded by users;
[0019] Analyze root cause feedback information to obtain causal paths in order to correct the abnormal event graph, including: enhancing the edge weights between nodes in the abnormal event graph that conform to the obtained causal path, adding new causal paths between nodes in the abnormal event graph, and reducing the edge weights between nodes in the abnormal event graph that contradict the obtained causal path.
[0020] Analyzing root cause feedback information to obtain causal paths in order to adjust the abnormal probability threshold or causal inference threshold of abnormal event nodes includes: determining whether each node in the abnormal event root cause localization is included in each node in the causal path obtained based on root cause feedback information; determining that nodes in the abnormal event root cause localization are not included in each node in the causal path obtained based on root cause feedback information, and whose occurrence frequency is greater than the preset frequency within a preset time, identifying such nodes as the first high-frequency false alarm nodes in the abnormal event root cause localization, and reducing the node abnormal probability threshold of the first high-frequency false alarm nodes.
[0021] After determining that each node in the root cause localization of anomalies is included in the nodes of the causal path obtained based on root cause feedback information, the similarity between each node in the root cause localization of anomalies and each node in the causal path obtained based on root cause feedback information is compared. If the similarity between the node and each node in the causal path obtained based on root cause feedback information is less than a preset similarity, and the frequency of occurrence of the similarity between the node and each node in the causal path obtained based on root cause feedback information being less than the preset similarity is greater than a preset frequency within a preset time, the node is identified as the second most frequent false alarm node in the root cause localization of anomalies, and the node causal inference threshold of the second most frequent false alarm node is reduced.
[0022] By adopting the above scheme, designing a structured root cause feedback template and receiving structured root cause feedback information uploaded by users, the actual situation of the correlation between multi-source operation and maintenance data and abnormal events can be obtained; the root cause feedback information can be analyzed to obtain causal paths to correct the abnormal event graph, making the graph more accurately reflect the relationship between abnormal events; the root cause feedback information can be analyzed to adjust the abnormal probability threshold or causal inference threshold of abnormal event nodes, reduce the abnormal probability threshold of high-frequency false alarm nodes, reduce false alarms, and improve the accuracy of abnormal event detection and the reliability of causal inference.
[0023] Preferred options also include:
[0024] In the confidence assessment of statistical root cause localization, the proportion of root cause localizations with output confidence scores higher than the preset confidence score is used to determine the percentage of all output confidence scores. When the determined statistical proportion is less than the preset proportion, a circuit breaker protection instruction is generated. This instruction is used to disable the first acceleration strategy during the differentiated processing of multi-source operation and maintenance data driven by abnormal event types that are in high-frequency anomalies within the preset circuit breaker protection period.
[0025] By adopting the above scheme, when the proportion of high-confidence results in root cause localization is lower than the preset proportion, a circuit breaker protection command can be generated in a timely manner and the first acceleration strategy can be turned off, avoiding the continued use of the acceleration strategy when the results are unreliable, and ensuring the accuracy and reliability of abnormal event detection results.
[0026] Preferred options also include:
[0027] Statistical analysis is performed on the finally identified abnormal events to obtain the proportion of different abnormal events within a preset time period; trigger conditions are set based on the obtained proportion of different abnormal events, and resource allocation is dynamically adjusted based on the differentiated processing of multi-source operation and maintenance data driven by abnormal event type and the differentiated causal reasoning driven by abnormal event type when resource allocation is triggered.
[0028] The triggering conditions include: periodic triggering, timely triggering, and emergency triggering;
[0029] The periodic triggering involves setting a preset adjustment period, calculating the proportional changes of different types of abnormal events within the preset adjustment period, matching the calculated proportional changes with preset abnormal event priority adjustment rules, triggering the adjustment of the priorities of the corresponding different types of abnormal events, and determining the resource allocation weights of different abnormal events according to the adjusted priorities of the different types of abnormal events to achieve the adjustment of resource allocation weights; the preset abnormal event priority adjustment rules include increasing the priority level of the corresponding type of abnormal event when the calculated proportional change is greater than a preset proportional threshold range.
[0030] The timely triggering refers to monitoring abnormal event types where the proportion of abnormal events exceeds the proportion of abnormal event threshold, directly triggering the increase of the priority of the corresponding type of abnormal event, and determining the resource allocation weight of different abnormal events according to the adjusted priority of different types of abnormal events to achieve the adjustment of resource allocation weight.
[0031] The emergency triggering process involves receiving instructions to adjust the priorities of different types of abnormal events, adjusting the priorities of the corresponding different types of abnormal events according to the instructions, and determining the resource allocation weights of different abnormal events based on the adjusted priorities of the different types of abnormal events to achieve the adjustment of resource allocation weights.
[0032] By adopting the above scheme, statistical analysis is performed on the final identified abnormal events to obtain the proportion of different abnormal events within a preset time period. Based on this, trigger conditions are set, and the resource allocation for differentiated processing and causal reasoning can be dynamically adjusted according to the proportion of different abnormal events when triggering resource allocation. At the same time, the priority of abnormal events and the weight of resource allocation are adjusted through three triggering methods: periodic, timely, and emergency, so that the resource allocation is more reasonable to adapt to different abnormal event situations and improve the efficiency and accuracy of abnormal event detection and processing.
[0033] Preferably, determining the resource allocation weights for different types of abnormal events according to the adjusted priorities of different types of abnormal events includes:
[0034] Based on the priority of different types of abnormal events, resource allocation weights are preset and matched with different priorities;
[0035] Total resources are allocated according to the resource allocation weights of different abnormal events;
[0036] Continue to adjust and execute the resource allocation for graph update, feature aggregation, and causal reasoning resources for a single type of abnormal event according to the pre-defined graph update resource ratio, feature aggregation resource ratio, and causal reasoning resource ratio.
[0037] By adopting the above scheme, pre-setting resource allocation weights based on the priority of abnormal events and setting a minimum weight, allocating total resources in combination with different abnormal event resource allocation weights, and then adjusting the resource allocation for a single type of abnormal event according to a preset ratio, resources can be dynamically and reasonably allocated according to the priority of abnormal event types, thereby improving the efficiency and accuracy of abnormal event detection and processing.
[0038] Preferably, the cross-type feature aggregation of all nodes in the graph using graph neural network technology includes: weight allocation based on the centrality of the graph structure for dynamically updated abnormal event graphs.
[0039] By adopting the above scheme, weight allocation is performed when performing cross-type feature aggregation on all nodes in the graph based on the centrality of the graph structure. This allows for a more reasonable consideration of the status and role of each node in the graph structure, thereby making the judgment of abnormal events, their types, and levels more accurate and improving the accuracy of root cause localization.
[0040] Secondly, this application provides an abnormal event detection system based on multi-source operation and maintenance data fusion, including:
[0041] The multi-source operation and maintenance data acquisition and processing module is used to acquire multi-source operation and maintenance data and perform preprocessing; the preprocessing includes noise reduction, data alignment and completion.
[0042] The feature map construction and update module is used to determine the abnormal event types mapped by the pre-built mapping relationship between multi-source operation and maintenance data and abnormal event types, complete the differential processing of multi-source operation and maintenance data driven by abnormal event types, and obtain a dynamically updated abnormal event map. The differential processing process includes: feature extraction based on the abnormal event type matching feature engine, relationship modeling based on the correlation relationship of operation and maintenance data based on the abnormal event type matching, generating map update instructions, and completing the dynamic update of the abnormal event map based on the map update strategy of abnormal event type matching.
[0043] The abnormal event root cause acquisition module is used to perform cross-type feature aggregation on all nodes in the graph based on the dynamically updated abnormal event map. It uses graph neural network technology to output the initial value of the abnormal probability of all nodes and the initial value of the probability distribution of each type of abnormal event. It makes a preliminary judgment on whether there is an abnormal event, the type of abnormal event and the level of abnormality. It completes differentiated causal reasoning based on the abnormal event type and obtains the final judgment of the abnormal event and the root cause location.
[0044] The abnormal event display and query module is used to perform statistical analysis on the finally determined abnormal events, automatically display or support querying to obtain the statistical content of abnormal events within a preset time period, including: the proportion of various types of abnormal events and the proportion of abnormal level of each type of abnormal event.
[0045] By adopting the above scheme, a multi-source operation and maintenance data acquisition and processing module is designed to preprocess multi-source operation and maintenance data by denoising, data alignment, and completion; a feature map construction and update module is designed to perform differentiated processing based on the mapping relationship between multi-source operation and maintenance data and abnormal event types, complete feature extraction, relationship modeling, and dynamic map update, and obtain dynamically updated abnormal event maps; an abnormal event root cause acquisition module is designed to use graph neural network technology to perform cross-type feature aggregation on map nodes, complete differentiated causal reasoning, and achieve accurate location of abnormal events; an abnormal event display and query module is designed to perform statistical analysis on the finally judged abnormal events, automatically display or support querying the proportion of various types of abnormal events and the proportion of abnormality level of each type of abnormal event within a preset time period.
[0046] Thirdly, this application provides a computer-readable storage medium including a stored computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to perform the method described above.
[0047] Fourthly, this application provides a computer device, the computer device including a memory, a processor and a program stored in the memory and executable thereon, the program being executed by the processor to implement the steps of the method described above.
[0048] In summary, this application has the following beneficial effects:
[0049] 1. Based on the mapping relationship, complete the differentiated processing of multi-source operation and maintenance data, obtain dynamically updated abnormal event graphs, and be able to mine potential data correlations from multiple dimensions and accurately locate abnormal events; use graph neural network technology to aggregate cross-type features of graph nodes, perform differentiated causal reasoning and root cause localization, ensure the accuracy of abnormal event correlation analysis, and meet the real-time response requirements of operation and maintenance.
[0050] 2. Considering the types of high-frequency abnormal events, the first and second acceleration strategies are used to accelerate the updating of the abnormal event map and the causal reasoning process, thereby improving the efficiency of abnormal event detection and root cause localization. At the same time, the confidence of root cause localization is evaluated to ensure that the output results have high credibility.
[0051] 3. By utilizing structured root cause feedback information uploaded by users, the abnormal event graph is corrected and the thresholds are adjusted, which improves the accuracy of abnormal event detection, reduces the false alarm rate, makes the correlation analysis of abnormal events more consistent with the actual situation, and enhances the system's ability to detect and handle abnormal events. Attached Figure Description
[0052] Figure 1 This is a flowchart of the abnormal event detection method based on multi-source operation and maintenance data fusion described in a specific embodiment;
[0053] Figure 2 This is a graph showing the abnormal event query results of the abnormal event detection method based on multi-source operation and maintenance data fusion applied in a specific embodiment.
[0054] Figure 3 This is a schematic diagram of the anomaly event detection system based on multi-source operation and maintenance data fusion as described in a specific embodiment;
[0055] Figure 4 This is a schematic diagram of the abnormal event detection system based on multi-source operation and maintenance data fusion described in a specific embodiment. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0057] This application mainly adopts multi-source operation and maintenance data fusion to detect abnormal events and locate root causes, thereby improving the accuracy and real-time performance of abnormal event correlation analysis and achieving accurate abnormal event detection. The following is a further detailed description of this application.
[0058] like Figure 1 As shown in the figure, this application discloses an abnormal event detection method based on multi-source operation and maintenance data fusion, including steps such as data acquisition and preprocessing, feature map construction and updating, abnormal event root cause acquisition, and abnormal event display and query. The specific steps are as follows:
[0059] S1. Collect multi-source operation and maintenance data and perform preprocessing.
[0060] Specifically, data acquisition and preprocessing includes collecting and preprocessing multi-source operation and maintenance data. Preprocessing further includes noise reduction, data alignment, and data completion. During data collection, various data acquisition tools, such as sensors and log collectors, can be used to obtain operation and maintenance data from different data sources. For noise reduction, filtering algorithms, such as Kalman filtering and median filtering, can be used to remove noise interference from the data. Data alignment unifies data from different data sources in time or space to ensure data consistency. Data completion can be performed using interpolation algorithms, such as linear interpolation and spline interpolation, to supplement missing data.
[0061] S2. Based on the differential processing of multi-source operation and maintenance data driven by abnormal event types, obtain dynamically updated abnormal event maps.
[0062] Specifically, to improve the accuracy of abnormal event detection, an abnormal event graph is constructed. By performing graph structure analysis on the abnormal event graph, an abnormal event correlation analysis method based on multi-source operation and maintenance data fusion is realized. This application considers the monitoring differences of different abnormal types (such as user application abnormalities, HPC cluster abnormalities, network abnormalities, cooling abnormalities, power distribution abnormalities, reporting abnormalities, etc.) and designs a dynamic updating abnormal event graph optimization scheme that takes into account both comprehensiveness and adaptability, thereby ensuring the accuracy of subsequent abnormal event detection. The construction process of the dynamic updating abnormal event graph is described in detail below.
[0063] First, construct a basic abnormal event graph.
[0064] The design incorporates a hierarchical heterogeneous graph structure, comprising a base layer and a dynamic layer. The base layer is a static topology layer, specifically constructed based on the physical / logical architecture (such as network device connections, cooling pipelines, and power topology) to create an initial static graph. The dynamic layer is a multimodal relational topology layer, which extracts feature data to construct time-series related edges, event propagation edges, and semantically similar edges.
[0065] Secondly, the basic abnormal event map is dynamically updated.
[0066] Considering the differences in characteristic relationships between different anomaly event types, it is necessary to determine the anomaly event types mapped to the preprocessed multi-source operation and maintenance data based on the pre-built mapping relationship between multi-source operation and maintenance data and anomaly event types; for example, the current multi-source operation and maintenance data includes: ,Sure , These correspond to network anomaly event types and cooling anomaly event types, respectively. Furthermore, to determine the anomaly event types mapped from the preprocessed multi-source operation and maintenance data, a deep learning algorithm can be used to construct an anomaly event type identification model for the multi-source operation and maintenance data, thus completing the determination of the anomaly event types associated with the multi-source operation and maintenance data.
[0067] Complete the differentiated processing of multi-source operation and maintenance data driven by abnormal event types to obtain a dynamically updated abnormal event map. Specifically, the differentiated processing flow includes: Feature extraction: Feature extraction is performed based on the matching feature engine according to the abnormal event type; different abnormal event types are set with matching feature engines, and specific feature extraction is performed according to the matching feature engine, such as: for network abnormal event types, matching topology association and traffic dynamic feature extraction engine; for cooling / power distribution abnormal event types, matching physical quantity time series and spatial association feature extraction engine; for AI cluster abnormal event types, matching computing power resources and training process feature extraction engine; for user job abnormal event types, matching job behavior and resource interaction feature extraction engine; for reported abnormal event types, matching alarm semantics and behavioral pattern feature extraction engine.
[0068] Relationship modeling: Relationship modeling is performed based on the matching of operation and maintenance data relationships according to the anomaly event type; different anomaly event types are set with matching operation and maintenance data relationships, such as: time sequence relationship, event propagation relationship, semantic similarity relationship, and relationship construction is completed according to the corresponding operation and maintenance data relationship.
[0069] Graph Update: Dynamic updates to the graph are performed based on graph update strategies that match anomaly event types. This includes generating graph update instructions and completing dynamic updates of the graph based on these strategies. Different anomaly event types have corresponding graph update strategies (update frequencies) to update nodes and reconstruct edge relationships. For example, for network anomaly events, a topology-driven real-time diffusion update strategy is used; for cooling / power distribution anomaly events, a spatially correlated temporal cumulative update strategy is used; for AI cluster anomaly events, a computing power collaborative task link update strategy is used; for HPC cluster anomaly events, a job-driven node collaboration update strategy is used; and for user job anomaly events, a behavior tracing and verification update strategy is used.
[0070] S3. Based on the dynamically updated abnormal event map, graph neural network technology is used to perform cross-type feature aggregation on all nodes in the map, complete the differentiated causal reasoning driven by the abnormal event type, and obtain the final judgment of the abnormal event and the root cause location.
[0071] First, obtain the dynamically updated anomaly event map of the currently collected data. For all the dynamically updated anomaly event maps, use graph neural network technology (GNN) to perform cross-type feature aggregation on all nodes in the map, output the initial value of the anomaly probability of all nodes and the initial value of the probability distribution of each type of anomaly event, and make a preliminary judgment on whether there are anomalies, the type of anomaly event, and the anomaly level.
[0072] Specifically, GNN layer technology is used to perform cross-type feature aggregation on all nodes in the updated anomaly event graph. This includes: First, performing intra-type aggregation on each type of anomaly graph to strengthen the association features of nodes of the same type; Second, establishing inter-type associations through meta-path guidance and relational attention to achieve cross-type feature transfer, including: meta-path definition, pre-setting cross-type propagation paths based on operational scenario knowledge; relational attention mechanism, dynamically assigning weights to edges of different types; Third, inputting intra-type enhanced features and cross-type intermediate features into a globally fused GNN, outputting the initial anomaly probability value of each node (the comprehensive anomaly probability of each node is 0-1) and the initial probability distribution value of each node belonging to each type of anomaly event (the probability distribution of each node belonging to each type of anomaly event), thereby determining whether an anomaly event exists and its type and level. For example, a preset threshold for the comprehensive anomaly event probability is set; if it is greater than the preset threshold, an anomaly event exists, and the anomaly event type with the largest probability distribution is initially determined; preset thresholds for the anomaly event type probability are set for different anomaly event types to match the anomaly level, and the anomaly level of the initially determined anomaly event type is determined.
[0073] Furthermore, in the process of using graph neural network technology to aggregate cross-type features of all nodes in the graph, for dynamically updated abnormal event graphs, weights are assigned based on the centrality of the graph structure, with higher weights assigned to nodes that are closer to each type of abnormal event graph.
[0074] Secondly, we complete differentiated causal reasoning driven by abnormal event types to obtain the final judgment of abnormal events and root cause localization.
[0075] To further determine the specific root cause of the anomalous event, during the causal reasoning stage, different reasoning algorithms are selected to complete root cause localization based on the initially determined anomalous event type. This includes identifying the root cause node, causal path, and scope of impact. Specifically, a causal reasoning strategy is matched to the anomalous event type and anomalous level of each node in each dynamically updated graph. Following the matched causal reasoning strategy, root cause localization is completed, and the root cause localization of the anomalous event corresponding to the current anomalous event type is output. For the reasoning results of each node, the final determined anomalous event and its root cause localization are statistically obtained.
[0076] Specifically, different causal reasoning strategies are set up to match different types of abnormal events. For example, for network abnormal events, which are usually manifested as faults propagating along the network topology, the propagation path reasoning is matched to obtain the propagation path node sequence; for cooling abnormal events, which are manifested as abnormal changes in temperature gradients or disruption of periodic patterns, the periodic gradient reasoning is matched to trace the thermodynamic conduction path; for power distribution abnormal events, which must strictly adhere to circuit topology constraints (such as series / parallel relationships), the physical rule reasoning is matched to complete the causal discovery of ontological constraints; and for AI cluster abnormal events, which generally originate from job dependency chains, the DAG reasoning is matched to the DAG of the computing task to complete the tracking of fault job points.
[0077] S4. Perform statistical analysis on the finally determined abnormal events, and automatically display or support querying to obtain the statistical content of abnormal events within a preset time period.
[0078] Specifically, the acquired abnormal events, their types, and levels will be statistically analyzed to obtain the proportion of each type of abnormal event and the proportion of each type of abnormal event at each level. Figure 2 As shown, search methods are set up to allow users to intuitively query and obtain statistical content of abnormal events within a preset time period.
[0079] In a specific embodiment, to further improve the accuracy and real-time performance of abnormal event detection, especially for processing large-scale datasets, special processing is applied to high-frequency abnormal event types. Acceleration strategies improve processing efficiency and reduce computation time. Key feature extraction and key correlation modeling focus on important information, avoiding unnecessary computation. Pruning inference reduces the number of nodes involved in root cause calculation, improving inference speed. The design of the confidence evaluation function ensures high reliability of the output root cause localization. The method includes: determining the types of high-frequency abnormal events based on the statistical content of abnormal events within a nearby preset time period; wherein, preset frequencies are set for different abnormal event types, and the current abnormal event type is considered to be high-frequency if its frequency is greater than the corresponding preset frequency. For example, if the frequency of power distribution abnormal events within the preset 2-hour period closest to the current time is greater than the preset frequency corresponding to the power distribution abnormal event type, the current power distribution abnormal event is determined to be high-frequency.
[0080] In the process of differentiated processing of multi-source operation and maintenance data driven by high-frequency anomaly event types, a first acceleration strategy is initiated to obtain a sub-graph of the anomaly event graph that is dynamically updated at an accelerated pace. The sub-graph of the anomaly event graph refers to a sub-graph constructed based on historical data or real-time features, using the first acceleration strategy to retain only features highly correlated with the current anomaly event. For example, for network anomalies, a device feature sub-graph is extracted from the broadcast domain or Layer 3 path where the current anomaly node is located. The first acceleration strategy includes: extracting key features based on a feature engine matching high-frequency anomaly event types; modeling relationships based on the correlation between operation and maintenance data matching high-frequency anomaly event types; removing relationship edges with weights below a preset threshold; generating a graph update command; and completing the accelerated dynamic update of the sub-graph of the anomaly event graph based on the graph update strategy matching high-frequency anomaly event types.
[0081] Based on the sub-graph of the dynamically updated anomaly event graph, a second acceleration strategy is initiated during the process of completing differentiated causal inference driven by anomaly event type to obtain the final identified anomaly event and root cause localization. The second acceleration strategy employs a pruning inference approach to accelerate causal inference, including retaining only the nodes with the top-K correlation with the anomaly node for root cause calculation. For example, in the causal inference acceleration process for network anomaly types, a lightweight contribution analysis based on path hop count can be used to replace the random walk algorithm; or, for example, in the causal inference acceleration process for cooling anomaly types, periodic extreme point gradient detection can be used to replace complex gradient backpropagation.
[0082] To balance the accuracy of key feature extraction, accelerated inference, and root cause localization, a confidence evaluation function is designed based on the root cause localization of abnormal event types. This function evaluates the confidence of root cause localization and outputs root cause localizations with confidence scores higher than the preset confidence scores. The confidence evaluation function can be designed with a multi-dimensional evaluation index system. The evaluation index parameters include historical pattern matching degree (X1) and subgraph coverage integrity (X2), which are weighted to obtain the final confidence score. The historical pattern matching degree calculation includes: obtaining historical root cause localization results for different abnormal event types and calculating the similarity between the historical root cause localization results and the current root cause localization. The subgraph coverage integrity calculation includes: obtaining the key node set and calculating the coverage rate of key nodes among the nodes included in the current root cause localization. Furthermore, the weights are set independently during the confidence evaluation of root cause localization for different abnormal event types.
[0083] Furthermore, to balance the accuracy of key feature extraction, accelerated inference, and root cause localization, a circuit breaker protection mechanism is further implemented to prevent poor accuracy in root cause localization; the method includes:
[0084] In the confidence assessment of statistical root cause localization, the proportion of root cause localizations with output confidence scores higher than the preset confidence score is less than the preset proportion. If the statistical proportion is less than the preset proportion, it indicates that the accuracy of the current root cause localization is poor, and a circuit breaker protection instruction is generated. This instruction is used to disable the first acceleration strategy during the differentiated processing of multi-source operation and maintenance data driven by abnormal event types that are in high-frequency anomalies within the preset circuit breaker protection period.
[0085] In one specific embodiment, the method involves receiving root cause feedback information from users to correct and optimize the abnormal event map, thereby improving the map's accuracy. For the identification and processing of high-frequency false alarm nodes, the method adjusts the abnormal probability threshold or causal inference threshold for abnormal event nodes to reduce false alarms and further improve the accuracy of abnormal event detection. The method also includes:
[0086] Design a structured root cause feedback template and receive structured root cause feedback information for different types of abnormal events uploaded by users. In order to enhance the causal relationship of user feedback intervention, a root cause feedback template interface can be set up so that users can use the root cause feedback template on the interface to output feedback information for the root cause location of abnormal events. In order to facilitate intelligent analysis, a structured root cause feedback template is designed.
[0087] From the perspective of graph enhancement, we analyze root cause feedback information to obtain causal paths in order to correct the abnormal event graph. This includes: enhancing the edge weights between nodes in the abnormal event graph that conform to the causal path, adding new causal paths between nodes in the abnormal event graph, and reducing the edge weights between nodes in the abnormal event graph that contradict the causal path.
[0088] From the perspective of anomaly probability thresholds and causal inference thresholds, this study analyzes root cause feedback information to obtain causal paths in order to adjust the anomaly probability thresholds or causal inference thresholds for anomalous event nodes. Specifically, this includes:
[0089] If each node in the root cause localization of an abnormal event is not included in the nodes of the causal path obtained based on the root cause feedback information, and the frequency of occurrence of the node not included in the nodes of the causal path obtained based on the root cause feedback information is greater than the preset frequency within a preset time, then the node is identified as the first high-frequency false alarm node in the root cause localization of the abnormal event, and the node abnormality probability threshold of the first high-frequency false alarm node is reduced.
[0090] After determining that each node in the root cause localization of anomalies is included in the nodes of the causal path obtained based on root cause feedback information, the similarity between each node in the root cause localization of anomalies and each node in the causal path obtained based on root cause feedback information is compared. If the similarity between the node and each node in the causal path obtained based on root cause feedback information is less than a preset similarity, and the frequency of occurrence of the similarity between the node and each node in the causal path obtained based on root cause feedback information being less than the preset similarity is greater than a preset frequency within a preset time, the node is identified as the second most frequent false alarm node in the root cause localization of anomalies, and the node causal inference threshold of the second most frequent false alarm node is reduced.
[0091] In a specific embodiment, by statistically analyzing the proportion of abnormal events, resource allocation is dynamically adjusted according to different triggering conditions to rationally utilize resources and improve processing efficiency. Allocating resources based on the priority of abnormal events ensures timely handling of important abnormal events, improving system response speed and reliability. The method further includes: statistically analyzing the finally determined abnormal events to obtain the proportion of different abnormal events within a preset time period; for example, the proportions of network abnormalities, cooling abnormalities, and power distribution abnormalities are 60%:30%:10%. Triggering conditions are set based on the obtained proportions of different abnormal events, and resource allocation is dynamically adjusted during resource allocation based on the differentiated processing of multi-source operation and maintenance data driven by abnormal event types and the differentiated causal inference based on abnormal event types. Specifically, the priority of abnormal event types is dynamically adjusted, and differentiated processing resources and causal inference resources are allocated according to priority. The triggering conditions include: periodic triggering, timely triggering, and emergency triggering.
[0092] Specifically, the periodic triggering involves setting a preset adjustment period, calculating the proportional changes of different types of abnormal events within the preset adjustment period, matching the calculated proportional changes with preset abnormal event priority adjustment rules, triggering the adjustment of the priorities of the corresponding different types of abnormal events, and determining the resource allocation weights of different abnormal events according to the adjusted priorities of the different types of abnormal events to achieve the adjustment of resource allocation weights; the preset abnormal event priority adjustment rules include increasing the priority level of the corresponding type of abnormal event when the calculated proportional change is greater than a preset proportional threshold range; for example, if the proportional change is greater than 20% within the preset adjustment period, the priority level of the current corresponding type of abnormal event will be increased accordingly.
[0093] The timely trigger is to detect abnormal event types where the proportion of abnormal events changes beyond the proportion change threshold. This indicates that the abnormal event type occurs frequently in the current period. The corresponding abnormal event type is directly triggered to increase its priority. The resource allocation weight of different abnormal events is determined according to the adjusted priority of different types of abnormal events to achieve the adjustment of resource allocation weight; for example, the proportion change exceeds 50% within the preset monitoring period.
[0094] The emergency trigger is to receive instructions on adjusting the priority of different types of abnormal events, and then adjust the priority of the corresponding different types of abnormal events according to the instructions. Based on the adjusted priorities of the different types of abnormal events, the resource allocation weight of the different abnormal events is determined to realize the adjustment of the resource allocation weight. The resource allocation trigger board can be set to receive user instructions on adjusting the priority of different types of abnormal events.
[0095] How to determine the resource allocation weights for different types of abnormal events based on the adjusted priorities described above, specifically including:
[0096] Based on the priority of different types of abnormal events, resource allocation weights are preset and matched with different priorities. For example, the resource allocation weights for priority levels are set from high to low as 0.5, 0.3, 0.2, and 0.1. For the N types of abnormal events, the resource allocation weights matched with the priority of each type of abnormal event are determined, and then normalization is performed to obtain the final weights, such as 0.703, 0.248, and 0.049. In order to avoid the weights being too low, a minimum weight for resource allocation of a single type of abnormal event can also be set. For example, the weight value below the minimum weight is adjusted to the minimum weight of 0.1, and then re-normalized.
[0097] Total resources are allocated according to the resource allocation weight of different abnormal events; for example, if the total allocation is 100 CPU cores, 67 are obtained for network abnormalities, 24 for cooling abnormalities, and 9 for power distribution abnormalities.
[0098] Continue to adjust and execute the resource allocation for graph update, feature aggregation, and causal inference for a single type of abnormal event according to the pre-defined resource allocation ratios for graph update, feature aggregation, and causal inference. For example, if the pre-defined resource allocation ratios for graph update, feature aggregation, and causal inference are 3:4:3, then 67*0.3 graph update resources, 67*0.4 feature aggregation resources, and 67*0.3 causal inference resources are allocated, which are approximately 20, 27, and 20 CPU cores respectively, and the corresponding resource allocation adjustment is completed.
[0099] like Figure 3As shown in the figure, this application discloses an abnormal event detection system based on multi-source operation and maintenance data fusion, including:
[0100] The multi-source operation and maintenance data acquisition and processing module 101 is used to acquire multi-source operation and maintenance data and perform preprocessing; the preprocessing includes noise reduction, data alignment and completion.
[0101] The feature map construction and update module 102 is used to determine the abnormal event types mapped by the pre-built multi-source operation and maintenance data and abnormal event type mapping relationship based on the pre-built multi-source operation and maintenance data and abnormal event type mapping relationship, complete the differential processing of multi-source operation and maintenance data driven by abnormal event type, and obtain a dynamically updated abnormal event map; the differential processing process includes: feature extraction based on abnormal event type matching feature engine, relationship modeling based on abnormal event type matching operation and maintenance data association relationship, generating map update instructions and completing the dynamic update of abnormal event map based on the map update strategy of abnormal event type matching;
[0102] The abnormal event root cause acquisition module 103 is used to perform cross-type feature aggregation on all nodes in the graph based on the dynamically updated abnormal event map, and output the initial value of the abnormal probability of all nodes and the initial value of the probability distribution of each type of abnormal event. It preliminarily judges whether there is an abnormal event, the type of abnormal event and the level of abnormality, completes differentiated causal reasoning driven by the abnormal event type, and obtains the final judgment of the abnormal event and the root cause location.
[0103] The abnormal event display and query module 104 is used to perform statistical analysis on the finally determined abnormal events, automatically display or support querying and obtaining the statistical content of abnormal events within a preset time period, including: the proportion of various types of abnormal events and the proportion of abnormal level of each type of abnormal event.
[0104] Through the collaborative work of various modules, the collection, processing, analysis, and application display of multi-source operation and maintenance data were achieved; for example... Figure 4 As shown, the multi-source operation and maintenance data acquisition and processing module, serving as the data acquisition and processing layer, provides the data foundation for subsequent analysis. The feature graph construction and update module, serving as the analysis engine layer, constructs a correlation graph of abnormal events. The abnormal event root cause acquisition module, serving as the application service layer, realizes the location and root cause analysis of abnormal events. The abnormal event display and query module presents the analysis results to the user, facilitating decision-making. This system effectively improves the accuracy and real-time performance of abnormal event correlation analysis, meeting the needs of complex operation and maintenance environments.
[0105] In one specific embodiment, the system further includes:
[0106] The abnormal event result statistics module 105 is used to determine the type of abnormal event that is in the high-frequency abnormality based on the statistical content of abnormal events within a nearby preset time period.
[0107] The feature graph construction, update, and optimization module 106 is also used to initiate a first acceleration strategy during the differentiated processing of multi-source operation and maintenance data driven by abnormal event types in high-frequency anomalies, and to obtain a sub-graph of the abnormal event graph that is dynamically updated in an accelerated manner. The first acceleration strategy includes: extracting key features based on the feature engine matching abnormal event types in high-frequency anomalies, performing relationship modeling based on the relationship between operation and maintenance data matching abnormal event types in high-frequency anomalies, removing relationship edges with weights lower than a preset threshold, generating a graph update instruction, and completing the accelerated dynamic update of the sub-graph of the abnormal event graph based on the graph update strategy matching abnormal event types in high-frequency anomalies.
[0108] The abnormal event root cause optimization acquisition module 107 is also used to activate a second acceleration strategy during the process of completing differentiated causal reasoning driven by abnormal event type based on the sub-map of the accelerated dynamically updated abnormal event map, to obtain the finally judged abnormal event and root cause location; to design a confidence evaluation function based on the root cause location of abnormal event type, to complete the confidence evaluation of root cause location, and to output the root cause location with a confidence level higher than the preset confidence level; the second acceleration strategy adopts a pruning reasoning method to accelerate causal reasoning, including only retaining the nodes with the top-K correlation with the abnormal node to participate in the root cause calculation.
[0109] In one specific embodiment, the system further includes:
[0110] The root cause feedback module 108 is used to design structured root cause feedback templates and receive structured root cause feedback information of different types of abnormal events uploaded by users.
[0111] The feature graph construction and correction module 109 is used to analyze root cause feedback information to obtain causal paths in order to correct the abnormal event graph. It includes: enhancing the edge weights between nodes in the abnormal event graph that conform to the causal path, adding new causal paths between nodes in the abnormal event graph, and reducing the edge weights between nodes in the abnormal event graph that contradict the causal path.
[0112] The abnormal event root cause correction module 110 is used to analyze root cause feedback information to obtain causal paths in order to adjust the abnormal probability threshold or causal inference threshold of abnormal event nodes. This includes: determining whether each node in the abnormal event root cause localization is included in all nodes obtained based on the causal path based on root cause feedback information; identifying nodes in the abnormal event root cause localization that are not all included in all nodes obtained based on the causal path based on root cause feedback information, and whose occurrence frequency within a preset time is greater than a preset frequency, as the first high-frequency false alarm node in the abnormal event root cause localization; and lowering the node abnormal probability threshold of the first high-frequency false alarm node. After determining that each node in the root cause localization of an abnormal event is included in the nodes of the causal path obtained based on root cause feedback information, the similarity between each node in the root cause localization of an abnormal event and each node in the causal path obtained based on root cause feedback information is compared. If the similarity between the node and each node in the causal path obtained based on root cause feedback information is less than a preset similarity, and the frequency of occurrence of the similarity between the node and each node in the causal path obtained based on root cause feedback information being less than the preset similarity is greater than a preset frequency within a preset time, the node is identified as the second most frequent false alarm node in the root cause localization of an abnormal event, and the node causal inference threshold of the second most frequent false alarm node is reduced.
[0113] In one specific embodiment, the abnormal event result statistics module 105 is further used to determine the type of abnormal event that is in a high frequency of abnormality based on the statistical content of abnormal events within a nearby preset time period.
[0114] The resource allocation module 111 is used to perform statistical analysis on the finally judged abnormal events, obtain the proportion of different abnormal events within a nearby preset time period, set trigger conditions based on the obtained proportion of different abnormal events, and dynamically adjust the resource allocation based on the differential processing of multi-source operation and maintenance data driven by the abnormal event type and the resource allocation based on the differential causal reasoning driven by the abnormal event type when resource allocation is triggered.
[0115] This application also discloses a computer-readable storage medium.
[0116] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed, such as the above-described abnormal event detection method based on multi-source operation and maintenance data fusion. The computer-readable storage medium includes, for example, various media that can store program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0117] This application also discloses a computer device.
[0118] Specifically, the computer device includes a memory and a processor. The memory stores a computer program that can be loaded and executed by the processor to perform the above-mentioned abnormal event detection method based on multi-source operation and maintenance data fusion.
[0119] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. An abnormal event detection method based on multi-source operation and maintenance data fusion, characterized in that, include: Collect multi-source operation and maintenance data and perform preprocessing; the preprocessing includes noise reduction, data alignment and completion. Based on the pre-built mapping relationship between multi-source operation and maintenance data and abnormal event types, the abnormal event types mapped by the pre-processed multi-source operation and maintenance data are determined, and the differential processing of multi-source operation and maintenance data driven by abnormal event types is completed to obtain a dynamically updated abnormal event map. The differentiated processing flow includes: feature extraction based on an anomaly event type matching feature engine; relationship modeling based on anomaly event type matching operation and maintenance data association; generating graph update instructions and completing dynamic updates of the anomaly event graph based on the graph update strategy matched by the anomaly event type; wherein, different anomaly event types are set with matching feature engines, and specific feature extraction is completed according to the matching feature engines; different anomaly event types are set with matching operation and maintenance data association, and relationship construction is completed according to the matching operation and maintenance data association; different anomaly event types are set with matching graph update strategies, thereby completing node updates and edge relationship reconstruction; Based on a dynamically updated abnormal event graph, graph neural network technology is used to aggregate cross-type features of all nodes in the graph, outputting the initial values of the abnormal probability of all nodes and the initial values of the probability distribution of each type of abnormal event. This allows for a preliminary determination of whether an abnormal event exists, its type, and its level. Differential causal reasoning driven by the abnormal event type is then completed to obtain the final judgment of the abnormal event and the root cause location. Statistical analysis is performed on the finally identified abnormal events, and the statistical content of abnormal events within the recent preset time period is automatically displayed or can be queried and retrieved, including: the proportion of various types of abnormal events and the proportion of abnormality level of each type of abnormal event.
2. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 1, characterized in that, Also includes: Based on the statistical data of abnormal events within a preset time period, determine the types of abnormal events that are high-frequency abnormalities; In the process of differentiated processing of multi-source operation and maintenance data driven by high-frequency abnormal event types, the first acceleration strategy is launched to obtain the sub-map of the abnormal event map that is dynamically updated at an accelerated pace. The first acceleration strategy includes: extracting key features based on the feature engine matching the abnormal event type in high frequency anomalies, performing relationship modeling based on the relationship between the operation and maintenance data matching the abnormal event type in high frequency anomalies, removing relationship edges with weights lower than a preset threshold, generating a graph update instruction, and completing the subgraph accelerated dynamic update of the abnormal event graph based on the graph update strategy matching the abnormal event type in high frequency anomalies. Based on the sub-graph of the dynamically updated anomaly event graph, a second acceleration strategy is initiated during the process of completing differentiated causal reasoning driven by anomaly event type to obtain the final identified anomaly event and root cause localization. A confidence evaluation function is designed based on the root cause localization of anomaly event type to complete the confidence evaluation of root cause localization and output the root cause localization with a confidence level higher than the preset confidence level. The second acceleration strategy adopts a pruning reasoning method to accelerate causal reasoning, including retaining only the nodes with the top-K correlation with the anomaly node to participate in the root cause calculation.
3. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 1, characterized in that, Also includes: Design a structured root cause feedback template and receive structured root cause feedback information for different types of abnormal events uploaded by users; Analyze root cause feedback information to obtain causal paths in order to correct the abnormal event graph, including: enhancing the edge weights between nodes in the abnormal event graph that conform to the obtained causal path, adding new causal paths between nodes in the abnormal event graph, and reducing the edge weights between nodes in the abnormal event graph that contradict the obtained causal path. Analyzing root cause feedback information to obtain causal paths in order to adjust the abnormal probability threshold or causal inference threshold of abnormal event nodes includes: determining whether each node in the abnormal event root cause localization is included in each node in the causal path obtained based on root cause feedback information; determining that nodes in the abnormal event root cause localization are not included in each node in the causal path obtained based on root cause feedback information, and whose occurrence frequency is greater than the preset frequency within a preset time, identifying such nodes as the first high-frequency false alarm nodes in the abnormal event root cause localization, and reducing the node abnormal probability threshold of the first high-frequency false alarm nodes. After determining that each node in the root cause localization of anomalies is included in the nodes of the causal path obtained based on root cause feedback information, the similarity between each node in the root cause localization of anomalies and each node in the causal path obtained based on root cause feedback information is compared. If the similarity between the node and each node in the causal path obtained based on root cause feedback information is less than a preset similarity, and the frequency of occurrence of the similarity between the node and each node in the causal path obtained based on root cause feedback information being less than the preset similarity is greater than a preset frequency within a preset time, the node is identified as the second most frequent false alarm node in the root cause localization of anomalies, and the node causal inference threshold of the second most frequent false alarm node is reduced.
4. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 2, characterized in that, Also includes: In the confidence assessment of statistical root cause localization, the proportion of root cause localizations with output confidence scores higher than the preset confidence score is used to determine the percentage of all output confidence scores. When the determined statistical proportion is less than the preset proportion, a circuit breaker protection instruction is generated. This instruction is used to disable the first acceleration strategy during the differentiated processing of multi-source operation and maintenance data driven by abnormal event types that are in high-frequency anomalies within the preset circuit breaker protection period.
5. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 1, characterized in that, Also includes: Statistical analysis is performed on the finally identified abnormal events to obtain the proportion of different abnormal events within a preset time period; Triggering conditions are set based on the proportion of different abnormal events acquired, and resource allocation based on the differential processing of multi-source operation and maintenance data driven by abnormal event type and the differential causal reasoning driven by abnormal event type are dynamically adjusted when triggering resource allocation; wherein, the triggering conditions include: periodic triggering, timely triggering and emergency triggering; The periodic triggering involves setting a preset adjustment period, calculating the proportional changes of different types of abnormal events within the preset adjustment period, matching the calculated proportional changes with preset abnormal event priority adjustment rules, triggering the adjustment of the priorities of the corresponding different types of abnormal events, and determining the resource allocation weights of different abnormal events according to the adjusted priorities of the different types of abnormal events to achieve the adjustment of resource allocation weights; the preset abnormal event priority adjustment rules include increasing the priority level of the corresponding type of abnormal event when the calculated proportional change is greater than a preset proportional threshold range. The timely triggering refers to monitoring abnormal event types where the proportion of abnormal events exceeds the proportion of abnormal event threshold, directly triggering an increase in the priority of the corresponding type of abnormal event, and determining the resource allocation weight of different abnormal events according to the adjusted priority of different types of abnormal events to achieve the adjustment of resource allocation weight; The emergency triggering process involves receiving instructions to adjust the priorities of different types of abnormal events, adjusting the priorities of the corresponding different types of abnormal events according to the instructions, and determining the resource allocation weights of different abnormal events based on the adjusted priorities of the different types of abnormal events to achieve the adjustment of resource allocation weights.
6. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 5, characterized in that, The determination of resource allocation weights for different types of abnormal events based on their adjusted priorities includes: Based on the priority of different types of abnormal events, resource allocation weights are preset and matched with different priorities; Allocate total resources according to the resource allocation weights of different anomalies; continue to adjust and execute the resource allocation of graph update resources, feature aggregation resources, and causal inference resources for a single type of anomaly according to the pre-divided graph update resource ratio, feature aggregation resource ratio, and causal inference resource ratio.
7. The abnormal event detection method based on multi-source operation and maintenance data fusion according to claim 1, characterized in that, The method of using graph neural network technology to perform cross-type feature aggregation on all nodes in the graph includes: for dynamically updated abnormal event graphs, weight allocation based on the centrality of the graph structure.
8. An abnormal event detection system based on multi-source operation and maintenance data fusion, characterized in that, include: The multi-source operation and maintenance data acquisition and processing module is used to acquire multi-source operation and maintenance data and perform preprocessing; the preprocessing includes noise reduction, data alignment and completion. The feature map construction and update module is used to determine the abnormal event types mapped by the pre-built mapping relationship between multi-source operation and maintenance data and abnormal event types, complete the differential processing of multi-source operation and maintenance data driven by abnormal event types, and obtain dynamically updated abnormal event maps. The differentiated processing flow includes: feature extraction based on an anomaly event type matching feature engine; relationship modeling based on anomaly event type matching operation and maintenance data association; generating graph update instructions and completing dynamic updates of the anomaly event graph based on the graph update strategy matched by the anomaly event type; wherein, different anomaly event types are set with matching feature engines, and specific feature extraction is completed according to the matching feature engines; different anomaly event types are set with matching operation and maintenance data association, and relationship construction is completed according to the matching operation and maintenance data association; different anomaly event types are set with matching graph update strategies, thereby completing node updates and edge relationship reconstruction; The abnormal event root cause acquisition module is used to perform cross-type feature aggregation on all nodes in the graph based on the dynamically updated abnormal event map. It uses graph neural network technology to output the initial value of the abnormal probability of all nodes and the initial value of the probability distribution of each type of abnormal event. It makes a preliminary judgment on whether there is an abnormal event, the type of abnormal event and the level of abnormality. It completes differentiated causal reasoning based on the abnormal event type and obtains the final judgment of the abnormal event and the root cause location. The abnormal event display and query module is used to perform statistical analysis on the finally determined abnormal events, automatically display or support querying to obtain the statistical content of abnormal events within a preset time period, including: the proportion of various types of abnormal events and the proportion of abnormal level of each type of abnormal event.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the method as described in any one of claims 1 to 7.
10. A computer device, characterized in that, The computer device includes a memory, a processor, and a program stored in and executable on the memory, the program being executed by the processor to implement the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge graph construction method and system for identifying new energy abnormal data
CN116340534A
Quick traceable multi-dimensional abnormal event root cause analysis algorithm
CN117827512A