Intelligent task alarm rule self-learning method and system based on supporting priority

Through the intelligent task alarm rules self-learning method, a two-way feature extraction network and an adversarial learning framework are used to generate an adaptive priority scoring system and a multi-level alarm strategy tree, which solves the problem that traditional alarm systems are difficult to adapt to dynamic changes and improves the accuracy and efficiency of alarms.

CN119441832BActive Publication Date: 2025-05-16北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510026462.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Traditional task alarm systems are difficult to adapt to dynamically changing environments, resulting in problems such as alarm storms, missed false alarms, and lack of full understanding of task context information, resulting in low alarm accuracy.

Method used

Using the intelligent task alarm rules self-learning method based on supporting priority, a two-way feature extraction network and an adversarial learning framework are built, and an adaptive priority scoring system and a multi-level alarm strategy tree are generated to achieve dynamic optimization of alarm rules.

Benefits of technology

It improves the accuracy and efficiency of alarms, reduces false alarms and missed reports, realizes intelligent alarms and responses, and enhances the generalization ability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441832B_ABST
    Figure CN119441832B_ABST
Patent Text Reader

Abstract

The present invention provides a self-learning method and system for intelligent task alarm rules based on priority support, which relates to the field of intelligent technology, including building a two-way feature extraction network to collect task operation data, and generating a task abnormality association model in combination with historical alarm data. The model is migrated and mapped with a business scenario knowledge base using a transfer learning algorithm to generate task priority quantitative indicators, and an adaptive priority scoring system is constructed. Based on the model and the scoring system, an alarm rule generator is trained to generate a multi-level alarm strategy tree to achieve differentiated alarms. The incremental learning engine monitors the task status in real time, dynamically adjusts the alarm response strategy and feeds back data to optimize the rules. The federated learning mechanism is used to collaboratively train the model to improve the generalization ability and adaptability of the alarm rules, thereby achieving more accurate and efficient task alarms, reducing the false alarm rate, and improving operation and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to intelligent technology, and in particular to an intelligent task alarm rule self-learning method and system based on supporting priorities. Background Art

[0002] Task alarms play a vital role in ensuring the stable operation of complex systems. As the scale of the system expands and the complexity of the business increases, traditional static alarm rules are difficult to adapt to the dynamically changing environment, resulting in frequent problems such as alarm storms, missed reports and false reports, which seriously affect the efficiency of operation and maintenance.

[0003] Lack of full understanding of task context information: Traditional alarm rules are usually based on simple threshold settings, which cannot effectively capture the complexity and relevance of task running status, resulting in low alarm accuracy. For example, a short-term increase in the resource usage of a task may be a normal fluctuation, but a simple threshold alarm mechanism will misjudge it as an abnormality.

[0004] The alarm priority distinction is not precise enough: Traditional alarm systems often use the same processing strategy for all alarms, lacking the distinction between the importance of different tasks and alarm levels, resulting in critical task alarms being submerged in a large number of low-priority alarms, delaying fault handling.

[0005] Alarm rules are difficult to adapt to dynamic changes: The system operating environment and business load will change continuously. Static alarm rules are difficult to adapt to this dynamic nature and require frequent manual intervention and adjustment, which increases operation and maintenance costs. For example, the resource utilization rate during business peak and trough periods varies greatly, and static thresholds are difficult to meet the needs of different periods at the same time. Summary of the invention

[0006] The embodiments of the present invention provide a method and system for self-learning intelligent task alarm rules based on supporting priorities, which can solve the problems in the prior art.

[0007] According to a first aspect of the embodiments of the present invention,

[0008] Provides an intelligent task alarm rule self-learning method based on support priority, including:

[0009] A bidirectional feature extraction network is constructed to collect task operation data. The bidirectional feature extraction network includes a forward feature extraction layer and a reverse verification layer. The forward feature extraction layer obtains the task operation time, resource occupancy rate, execution status and business importance to construct a real-time feature matrix. The reverse verification layer analyzes the historical alarm data based on the graph neural network to obtain an alarm impact link diagram. The real-time feature matrix and the alarm impact link diagram are feature-fused to obtain a task anomaly association model. The task anomaly association model is migrated and mapped to a business scenario knowledge base using a transfer learning algorithm to generate a task priority quantitative index that considers the business scenario, and an adaptive priority scoring system is constructed based on the task priority quantitative index.

[0010] An alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link diagram, and continuously optimizes the accuracy of the alarm rule through adversarial training; the verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds, and each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and response mechanism of different priorities;

[0011] An incremental learning engine is started to monitor the task running status in real time. When it is detected that the task running indicators trigger the alarm conditions in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated during the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to achieve dynamic optimization of the alarm rules; based on the federated learning mechanism, the task anomaly association model is collaboratively trained among multiple business nodes to improve the generalization ability and adaptability of the alarm rules while protecting data privacy.

[0012] The forward feature extraction layer obtains the task's running time, resource occupancy rate, execution status and business importance to build a real-time feature matrix. The reverse verification layer analyzes historical alarm data based on a graph neural network to obtain an alarm impact link graph. The real-time feature matrix and the alarm impact link graph are subjected to feature fusion to obtain a task abnormality association model, including:

[0013] The running time, the resource occupancy rate, the execution status and the business importance are constructed into a four-dimensional feature vector; the forward feature extraction layer performs attention-weighted fusion on the four-dimensional feature vector to generate a real-time feature matrix representing the task running status;

[0014] The reverse verification layer uses a graph neural network to construct an alarm impact link graph, constructs the historical alarm data into graph nodes with alarm attributes, constructs graph edge connections based on the association relationship between alarms, extracts deep features of the graph nodes through graph convolution operations, and fuses the topological information of the graph edge connections to generate an alarm impact link graph that characterizes the alarm propagation mode; converts the node features in the alarm impact link graph into a feature space with the same dimension as the real-time feature matrix through feature projection;

[0015] A cross-modal feature fusion network is constructed. The cross-modal feature fusion network calculates the similarity between the real-time feature matrix and the node features of the alarm impact link graph to obtain the attention weight. Based on the attention weight, the real-time feature matrix and the node features of the alarm impact link graph are adaptively fused to generate a task anomaly association model.

[0016] Using a transfer learning algorithm to perform migration mapping between the task anomaly association model and the business scenario knowledge base, generating a task priority quantitative index considering the business scenario, and constructing an adaptive priority scoring system based on the task priority quantitative index includes:

[0017] Construct a business scenario knowledge base, which includes a business process topology, system component dependencies, and operation and maintenance processing strategies, and use knowledge graph technology to construct the information in the business scenario knowledge base into knowledge entities and relationship edges, and map the knowledge entities and the relationship edges into knowledge feature vectors through a graph embedding algorithm;

[0018] A feature migration mapping network is constructed by using a migration learning algorithm. The feature migration mapping network calculates the semantic similarity between the feature representation of the task anomaly association model and the knowledge feature vector to generate an attention weight matrix. Based on the attention weight matrix, the knowledge feature vector is migrated and mapped to the feature space of the task anomaly association model. Residual connections are used to maintain the original feature information to generate a task feature representation that integrates business scenario knowledge.

[0019] Based on the task feature representation of the integrated business scenario knowledge, a task priority quantitative index is constructed, the task priority quantitative index includes a task urgency index, a business impact index and a resource consumption index, a hierarchical analysis method is used to calculate the weight coefficient of the task priority quantitative index, and the task priority quantitative index and the weight coefficient are weightedly combined to obtain a priority score;

[0020] An adaptive priority scoring system is constructed based on the priority scoring. The adaptive priority scoring system takes the task feature representation integrating the business scenario knowledge as the input state, takes the priority score as the output action, designs a reward function based on historical processing effects, and uses a deep reinforcement learning algorithm to continuously optimize the scoring strategy.

[0021] The alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link graph. The accuracy of the alarm rule is continuously optimized through adversarial training, including:

[0022] Constructing an initial training set for a rule generator, taking the abnormal feature matrix in the task abnormality association model as the basic data of the rule triggering condition, taking the scoring result of the adaptive priority scoring system as the basis for dividing the rule alarm level, taking the historical alarm processing record as the reference data for the rule processing action, and vectorizing and encoding the basic data, the division basis and the reference data to obtain a rule training vector;

[0023] Constructing a generative network in an adversarial learning framework, wherein the generative network adopts a multi-layer neural network structure, inputs the rule training vector to generate an initial alarm rule candidate set, wherein each rule in the initial alarm rule candidate set includes a trigger condition threshold, an alarm level, and a processing action sequence;

[0024] Constructing a discriminant network in an adversarial learning framework, wherein the discriminant network constructs a rule evaluation model based on the alarm impact link diagram, wherein the rule evaluation model calculates the similarity between the trigger condition threshold of the rule in the initial alarm rule candidate set and the historical abnormal characteristics to obtain a rule accuracy score, calculates the consistency between the alarm level of the rule and the result of the adaptive priority scoring system to obtain a rule rationality score, and verifies the executableness of the processing action sequence of the rule in the current system environment to obtain a rule effectiveness score;

[0025] Based on the rule accuracy score, the rule rationality score and the rule effectiveness score, the discriminant network screens the initial alarm rule candidate set, and inputs the scoring result as a feedback signal into the generation network to guide the generation network to adjust the rule generation strategy, and obtains the accuracy of the final alarm rule through multiple rounds of adversarial training iterative optimization.

[0026] The verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds. Each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and a response mechanism of different priorities, including:

[0027] Establishing a mapping relationship between the trigger conditions in the verified alarm rule set and the task priority quantitative index, stratifying the alarm rule set according to the adaptive priority score, and constructing a basic framework of a multi-level alarm strategy tree, wherein each node of the multi-level alarm strategy tree includes a priority interval, an alarm trigger condition, and a response processing mechanism, and organizing alarm rules with similar adaptive priority scores into the same node;

[0028] For each node of the multi-level alarm strategy tree, based on the historical trigger data of the alarm rules in the node, a density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node;

[0029] A response processing mechanism is configured for each node of the multi-level alarm strategy tree, independent processing resources are allocated to priority nodes above the alarm threshold and a quick response channel is set up, a resource pool dynamic allocation mechanism is adopted for priority nodes aligned with the alarm threshold, and a batch processing mode is used for priority nodes below the alarm threshold, thereby building an alarm upgrade path between nodes.

[0030] Based on the historical trigger data of the alarm rules in the node, the density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node, including:

[0031] Extract historical trigger data of alarm rules from the node, construct the indicator value, trigger frequency, duration and impact range in the historical trigger data into an initial feature vector, and perform data cleaning and standardization on the initial feature vector to obtain a standardized feature vector;

[0032] A density clustering method is used to perform cluster analysis on the standardized feature vector, wherein the density clustering method dynamically calculates the density accessibility of sample points in the feature space through an adaptive neighborhood radius mechanism to obtain a cluster set of alarm feature distribution;

[0033] For each cluster in the cluster set, the center point position, density distribution and boundary range are calculated to obtain a cluster feature set, and an alarm threshold calculation model is established based on the cluster feature set, wherein the alarm threshold calculation model includes a cluster feature fusion unit and a threshold dynamic adjustment unit;

[0034] Obtaining priority interval information of the node, the alarm threshold calculation model assigns weight coefficients to different clusters in the cluster feature set according to the priority interval information, and generates initial alarm trigger conditions with differentiated sensitivity by weighted fusion of various cluster features;

[0035] Collect alarm trigger data within the sliding time window, calculate the false alarm rate, missed alarm rate and timeliness index of the initial alarm trigger condition to obtain a monitoring index set, and when the index value in the monitoring index set deviates from the expected threshold, the expected threshold dynamic adjustment unit calculates the adjustment parameter according to the deviation degree and the node priority;

[0036] The initial alarm trigger condition is adaptively adjusted based on the adjustment parameter, and the historical data of the node business load is analyzed to identify the load cycle characteristics, a mapping relationship model between the initial alarm trigger condition and the load state is established, and the sensitivity of the initial alarm trigger condition is optimized according to the current load level.

[0037] The incremental learning engine is started to monitor the task running status in real time. When it is detected that the task running indicator triggers the alarm condition in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated in the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to realize the dynamic optimization of the alarm rules, including:

[0038] Starting an incremental learning engine to monitor task running status data in real time, the incremental learning engine obtains task running indicators through a multidimensional time series analysis method, and uses a sliding window mechanism to preprocess the task running indicators to obtain an indicator feature sequence;

[0039] Inputting the indicator feature sequence into the multi-level alarm strategy tree for alarm condition matching, and when the indicator feature sequence triggers the alarm condition in the multi-level alarm strategy tree, constructing an alarm feature vector of the current alarm scenario based on the indicator feature sequence;

[0040] Performing scene recognition in a pre-established scene feature library according to the alarm feature vector, and selecting a corresponding response strategy template from a response strategy library based on the scene recognition result, wherein the response strategy template includes a processing step sequence, a resource configuration scheme, and a time limit parameter;

[0041] The incremental learning engine dynamically adjusts the selected response strategy template according to the current system load status and resource utilization, generates an alarm response strategy adapted to the current scenario, and sends the alarm response strategy to the alarm processing module;

[0042] Collecting state change data, operation sequence data and processing result data generated during the alarm processing process, and transmitting the state change data, the operation sequence data and the processing result data to a bidirectional feature extraction network in real time;

[0043] The bidirectional feature extraction network extracts features from the received data, updates the extracted new feature patterns to the feature library, and feeds back the updated feature data to the alarm rule generator; the alarm rule generator starts an incremental training process based on the received feature data and optimizes the existing alarm rules using a progressive learning strategy.

[0044] According to a second aspect of the embodiments of the present invention,

[0045] Provides an intelligent task alarm rule self-learning system based on support priorities, including:

[0046] The first unit is used to build a bidirectional feature extraction network to collect task operation data. The bidirectional feature extraction network includes a forward feature extraction layer and a reverse verification layer. The forward feature extraction layer obtains the task's running time, resource occupancy rate, execution status and business importance to build a real-time feature matrix. The reverse verification layer analyzes historical alarm data based on a graph neural network to obtain an alarm impact link diagram, and performs feature fusion on the real-time feature matrix and the alarm impact link diagram to obtain a task anomaly association model; the task anomaly association model is migrated and mapped with a business scenario knowledge base using a transfer learning algorithm to generate a task priority quantitative index that considers the business scenario, and an adaptive priority scoring system is built based on the task priority quantitative index;

[0047] The second unit is used to train an alarm rule generator based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link diagram, and continuously optimizes the accuracy of the alarm rule through adversarial training; associates the verified alarm rules with the task priority quantitative indicators, and constructs a multi-level alarm strategy tree with differentiated alarm thresholds. Each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and response mechanism of different priorities;

[0048] The third unit is used to start the incremental learning engine to monitor the task running status in real time. When it is detected that the task running indicators trigger the alarm conditions in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and at the same time feeds back the new data generated during the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to realize dynamic optimization of the alarm rules.

[0049] According to a third aspect of the embodiments of the present invention,

[0050] An electronic device is provided, comprising:

[0051] processor;

[0052] a memory for storing processor-executable instructions;

[0053] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0054] A fourth aspect of the embodiments of the present invention is:

[0055] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0056] The beneficial effects of this application are as follows:

[0057] 1. Improve alarm accuracy: This invention can more accurately identify task anomalies and generate more effective alarm rules through a bidirectional feature extraction network and an adversarial learning framework. The reverse verification layer uses historical alarm data to build an alarm impact link diagram, avoiding misjudgments that may be caused by relying solely on real-time data. The adversarial learning framework continuously optimizes the accuracy of alarm rules and reduces false positives and negatives through adversarial training of the generation network and the discriminant network.

[0058] 2. Realize intelligent alarm and response: The present invention realizes intelligent task alarm based on support priority. The task anomaly association model is migrated and mapped with the business scenario knowledge base through the transfer learning algorithm, and the task priority quantitative index considering the business scenario is generated, and an adaptive priority scoring system is constructed. Combined with the multi-level alarm strategy tree, differentiated alarm thresholds and response mechanisms can be implemented according to task priorities, effectively improving alarm efficiency and resource utilization. The incremental learning engine can dynamically adjust the alarm response strategy according to the current alarm scenario, further improving the intelligence level of the alarm.

[0059] 3. Enhance the generalization ability and adaptability of the alarm system: This invention uses the federated learning mechanism to collaboratively train the task anomaly association model between multiple business nodes. While protecting data privacy, it can improve the generalization ability and adaptability of the alarm rules, so that it can better adapt to different business scenarios and environmental changes. At the same time, the introduction of the incremental learning engine enables the system to continuously learn new data and scenarios, continuously optimize the alarm rules, and enhance the robustness and long-term effectiveness of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 A flowchart of a method for self-learning intelligent task alarm rules based on priority support according to an embodiment of the present invention;

[0061] Figure 2 The present invention is a schematic diagram of the structure of an intelligent task alarm rule self-learning system based on priority support according to an embodiment of the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0064] Figure 1 FIG. 1 is a flow chart of a method for self-learning intelligent task alarm rules based on priority support according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0065] S11. Construct a bidirectional feature extraction network to collect task operation data, wherein the bidirectional feature extraction network includes a forward feature extraction layer and a reverse verification layer, wherein the forward feature extraction layer obtains the task operation time, resource occupancy rate, execution status and business importance to construct a real-time feature matrix, and the reverse verification layer analyzes the historical alarm data based on the graph neural network to obtain an alarm impact link diagram, and performs feature fusion between the real-time feature matrix and the alarm impact link diagram to obtain a task anomaly association model; utilizes a transfer learning algorithm to perform migration mapping between the task anomaly association model and the business scenario knowledge base, generates a task priority quantitative index considering the business scenario, and constructs an adaptive priority scoring system based on the task priority quantitative index;

[0066] S12. Based on the task anomaly association model and the adaptive priority scoring system, an alarm rule generator is trained. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link diagram, and continuously optimizes the accuracy of the alarm rules through adversarial training; the verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds, and each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and response mechanism of different priorities;

[0067] S13. Start the incremental learning engine to monitor the task running status in real time. When it is detected that the task running indicators trigger the alarm conditions in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated in the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to achieve dynamic optimization of the alarm rules; based on the federated learning mechanism, the task anomaly association model is collaboratively trained among multiple business nodes to improve the generalization ability and adaptability of the alarm rules while protecting data privacy.

[0068] In an optional implementation, the forward feature extraction layer acquires the task running time, resource occupancy rate, execution status and business importance to construct a real-time feature matrix, the reverse verification layer analyzes the historical alarm data based on the graph neural network to obtain an alarm impact link diagram, and the real-time feature matrix is ​​fused with the alarm impact link diagram to obtain a task abnormality association model, including:

[0069] The running time, the resource occupancy rate, the execution status and the business importance are constructed into a four-dimensional feature vector; the forward feature extraction layer performs attention-weighted fusion on the four-dimensional feature vector to generate a real-time feature matrix representing the task running status;

[0070] The reverse verification layer uses a graph neural network to construct an alarm impact link graph, constructs the historical alarm data into graph nodes with alarm attributes, constructs graph edge connections based on the association relationship between alarms, extracts deep features of the graph nodes through graph convolution operations, and fuses the topological information of the graph edge connections to generate an alarm impact link graph that characterizes the alarm propagation mode; converts the node features in the alarm impact link graph into a feature space with the same dimension as the real-time feature matrix through feature projection;

[0071] A cross-modal feature fusion network is constructed. The cross-modal feature fusion network calculates the similarity between the real-time feature matrix and the node features of the alarm impact link graph to obtain the attention weight. Based on the attention weight, the real-time feature matrix and the node features of the alarm impact link graph are adaptively fused to generate a task anomaly association model.

[0072] A task anomaly correlation analysis method based on real-time features and historical alarms aims to accurately identify and predict potential task anomalies. This method integrates the real-time running status of tasks and historical alarm information to build a task anomaly correlation model, thereby improving the accuracy and efficiency of anomaly detection.

[0073] First, extract the features of the real-time running status of the task. Collect data such as the running time, resource utilization, execution status, and business importance of the task. For example, the running time of task A is 10 minutes, the resource utilization rate is 80%, the execution status is "running", and the business importance is "high". Quantify these indicators, for example, assign the execution status "running" to 1, and "failure" to 0. Assign the business importance "high" to 3, "medium" to 2, and "low" to 1. Construct these quantified indicators into a four-dimensional feature vector, for example, the feature vector of task A is [10,80,1,3]. Process the feature vector, for example, use the attention mechanism to assign different weights to each dimension, and finally generate a real-time feature matrix that represents the running status of the task. Assume that after attention weighted fusion, the feature vector of task A becomes [8,90,0.8,2.5], and the feature vectors of multiple tasks are combined to form a real-time feature matrix.

[0074] Next, analyze the historical alarm data and construct an alarm impact link graph. The historical alarm data is represented as graph nodes, and each node contains the attribute information of the alarm, such as the alarm type, occurrence time, and impact range. For example, an alarm node represents "database connection failure", the occurrence time is "October 26, 2024 10:00", and the impact range is "Server A". According to the association between alarms, such as alarm A triggers alarm B, connect edges between nodes to form an alarm impact link graph. The graph neural network is used to analyze the graph, and the deep features of the nodes are extracted through graph convolution operations. The topological information of the graph is integrated to generate an alarm impact link graph that represents the alarm propagation mode. For example, through the graph neural network, it is learned that "database connection failure" usually leads to "application crash", and this relationship is reflected in the alarm impact link graph.

[0075] Then, the real-time feature matrix and the alarm impact link graph are feature fused to construct a task anomaly association model. The node features in the alarm impact link graph are converted to the same dimensional space as the real-time feature matrix through feature projection operation. For example, the feature vector of the alarm node is also converted into a four-dimensional vector. The similarity between the node features in the real-time feature matrix and the alarm impact link graph is calculated to obtain the attention weight. For example, if the real-time feature vector of task A has a high similarity with the feature vector of the "database connection failure" alarm node, a higher attention weight is assigned. Based on the attention weight, the node features of the real-time feature matrix and the alarm impact link graph are adaptively fused to generate a task anomaly association model. This model reflects the correlation between the real-time running status of the task and the historical alarm, and can be used to identify and predict potential task anomalies.

[0076] The solution of this application can:

[0077] Improve the accuracy of anomaly detection: By integrating the real-time running status of tasks and historical alarm information, the abnormal characteristics of tasks can be more comprehensively characterized, thereby improving the accuracy of anomaly detection and reducing false positives and missed positives. Improve the efficiency of anomaly location: The alarm impact link diagram provides path information for alarm propagation, which can help quickly locate the root cause of the anomaly and shorten the troubleshooting time. Achieve prediction and early warning of anomalies: By analyzing the correlation between the real-time running status of tasks and historical alarms, potential abnormal risks can be predicted, and preventive measures can be taken in advance to avoid abnormalities.

[0078] In an optional implementation, the task anomaly association model is migrated and mapped with the business scenario knowledge base using a transfer learning algorithm to generate a task priority quantitative index considering the business scenario, and an adaptive priority scoring system is constructed based on the task priority quantitative index, including:

[0079] Construct a business scenario knowledge base, which includes a business process topology, system component dependencies, and operation and maintenance processing strategies, and use knowledge graph technology to construct the information in the business scenario knowledge base into knowledge entities and relationship edges, and map the knowledge entities and the relationship edges into knowledge feature vectors through a graph embedding algorithm;

[0080] A feature migration mapping network is constructed by using a migration learning algorithm. The feature migration mapping network calculates the semantic similarity between the feature representation of the task anomaly association model and the knowledge feature vector to generate an attention weight matrix. Based on the attention weight matrix, the knowledge feature vector is migrated and mapped to the feature space of the task anomaly association model. Residual connections are used to maintain the original feature information to generate a task feature representation that integrates business scenario knowledge.

[0081] Based on the task feature representation of the integrated business scenario knowledge, a task priority quantitative index is constructed, the task priority quantitative index includes a task urgency index, a business impact index and a resource consumption index, a hierarchical analysis method is used to calculate the weight coefficient of the task priority quantitative index, and the task priority quantitative index and the weight coefficient are weightedly combined to obtain a priority score;

[0082] An adaptive priority scoring system is constructed based on the priority scoring. The adaptive priority scoring system takes the task feature representation integrating the business scenario knowledge as the input state, takes the priority score as the output action, designs a reward function based on historical processing effects, and uses a deep reinforcement learning algorithm to continuously optimize the scoring strategy.

[0083] In order to build an adaptive system that can dynamically adjust task priorities according to business scenarios, it is necessary to first build a comprehensive business scenario knowledge base. This knowledge base contains the topological structure of the business process, such as the sequence and dependencies between various business links; the dependencies of system components, such as the call relationships between databases, servers, and applications; and operation and maintenance processing strategies, such as emergency plans for different types of failures. This information is organized using knowledge graph technology, abstracting concepts such as business processes, system components, and operation and maintenance strategies into knowledge entities, and using relationship edges to describe the connections between them. For example, "order processing" and "payment system" are two entities that can be connected by a "dependency" relationship. In order to facilitate subsequent calculations, a graph embedding algorithm is used to convert knowledge entities and relationship edges into numerical knowledge feature vectors. For example, the feature vectors of "order processing" and "payment system" can be expressed as [0.2, 0.5, 0.8] and [0.1, 0.7, 0.6], respectively.

[0084] Next, the transfer learning algorithm is used to associate the existing task anomaly association model with the constructed business scenario knowledge base. The task anomaly association model describes the association relationship between different types of task anomalies. For example, a database connection failure may lead to order processing failure. The goal of the transfer learning algorithm is to integrate the information in the business scenario knowledge base into the task anomaly association model. Specifically, a feature transfer mapping network is constructed, which calculates the semantic similarity between the feature representation of the task anomaly association model and the knowledge feature vector, and generates an attention weight matrix. For example, if the feature vector of "database connection failure" has a high similarity with the feature vector of "payment system", the corresponding attention weight will be relatively large. Then, according to the attention weight matrix, the knowledge feature vector is migrated and mapped to the feature space of the task anomaly association model, and the original task anomaly feature information is retained through residual connection, and finally the task feature representation that integrates the business scenario knowledge is generated. For example, the feature representation of "database connection failure" after integrating the business scenario knowledge can be [0.3, 0.6, 0.9], which contains both the original anomaly features and the features of "payment system".

[0085] Then, task priority quantitative indicators are constructed based on the task feature representation integrated with business scenario knowledge. These indicators include task urgency indicators, such as the deadline of the task; business impact indicators, such as the number of affected users; and resource consumption indicators, such as the CPU and memory resources required to complete the task. The weight coefficients of these indicators are determined using the hierarchical analysis method. For example, the weight coefficient of business impact may be higher than the weight coefficient of resource consumption. The task priority quantitative indicators are weighted and combined with the corresponding weight coefficients to obtain the final priority score. For example, the urgency, business impact, and resource consumption indicators of a task are 8, 9, and 7, respectively, and the corresponding weight coefficients are 0.3, 0.5, and 0.2, respectively. The priority score of the task is 8*0.3+9*0.5+7*0.2=8.3.

[0086] Finally, an adaptive priority scoring system is built based on priority scoring. The system takes the task feature representation that integrates business scenario knowledge as the input state and the priority score as the output action. By analyzing the historical processing effects, a reward function is designed to evaluate the scoring strategy of the system. For example, if the system gives a low score to a high-priority task, resulting in a delay in the processing of the task, a negative reward will be obtained. The deep reinforcement learning algorithm is used to continuously optimize the scoring strategy, so that the system can dynamically adjust the task priority according to the changes in the business scenario. For example, during peak hours, the system may pay more attention to business impact indicators, while during off-peak hours, the system may pay more attention to resource consumption indicators.

[0087] Suppose there is a task called "processing user payment failure", which is associated with the "payment system". The knowledge feature vector of the "payment system" is [0.1, 0.7, 0.6]. The feature vector of "payment failure" in the task anomaly association model is [0.2, 0.3, 0.5]. After the feature migration mapping network calculation, the task feature that integrates the business scenario knowledge is expressed as [0.15, 0.4, 0.55]. The urgency, business impact and resource consumption indicators of this task are 7, 9 and 6 respectively, and the corresponding weight coefficients are 0.2, 0.6 and 0.2 respectively. The priority score of this task is 7*0.2+9*0.6+6*0.2=7.8.

[0088] The beneficial effects can be summarized into the following three aspects:

[0089] Improve task processing efficiency: By giving priority to high-priority tasks, critical business interruption time can be reduced and overall operational efficiency can be improved.

[0090] Optimize resource allocation: Dynamically allocate resources according to task priorities to avoid resource waste and improve resource utilization.

[0091] Enhanced system adaptability: The introduction of deep reinforcement learning algorithms enables the system to dynamically adjust scoring strategies according to changes in business scenarios, enhancing the adaptability and robustness of the system.

[0092] In an optional implementation, an alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link graph. The accuracy of the alarm rule is continuously optimized through adversarial training, including:

[0093] Constructing an initial training set for a rule generator, taking the abnormal feature matrix in the task abnormality association model as the basic data of the rule triggering condition, taking the scoring result of the adaptive priority scoring system as the basis for dividing the rule alarm level, taking the historical alarm processing record as the reference data for the rule processing action, and vectorizing and encoding the basic data, the division basis and the reference data to obtain a rule training vector;

[0094] Constructing a generative network in an adversarial learning framework, wherein the generative network adopts a multi-layer neural network structure, inputs the rule training vector to generate an initial alarm rule candidate set, wherein each rule in the initial alarm rule candidate set includes a trigger condition threshold, an alarm level, and a processing action sequence;

[0095] Constructing a discriminant network in an adversarial learning framework, wherein the discriminant network constructs a rule evaluation model based on the alarm impact link diagram, wherein the rule evaluation model calculates the similarity between the trigger condition threshold of the rule in the initial alarm rule candidate set and the historical abnormal characteristics to obtain a rule accuracy score, calculates the consistency between the alarm level of the rule and the result of the adaptive priority scoring system to obtain a rule rationality score, and verifies the executableness of the processing action sequence of the rule in the current system environment to obtain a rule effectiveness score;

[0096] Based on the rule accuracy score, the rule rationality score and the rule effectiveness score, the discriminant network screens the initial alarm rule candidate set, and inputs the scoring result as a feedback signal into the generation network to guide the generation network to adjust the rule generation strategy, and obtains the accuracy of the final alarm rule through multiple rounds of adversarial training iterative optimization.

[0097] The alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The generator adopts an adversarial learning framework. Its generation network continuously generates candidate sets of alarm rules, and its discriminant network evaluates the effectiveness of the candidate sets based on the alarm impact link graph. The accuracy of the alarm rules is continuously optimized through adversarial training. The implementation method and beneficial effects are described in detail below.

[0098] First, construct the initial training set of the rule generator. Extract the abnormal feature matrix from the task abnormality association model as the basic data of the rule triggering condition. For example, the abnormal features of a task may include indicators such as CPU usage, memory occupancy, and disk IO. Assuming that we have collected 1,000 historical abnormal samples, each sample contains 5 features, then we can form an abnormal feature matrix with 1,000 rows and 5 columns. At the same time, obtain the scoring results from the adaptive priority scoring system as the basis for dividing the rule alarm level. For example, the scoring range is 0-100, 0-30 corresponds to low-level alarms, 30-70 corresponds to medium-level alarms, and 70-100 corresponds to high-level alarms. In addition, extract the processing action sequence from the historical alarm processing record as the reference data for the rule processing action. For example, the processing action sequence of a certain alarm may include "restarting the service", "sending a notification", "recording a log", etc. The above basic data, division basis, and reference data are vectorized and encoded, such as using methods such as one-hot encoding or word embedding, to obtain the rule training vector. For example, a rule training vector can be represented as [CPU usage, memory usage, disk IO, priority score, processing action 1, processing action 2, ...].

[0099] Next, construct a generative network in the adversarial learning framework. The generative network uses a multi-layer neural network structure, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM), inputs a rule training vector, and generates an initial alarm rule candidate set. Each candidate rule contains a trigger condition threshold, an alarm level, and a processing action sequence. For example, a candidate rule can be expressed as: "When the CPU usage exceeds 90% and the memory usage exceeds 80%, a high-level alarm is triggered, the service is restarted, and a notification is sent." The output of the generative network is a set of multiple candidate rules.

[0100] Then, a discriminant network in the adversarial learning framework is constructed. The discriminant network constructs a rule evaluation model based on the alarm impact link graph. The alarm impact link graph describes the association between different alarms and their impact on the system. For example, one alarm may trigger another alarm, or one alarm may cause system performance to degrade. The rule evaluation model calculates three scores for candidate rules: accuracy score, rationality score, and effectiveness score. The accuracy score is obtained by calculating the similarity between the rule trigger condition threshold and the historical abnormal characteristics. For example, the cosine similarity or Euclidean distance can be used to measure the similarity. The rationality score is obtained by calculating the consistency between the rule alarm level and the result of the adaptive priority scoring system. For example, the cross entropy loss function can be used to measure the consistency. The effectiveness score is obtained by verifying the executableness of the rule processing action sequence in the current system environment. For example, it can be checked whether the resources required for processing the action are available.

[0101] Finally, the discriminant network screens the candidate rules according to the three scores and inputs the score results as feedback signals into the generative network to guide the generative network to adjust the rule generation strategy. Through multiple rounds of adversarial training and iterative optimization, accurate alarm rules are finally obtained.

[0102] The solution of this application can:

[0103] Improve the accuracy of alarm rules: Through the adversarial learning framework, the generator network and the discriminator network compete with each other to continuously optimize the alarm rules so that they can more accurately identify abnormal situations. Improve the rationality of alarm rules: Based on the adaptive priority scoring system, different alarm levels can be set according to the severity of the abnormality to avoid false positives and false negatives. Improve the effectiveness of alarm rules: By verifying the executability of the processing action, it can be ensured that the alarm rules can effectively handle abnormal situations.

[0104] In an optional implementation, the verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds, and each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and a response mechanism of different priorities, including:

[0105] Establishing a mapping relationship between the trigger conditions in the verified alarm rule set and the task priority quantitative index, stratifying the alarm rule set according to the adaptive priority score, and constructing a basic framework of a multi-level alarm strategy tree, wherein each node of the multi-level alarm strategy tree includes a priority interval, an alarm trigger condition, and a response processing mechanism, and organizing alarm rules with similar adaptive priority scores into the same node;

[0106] For each node of the multi-level alarm strategy tree, based on the historical trigger data of the alarm rules in the node, a density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node;

[0107] A response processing mechanism is configured for each node of the multi-level alarm strategy tree, independent processing resources are allocated to priority nodes above the alarm threshold and a quick response channel is set up, a resource pool dynamic allocation mechanism is adopted for priority nodes aligned with the alarm threshold, and a batch processing mode is used for priority nodes below the alarm threshold, thereby building an alarm upgrade path between nodes.

[0108] Based on the task priority quantitative indicators and alarm rules, a multi-level alarm strategy tree is constructed to achieve differentiated alarms and responses. The specific implementation methods are as follows:

[0109] First, prepare the alarm rule set and task priority quantitative indicators. The alarm rule set contains various trigger conditions, such as CPU utilization exceeding 90%, disk space less than 20%, etc. Task priority quantitative indicators such as availability, real-time performance, and impact range are obtained through expert scoring or machine learning model calculation, ranging from 0 to 10 points, and the higher the score, the higher the priority.

[0110] Next, the alarm rules are layered according to the adaptive priority score to build the basic framework of the multi-level alarm strategy tree. The priority score is divided into different intervals, such as 0-3 low priority, 4-6 medium priority, and 7-10 high priority. Alarm rules with the same priority interval are grouped into the same node. In this example, the strategy tree contains three nodes: high priority node (rule 1, rule 4), medium priority node (rule 2, rule 3), and low priority node (rule 5). Each node contains the priority interval, alarm trigger condition, and response processing mechanism.

[0111] Then, for each node, based on historical trigger data, the density clustering method is used to analyze the distribution of alarm features and establish a differentiated alarm threshold calculation model. For example, high-priority nodes contain CPU utilization and memory usage alarms. Analysis of historical data shows that CPU utilization usually fluctuates above 80%, occasionally reaching 90%, while memory usage is usually stable below 90% and rarely exceeds 95%. Therefore, the CPU utilization alarm threshold is set to 95%, and the memory usage alarm threshold is set to 98% to improve the accuracy of the alarm. The disk space alarm threshold of the medium-priority node is set to 15%, and the network delay alarm threshold is set to 120ms. The connection number alarm threshold of the low-priority node is too high and is 1.2 times the current peak value.

[0112] Finally, configure the response processing mechanism for each node. High-priority nodes are allocated independent processing resources and fast response channels, such as SMS and phone notifications, to ensure immediate response and processing. Medium-priority nodes use a dynamic resource pool allocation mechanism to dynamically adjust resource allocation based on the number and severity of alarms. Low-priority nodes use batch processing mode, such as regular summary processing, to reduce processing costs. At the same time, establish an alarm escalation path between nodes. For example, if the alarm of a medium-priority node continues to be unresolved for a period of time, it will be upgraded to a high-priority node, triggering a higher-level response mechanism.

[0113] The solution of this application can:

[0114] Improve the accuracy of alarms. Through density clustering analysis of alarm feature distribution, a differentiated alarm threshold calculation model is established to avoid false alarms and missed alarms of traditional fixed threshold methods, improve the accuracy of alarms, and enable operation and maintenance personnel to focus more on handling truly important alarms. Optimize resource allocation. The multi-level alarm strategy tree allocates different processing resources and response mechanisms according to priority, concentrates limited resources on high-priority alarms, and ensures the stable operation of key businesses. Low-priority alarms use batch processing mode to reduce processing costs and avoid resource waste. Improve alarm processing efficiency. The multi-level alarm strategy tree and differentiated response mechanism enable operation and maintenance personnel to quickly identify and handle high-priority alarms, shorten fault recovery time, improve alarm processing efficiency, and ensure business continuity.

[0115] In an optional implementation, based on the historical trigger data of the alarm rules in the node, a density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node, including:

[0116] Extract historical trigger data of alarm rules from the node, construct the indicator value, trigger frequency, duration and impact range in the historical trigger data into an initial feature vector, and perform data cleaning and standardization on the initial feature vector to obtain a standardized feature vector;

[0117] A density clustering method is used to perform cluster analysis on the standardized feature vector, wherein the density clustering method dynamically calculates the density accessibility of sample points in the feature space through an adaptive neighborhood radius mechanism to obtain a cluster set of alarm feature distribution;

[0118] For each cluster in the cluster set, the center point position, density distribution and boundary range are calculated to obtain a cluster feature set, and an alarm threshold calculation model is established based on the cluster feature set, wherein the alarm threshold calculation model includes a cluster feature fusion unit and a threshold dynamic adjustment unit;

[0119] Obtaining priority interval information of the node, the alarm threshold calculation model assigns weight coefficients to different clusters in the cluster feature set according to the priority interval information, and generates initial alarm trigger conditions with differentiated sensitivity by weighted fusion of various cluster features;

[0120] Collect alarm trigger data within the sliding time window, calculate the false alarm rate, missed alarm rate and timeliness index of the initial alarm trigger condition to obtain a monitoring index set, and when the index value in the monitoring index set deviates from the expected threshold, the expected threshold dynamic adjustment unit calculates the adjustment parameter according to the deviation degree and the node priority;

[0121] The initial alarm trigger condition is adaptively adjusted based on the adjustment parameter, and the historical data of the node business load is analyzed to identify the load cycle characteristics, a mapping relationship model between the initial alarm trigger condition and the load state is established, and the sensitivity of the initial alarm trigger condition is optimized according to the current load level.

[0122] The construction and application method of the alarm threshold calculation model are specifically implemented as follows:

[0123] First, extract the historical trigger data of node-related alarm rules from the monitoring system. This data contains information such as the timestamp, indicator value, trigger frequency, duration, and impact range of each alarm trigger. For example, the historical data of a node CPU usage alarm includes: 2024-07-27 10:00:00, CPU usage 95%, trigger frequency 1 time / minute, duration 5 minutes, impact range single node; 2024-07-27 10:05:00, CPU usage 98%, trigger frequency 2 times / minute, duration 10 minutes, impact range single node, etc.

[0124] Next, the extracted historical data is cleaned and standardized. Data cleaning includes operations such as removing duplicate data and filling missing values. For example, if the CPU usage data at a certain point in time is missing, it can be filled with the CPU usage at the previous point in time. Standardization is to convert indicator data of different dimensions into a unified numerical range, for example, normalizing indicator values ​​such as CPU usage and memory usage to between 0 and 1. The cleaned and standardized data constitute a standardized feature vector, for example: [0.95, 1, 5, 1], [0.98, 2, 10, 1].

[0125] Then, the density clustering method is used to perform cluster analysis on the standardized feature vectors. This method dynamically calculates the density accessibility of sample points in the feature space through an adaptive neighborhood radius mechanism. Specifically, an initial neighborhood radius is set first, and then the radius size is dynamically adjusted according to the number of neighbors around the sample point, so that sample points in areas with higher density are more likely to be aggregated together, while sample points in areas with lower density are more likely to form independent clusters. Finally, a cluster set of alarm feature distribution is obtained. For example, the above CPU usage alarm data is clustered into high-load clusters and low-load clusters.

[0126] For each cluster, its center point position, density distribution, and boundary range are calculated to obtain the cluster feature set. For example, the center point of the high-load cluster is [0.98, 2, 10, 1], the density distribution is concentrated, and the boundary range is small; the center point of the low-load cluster is [0.8, 0.5, 2, 1], the density distribution is relatively dispersed, and the boundary range is large.

[0127] An alarm threshold calculation model is established based on the cluster feature set. The model includes a cluster feature fusion unit and a threshold dynamic adjustment unit. The cluster feature fusion unit assigns weight coefficients to different clusters according to the priority interval information of the node. For example, for high-priority nodes, the weight coefficient of the high-load cluster is set to 0.8, and the weight coefficient of the low-load cluster is set to 0.2; for low-priority nodes, the weight coefficient of the high-load cluster is set to 0.5, and the weight coefficient of the low-load cluster is set to 0.5. The initial alarm trigger condition with differentiated sensitivity is generated by weighted fusion of various cluster features.

[0128] In the sliding time window (for example, the past hour), collect alarm trigger data, calculate the false alarm rate, missed alarm rate and timeliness of the initial alarm trigger condition, and obtain the monitoring indicator set. For example, the false alarm rate is 2%, the missed alarm rate is 1%, and the average alarm response time is 1 minute. When the indicator value in the monitoring indicator set deviates from the expected threshold (for example, the false alarm rate exceeds 5%), the threshold dynamic adjustment unit calculates the adjustment parameter according to the degree of deviation and the node priority. For example, if the false alarm rate of the high priority node reaches 8%, the calculated adjustment parameter is -0.1.

[0129] Adaptively adjust the initial alarm trigger conditions based on the adjustment parameters. For example, lower the alarm threshold of the high-load cluster by 10%. At the same time, analyze the historical data of the node business load to identify the load cycle characteristics. For example, 9 am to 12 pm every day is the business peak period. Establish a mapping relationship model between the initial alarm trigger conditions and the load status, and optimize the sensitivity of the initial alarm trigger conditions according to the current load level. For example, increase the alarm threshold during the business peak period and lower the alarm threshold during the business trough period.

[0130] The solution of this application can:

[0131] Improve alarm accuracy. Analyze the distribution of alarm features through density clustering method and establish a differentiated alarm threshold calculation model, which can effectively reduce the false alarm rate and missed alarm rate and improve the accuracy of alarm. Enhance alarm flexibility. Dynamically adjust the sensitivity of alarm trigger conditions according to the priority interval of the node, which can make the alarm strategy more flexible and better meet the needs of different scenarios. Optimize alarm timeliness. By adaptively adjusting the alarm threshold and combining the node load status, abnormal situations can be discovered and handled in time, improving the timeliness of alarms.

[0132] In an optional implementation, an incremental learning engine is started to monitor the task running status in real time. When it is detected that the task running indicator triggers the alarm condition in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated in the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time, so as to realize the dynamic optimization of the alarm rules, including:

[0133] Starting an incremental learning engine to monitor task running status data in real time, the incremental learning engine obtains task running indicators through a multidimensional time series analysis method, and uses a sliding window mechanism to preprocess the task running indicators to obtain an indicator feature sequence;

[0134] Inputting the indicator feature sequence into the multi-level alarm strategy tree for alarm condition matching, and when the indicator feature sequence triggers the alarm condition in the multi-level alarm strategy tree, constructing an alarm feature vector of the current alarm scenario based on the indicator feature sequence;

[0135] Performing scene recognition in a pre-established scene feature library according to the alarm feature vector, and selecting a corresponding response strategy template from a response strategy library based on the scene recognition result, wherein the response strategy template includes a processing step sequence, a resource configuration scheme, and a time limit parameter;

[0136] The incremental learning engine dynamically adjusts the selected response strategy template according to the current system load status and resource utilization, generates an alarm response strategy adapted to the current scenario, and sends the alarm response strategy to the alarm processing module;

[0137] Collecting state change data, operation sequence data and processing result data generated during the alarm processing process, and transmitting the state change data, the operation sequence data and the processing result data to a bidirectional feature extraction network in real time;

[0138] The bidirectional feature extraction network extracts features from the received data, updates the extracted new feature patterns to the feature library, and feeds back the updated feature data to the alarm rule generator; the alarm rule generator starts an incremental training process based on the received feature data and optimizes the existing alarm rules using a progressive learning strategy.

[0139] A dynamic alarm system and method based on incremental learning are used to monitor the task running status in real time and dynamically adjust the alarm response strategy.

[0140] First, initialize the system. The system includes modules such as incremental learning engine, multi-level alarm strategy tree, scenario feature library, response strategy library, alarm processing module, bidirectional feature extraction network and alarm rule generator. Among them, the multi-level alarm strategy tree predefines various alarm conditions and corresponding alarm levels; the scenario feature library stores feature vectors of different alarm scenarios; the response strategy library contains response strategy templates corresponding to various alarm scenarios, and each template contains a processing step sequence, resource allocation scheme and time limit parameters.

[0141] Next, start the incremental learning engine and start real-time monitoring of the task running status data. The incremental learning engine collects various indicators of task operation, such as CPU usage, memory occupancy, disk IO, etc., through monitoring systems such as Prometheus. The incremental learning engine uses a sliding window mechanism to preprocess the collected indicators. For example, set a time window of 60 seconds and collect data every 10 seconds, then a window contains 6 data points. These 6 data points constitute an indicator feature sequence. Assuming that the 6 data points of CPU usage are 60%, 65%, 70%, 75%, 80%, and 85%, respectively, the indicator feature sequence is [60, 65, 70, 75, 80, 85].

[0142] Then, the indicator feature sequence is input into the multi-level alarm strategy tree for alarm condition matching. For example, the multi-level alarm strategy tree can be set as follows: the first-level alarm condition is that the CPU usage exceeds 70% for three consecutive times, the second-level alarm condition is that the CPU usage exceeds 80% for three consecutive times, and the third-level alarm condition is that the CPU usage exceeds 90% for three consecutive times. In the above example, the indicator feature sequence [60, 65, 70, 75, 80, 85] triggers the first-level and second-level alarm conditions.

[0143] When the indicator feature sequence triggers an alarm condition, an alarm feature vector of the current alarm scenario is constructed based on the sequence. For example, the statistical features such as the average value, maximum value, minimum value, and variance of the indicator feature sequence can be used as the alarm feature vector.

[0144] The scene recognition is performed in the scene feature library based on the alarm feature vector. For example, the similarity between the alarm feature vector and each feature vector in the scene feature library can be calculated, and the scene with the highest similarity is selected as the recognition result. Assume that there is a feature vector [70, 80, 55, 90] in the scene feature library, and its corresponding scene is "CPU load is too high". If the calculated similarity is the highest, the current alarm scene is identified as "CPU load is too high".

[0145] Based on the scenario recognition results, select the corresponding response policy template from the response policy library. Assume that the response policy template corresponding to the "CPU load is too high" scenario is: 1. Send an alarm notification; 2. Restart related services; 3. Increase computing resources. The time limit is 10 minutes. The resource configuration plan is to add 2 CPU cores.

[0146] The incremental learning engine dynamically adjusts the selected response strategy template according to the current system load status and resource utilization, and generates an alarm response strategy that adapts to the current scenario. For example, if the current system load is high and resources are tight, the response strategy can be adjusted to: 1. Send an alarm notification; 2. Limit some non-critical services; 3. Try to release the cache. The time limit is 5 minutes. The resource configuration plan is to add 1 CPU core.

[0147] The generated alarm response strategy is sent to the alarm processing module for execution. The alarm processing module performs corresponding operations according to the received strategy, such as sending alarm notifications, restarting services, and increasing computing resources.

[0148] Collect state change data, operation sequence data, and processing result data generated during the alarm processing process. For example, record the state change of service restart, the executed operation sequence, and the final processing result.

[0149] The collected data is transmitted to the bidirectional feature extraction network in real time. The bidirectional feature extraction network extracts features from the received data, such as the frequency of state changes, the mode of operation sequence, the success rate of processing results, etc.

[0150] The extracted new feature patterns are updated to the feature library, and the updated feature data is fed back to the alarm rule generator. The alarm rule generator starts the incremental training process based on the received feature data and optimizes the existing alarm rules using a progressive learning strategy. For example, if a certain operation sequence is found to effectively reduce the CPU load, the operation sequence can be added to the response strategy library and the corresponding alarm rule can be updated.

[0151] The solution of this application can:

[0152] Improve alarm accuracy: Through incremental learning and dynamic adjustment, the system can automatically optimize alarm rules and response strategies according to actual conditions, thereby reducing false alarms and missed alarms and improving alarm accuracy. Enhance alarm response efficiency: The system can dynamically adjust the response strategy according to the current system status and select the most appropriate processing solution, thereby shortening the alarm processing time and improving response efficiency. Reduce operation and maintenance costs: The system can automatically learn and optimize, reduce manual intervention, and thus reduce operation and maintenance costs.

[0153] Figure 2 FIG. 1 is a schematic diagram of the structure of an intelligent task alarm rule self-learning system based on priority support according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0154] The first unit is used to build a bidirectional feature extraction network to collect task operation data. The bidirectional feature extraction network includes a forward feature extraction layer and a reverse verification layer. The forward feature extraction layer obtains the task's running time, resource occupancy rate, execution status and business importance to build a real-time feature matrix. The reverse verification layer analyzes historical alarm data based on a graph neural network to obtain an alarm impact link diagram, and performs feature fusion on the real-time feature matrix and the alarm impact link diagram to obtain a task anomaly association model; the task anomaly association model is migrated and mapped with a business scenario knowledge base using a transfer learning algorithm to generate a task priority quantitative index that considers the business scenario, and an adaptive priority scoring system is built based on the task priority quantitative index;

[0155] The second unit is used to train an alarm rule generator based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link diagram, and continuously optimizes the accuracy of the alarm rule through adversarial training; associates the verified alarm rules with the task priority quantitative indicators, and constructs a multi-level alarm strategy tree with differentiated alarm thresholds. Each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and response mechanism of different priorities;

[0156] The third unit is used to start the incremental learning engine to monitor the task running status in real time. When it is detected that the task running indicators trigger the alarm conditions in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and at the same time feeds back the new data generated during the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to realize dynamic optimization of the alarm rules.

[0157] According to a third aspect of the embodiments of the present invention,

[0158] An electronic device is provided, comprising:

[0159] processor;

[0160] a memory for storing processor-executable instructions;

[0161] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0162] A fourth aspect of the embodiments of the present invention is:

[0163] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0164] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent task alarm rule self-learning method based on supporting priority, characterized in that: include: A bidirectional feature extraction network is constructed to collect task operation data. The bidirectional feature extraction network includes a forward feature extraction layer and a reverse verification layer. The forward feature extraction layer obtains the task operation time, resource occupancy rate, execution status and business importance to construct a real-time feature matrix. The reverse verification layer analyzes the historical alarm data based on the graph neural network to obtain an alarm impact link diagram. The real-time feature matrix and the alarm impact link diagram are feature-fused to obtain a task anomaly association model. The task anomaly association model is migrated and mapped to a business scenario knowledge base using a transfer learning algorithm to generate a task priority quantitative index that considers the business scenario, and an adaptive priority scoring system is constructed based on the task priority quantitative index. An alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link diagram, and continuously optimizes the accuracy of the alarm rule through adversarial training; the verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds, and each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and response mechanism of different priorities; An incremental learning engine is started to monitor the task running status in real time. When it is detected that the task running indicators trigger the alarm conditions in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated during the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to achieve dynamic optimization of the alarm rules; based on the federated learning mechanism, the task anomaly association model is collaboratively trained among multiple business nodes to improve the generalization ability and adaptability of the alarm rules while protecting data privacy.

2. The method according to claim 1, characterized in that The forward feature extraction layer obtains the task's running time, resource occupancy rate, execution status and business importance to build a real-time feature matrix. The reverse verification layer analyzes historical alarm data based on a graph neural network to obtain an alarm impact link graph. The real-time feature matrix and the alarm impact link graph are subjected to feature fusion to obtain a task abnormality association model, including: The running time, the resource occupancy rate, the execution status and the business importance are constructed into a four-dimensional feature vector; the forward feature extraction layer performs attention-weighted fusion on the four-dimensional feature vector to generate a real-time feature matrix representing the task running status; The reverse verification layer uses a graph neural network to construct an alarm impact link graph, constructs the historical alarm data into graph nodes with alarm attributes, constructs graph edge connections based on the association relationship between alarms, extracts deep features of the graph nodes through graph convolution operations, and fuses the topological information of the graph edge connections to generate an alarm impact link graph that characterizes the alarm propagation mode; converts the node features in the alarm impact link graph into a feature space with the same dimension as the real-time feature matrix through feature projection; A cross-modal feature fusion network is constructed. The cross-modal feature fusion network calculates the similarity between the real-time feature matrix and the node features of the alarm impact link graph to obtain the attention weight. Based on the attention weight, the real-time feature matrix and the node features of the alarm impact link graph are adaptively fused to generate a task anomaly association model.

3. The method according to claim 1, characterized in that Using a transfer learning algorithm to perform migration mapping between the task anomaly association model and the business scenario knowledge base, generating a task priority quantitative index considering the business scenario, and constructing an adaptive priority scoring system based on the task priority quantitative index includes: Construct a business scenario knowledge base, which includes a business process topology, system component dependencies, and operation and maintenance processing strategies, and use knowledge graph technology to construct the information in the business scenario knowledge base into knowledge entities and relationship edges, and map the knowledge entities and the relationship edges into knowledge feature vectors through a graph embedding algorithm; A feature migration mapping network is constructed by using a migration learning algorithm. The feature migration mapping network calculates the semantic similarity between the feature representation of the task anomaly association model and the knowledge feature vector to generate an attention weight matrix. Based on the attention weight matrix, the knowledge feature vector is migrated and mapped to the feature space of the task anomaly association model. Residual connections are used to maintain the original feature information to generate a task feature representation that integrates business scenario knowledge. Based on the task feature representation of the integrated business scenario knowledge, a task priority quantitative index is constructed, the task priority quantitative index includes a task urgency index, a business impact index and a resource consumption index, a hierarchical analysis method is used to calculate the weight coefficient of the task priority quantitative index, and the task priority quantitative index and the weight coefficient are weightedly combined to obtain a priority score; An adaptive priority scoring system is constructed based on the priority scoring. The adaptive priority scoring system takes the task feature representation integrating the business scenario knowledge as the input state, takes the priority score as the output action, designs a reward function based on historical processing effects, and uses a deep reinforcement learning algorithm to continuously optimize the scoring strategy.

4. The method according to claim 1, characterized in that The alarm rule generator is trained based on the task anomaly association model and the adaptive priority scoring system. The alarm rule generator adopts an adversarial learning framework, and its generation network continuously generates an alarm rule candidate set. Its discriminant network evaluates the effectiveness of the alarm rule candidate set based on the alarm impact link graph. The accuracy of the alarm rule is continuously optimized through adversarial training, including: Constructing an initial training set for a rule generator, taking the abnormal feature matrix in the task abnormality association model as the basic data of the rule triggering condition, taking the scoring result of the adaptive priority scoring system as the basis for dividing the rule alarm level, taking the historical alarm processing record as the reference data for the rule processing action, and vectorizing and encoding the basic data, the division basis and the reference data to obtain a rule training vector; Constructing a generative network in an adversarial learning framework, wherein the generative network adopts a multi-layer neural network structure, inputs the rule training vector to generate an initial alarm rule candidate set, wherein each rule in the initial alarm rule candidate set includes a trigger condition threshold, an alarm level, and a processing action sequence; Constructing a discriminant network in an adversarial learning framework, wherein the discriminant network constructs a rule evaluation model based on the alarm impact link diagram, wherein the rule evaluation model calculates the similarity between the trigger condition threshold of the rule in the initial alarm rule candidate set and the historical abnormal characteristics to obtain a rule accuracy score, calculates the consistency between the alarm level of the rule and the result of the adaptive priority scoring system to obtain a rule rationality score, and verifies the executableness of the processing action sequence of the rule in the current system environment to obtain a rule effectiveness score; Based on the rule accuracy score, the rule rationality score and the rule effectiveness score, the discriminant network screens the initial alarm rule candidate set, and inputs the scoring result as a feedback signal into the generation network to guide the generation network to adjust the rule generation strategy, and obtains the accuracy of the final alarm rule through multiple rounds of adversarial training iterative optimization.

5. The method according to claim 1, characterized in that The verified alarm rules are associated with the task priority quantitative indicators to construct a multi-level alarm strategy tree with differentiated alarm thresholds. Each node of the multi-level alarm strategy tree corresponds to an alarm trigger condition and a response mechanism of different priorities, including: Establishing a mapping relationship between the trigger conditions in the verified alarm rule set and the task priority quantitative index, stratifying the alarm rule set according to the adaptive priority score, and constructing a basic framework of a multi-level alarm strategy tree, wherein each node of the multi-level alarm strategy tree includes a priority interval, an alarm trigger condition, and a response processing mechanism, and organizing alarm rules with similar adaptive priority scores into the same node; For each node of the multi-level alarm strategy tree, based on the historical trigger data of the alarm rules in the node, a density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node; A response processing mechanism is configured for each node of the multi-level alarm strategy tree, independent processing resources are allocated to priority nodes above the alarm threshold and a quick response channel is set up, a resource pool dynamic allocation mechanism is adopted for priority nodes aligned with the alarm threshold, and a batch processing mode is used for priority nodes below the alarm threshold, thereby building an alarm upgrade path between nodes.

6. The method according to claim 5, characterized in that Based on the historical trigger data of the alarm rules in the node, the density clustering method is used to analyze the distribution of alarm features, and a differentiated alarm threshold calculation model is established. The alarm threshold calculation model dynamically adjusts the sensitivity of the alarm trigger condition according to the priority interval of the node, including: Extract historical trigger data of alarm rules from the node, construct the indicator value, trigger frequency, duration and impact range in the historical trigger data into an initial feature vector, and perform data cleaning and standardization on the initial feature vector to obtain a standardized feature vector; A density clustering method is used to perform cluster analysis on the standardized feature vector, wherein the density clustering method dynamically calculates the density accessibility of sample points in the feature space through an adaptive neighborhood radius mechanism to obtain a cluster set of alarm feature distribution; For each cluster in the cluster set, the center point position, density distribution and boundary range are calculated to obtain a cluster feature set, and an alarm threshold calculation model is established based on the cluster feature set, wherein the alarm threshold calculation model includes a cluster feature fusion unit and a threshold dynamic adjustment unit; Obtaining priority interval information of the node, the alarm threshold calculation model assigns weight coefficients to different clusters in the cluster feature set according to the priority interval information, and generates initial alarm trigger conditions with differentiated sensitivity by weighted fusion of various cluster features; Collect alarm trigger data within the sliding time window, calculate the false alarm rate, missed alarm rate and timeliness index of the initial alarm trigger condition to obtain a monitoring index set, and when the index value in the monitoring index set deviates from the expected threshold, the threshold dynamic adjustment unit calculates the adjustment parameter according to the deviation degree and the node priority; The initial alarm trigger condition is adaptively adjusted based on the adjustment parameter, and the historical data of the node business load is analyzed to identify the load cycle characteristics, a mapping relationship model between the initial alarm trigger condition and the load state is established, and the sensitivity of the initial alarm trigger condition is optimized according to the current load level.

7. The method according to claim 5, characterized in that The incremental learning engine is started to monitor the task running status in real time. When it is detected that the task running indicator triggers the alarm condition in the multi-level alarm strategy tree, the incremental learning engine dynamically adjusts the alarm response strategy according to the current alarm scenario, and feeds back the new data generated in the alarm processing process to the bidirectional feature extraction network and the alarm rule generator in real time to realize the dynamic optimization of the alarm rules, including: Starting an incremental learning engine to monitor task running status data in real time, the incremental learning engine obtains task running indicators through a multidimensional time series analysis method, and uses a sliding window mechanism to preprocess the task running indicators to obtain an indicator feature sequence; Inputting the indicator feature sequence into the multi-level alarm strategy tree for alarm condition matching, and when the indicator feature sequence triggers the alarm condition in the multi-level alarm strategy tree, constructing an alarm feature vector of the current alarm scenario based on the indicator feature sequence; Performing scene recognition in a pre-established scene feature library according to the alarm feature vector, and selecting a corresponding response strategy template from a response strategy library based on the scene recognition result, wherein the response strategy template includes a processing step sequence, a resource configuration scheme, and a time limit parameter; The incremental learning engine dynamically adjusts the selected response strategy template according to the current system load status and resource utilization, generates an alarm response strategy adapted to the current scenario, and sends the alarm response strategy to the alarm processing module; The state change data, operation sequence data and processing result data generated during the alarm processing are collected, and the state change data, the operation sequence data and the processing result data are transmitted to the bidirectional feature extraction network in real time; the bidirectional feature extraction network extracts features from the received data, updates the extracted new feature patterns to the feature library, and feeds back the updated feature data to the alarm rule generator; the alarm rule generator starts an incremental training process based on the received feature data, and optimizes the existing alarm rules using a progressive learning strategy.

Citation Information

Patent Citations

  • Perceptual security protection method, system and equipment based on network port protection device

    CN118611997A

  • Advanced detection of identity-based attacks to assure identity fidelity in information technology environments

    US20210084073A1