Intelligent monitoring method and system based on task operation history information
By building task operation profiles and pattern matching algorithms, we monitor task health in real time and generate processing strategies, solving the problems of dynamic adjustment and automated diagnosis of task monitoring systems in existing technologies and achieving efficient anomaly identification and processing.
Patent Information
- Application Number
- CN202510919193.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing task monitoring system lacks in-depth analysis of the historical operation patterns of tasks and is unable to dynamically adjust monitoring strategies, resulting in high false alarm rates and frequent missed alarms. It also lacks an automated intelligent diagnosis mechanism, relies on manual experience and is inefficient, and is unable to learn and optimize from historical anomalies.
Build task operation profiles based on task operation history information, monitor and calculate health scores in real time, use pattern matching algorithms to identify abnormal types and generate targeted processing strategies, and continuously optimize the intelligent diagnosis mechanism.
It achieves all-round real-time monitoring of tasks, quickly identifies anomalies and handles them automatically, reduces manual workload, improves system operation and maintenance efficiency, reduces the impact of anomalies, and extends the system life cycle.
Smart Images

Figure CN120410162B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to intelligent monitoring technology, and in particular to an intelligent monitoring method and system based on task operation history information. Background Art
[0002] With the rapid development of information technology, the business processes of enterprises and organizations are increasingly dependent on the normal operation of various computer tasks. In large-scale computer systems, task scheduling and monitoring have become critical links to ensure stable system operation. Traditional task monitoring systems mainly monitor task status and issue alerts based on fixed thresholds and rules. This monitoring approach is often inefficient in handling complex and changing task operation environments.
[0003] As business complexity continues to increase, task operational characteristics are becoming highly diverse and dynamic. Relying solely on manual experience to monitor tasks is no longer sufficient. In recent years, intelligent monitoring methods based on data analysis and machine learning have gradually emerged. By mining and analyzing historical operational data and building task operational feature models, they can more accurately identify task anomalies and provide solutions.
[0004] Existing technologies in the field of task monitoring have the following major shortcomings: First, traditional monitoring methods often use static threshold judgments, lack in-depth analysis of the historical operating patterns of tasks, and are unable to dynamically adjust monitoring strategies based on the characteristics of different tasks, resulting in high false alarm rates and frequent missed alarms. Second, existing monitoring systems lack automated intelligent diagnostic mechanisms after discovering an anomaly, and mostly rely on the experience of operations and maintenance personnel to diagnose and handle faults, which is inefficient and easily affected by human factors. Third, existing systems generally lack mechanisms for accumulating and reusing exception handling experience, and are unable to learn from historical exception handling and continuously optimize exception identification and handling capabilities, resulting in the repeated occurrence of the same type of anomalies and inconsistent handling results. Summary of the Invention
[0005] The embodiments of the present invention provide an intelligent monitoring method and system based on task execution history information, which can solve the problems in the prior art.
[0006] A first aspect of an embodiment of the present invention provides an intelligent monitoring method based on task execution history information, comprising:
[0007] Collecting the target task's operating data, storing the operating data in a historical information database, and constructing a task operation profile based on the historical information database;
[0008] Based on the task operation profile, the target task is monitored in real time to obtain real-time operation data, the real-time operation data is converted according to the characteristic dimensions of the task operation profile, and the task health score is calculated in real time;
[0009] When it is detected that the task health score is lower than the preset health threshold, the intelligent diagnosis mechanism is triggered, and the pattern matching algorithm is used to identify the current abnormality type. The corresponding abnormality handling solution is matched according to the current abnormality type, and a targeted handling strategy is generated;
[0010] Executing the processing strategy, continuously monitoring the changing trend of the task health score, and terminating the execution of the processing strategy when the task health score recovers to above the preset health threshold;
[0011] The abnormal scene characteristics, task health score change curve, execution process of the processing strategy and processing effect are recorded as new abnormality records and stored in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
[0012] Converting the real-time operation data according to the characteristic dimensions of the task operation profile and calculating the task health score in real time includes:
[0013] Extracting a feature dimension template from the task operation profile, the feature dimension template including a time dimension benchmark value, a resource dimension benchmark value, and a performance dimension benchmark value; comparing the real-time operation data with the benchmark values of each dimension in the feature dimension template, and calculating the deviation value of the real-time operation data relative to the benchmark values of each dimension;
[0014] Based on a preset scoring rule, the deviation value is mapped to a preset scoring interval to obtain a score for each dimension, where the upper limit of the preset scoring interval corresponds to a no-deviation state, and the lower limit of the preset scoring interval corresponds to a maximum deviation state;
[0015] Based on the real-time fluctuations of each dimension, the weight coefficient of the deviation value is dynamically calculated. The weight coefficient changes nonlinearly with the degree of deviation. The scores of each dimension and the corresponding weight coefficient are weighted to obtain a weighted dimension score. The weighted dimension scores are combined and calculated. When a significant deviation is detected in the score of any dimension, the influence of the dimension in the final score is magnified to generate a task health score.
[0016] When it is detected that the task health score is lower than the preset health threshold, the intelligent diagnosis mechanism is triggered, and the pattern matching algorithm is used to identify the current abnormality type. According to the current abnormality type, the corresponding abnormality handling solution is matched, and a targeted handling strategy is generated, including:
[0017] Obtain multi-dimensional monitoring indicator data of the target task, calculate the health mean and health standard deviation based on the historical health data of the target task, and periodically correct the health mean and health standard deviation in combination with the task execution cycle and the periodic fluctuation coefficient to obtain a dynamic health threshold;
[0018] When the task health score is lower than the dynamic health threshold, constructing a time series feature sequence based on the multi-dimensional monitoring indicator data;
[0019] Obtaining a historical anomaly feature vector from an anomaly pattern library, the historical anomaly feature vector having a corresponding anomaly type identifier, using a pattern matching algorithm to calculate a similarity score between the temporal feature sequence and the historical anomaly feature vector, and determining a current anomaly type based on the similarity score; using the pattern matching algorithm to calculate a matching degree between the temporal feature sequence and a historical processing solution, and generating an emergency processing strategy set based on the matching degree and the current anomaly type;
[0020] Calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm;
[0021] The emergency treatment strategy set and the root cause treatment strategy set are combined to form a candidate treatment solution set, the candidate treatment solution set is prioritized, and the treatment solution with the highest priority is selected as the targeted treatment strategy.
[0022] Calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm include:
[0023] Calculating the time series correlation of each monitoring indicator in the time series feature sequence, wherein the time series correlation is obtained by calculating the standardized covariance of the monitoring indicator and its historical value under different time delays, wherein the standardized covariance is calculated based on the mean and standard deviation of the current moment and the historical moment;
[0024] Calculating a root cause posterior probability distribution based on the time series correlation and a preset root cause prior probability, wherein the root cause posterior probability distribution represents the distribution probability of different root cause types; calculating the information entropy and sensitivity of each monitoring indicator in the time series feature sequence, calculating the feature importance based on the information entropy and the sensitivity, and performing time series smoothing on the feature importance to obtain a time-varying feature weight;
[0025] Calculating the strength of the mapping relationship between the root cause type and the processing strategy based on the time-varying feature weight; generating a candidate strategy set based on the root cause posterior probability distribution and the mapping relationship strength, and calculating the degree of resource competition between the strategies in the candidate strategy set;
[0026] A strategy combination score is calculated based on the root cause posterior probability distribution, the resource competition degree, and the overlap between strategies; and according to the strategy combination score, the root cause posterior probability distribution is optimized and screened under the condition that a resource budget constraint and a conflict threshold constraint are satisfied.
[0027] Executing the processing strategy, continuously monitoring the changing trend of the task health score, and terminating the execution of the processing strategy when the task health score recovers to above the preset health threshold includes:
[0028] Construct a time sliding window, construct a health time series based on the task health score, calculate the mean and variance of the health time series, and use the mean and variance to represent the overall health level of the task; calculate the health difference between adjacent moments in the health time series, and calculate the short-term change trend based on the health difference. The short-term change trend is used to predict the development direction of the task health;
[0029] Recording the task health score after executing the processing strategy, combining the task health score with the short-term change trend to obtain a strategy execution effect evaluation index; constructing a recovery progress curve based on the strategy execution effect evaluation index, and continuing to execute the processing strategy when the slope of the recovery progress curve is positive;
[0030] Continuously comparing the task health score with a preset health threshold, and when the task health score exceeds the preset health threshold, starting a stability check process; calculating the health fluctuation range within the time sliding window in the stability check process, and confirming that the task health has returned to stability when the health fluctuation range is lower than the preset fluctuation threshold;
[0031] When the task health returns to stability and the task health score is continuously higher than the preset health threshold, the execution process of the processing strategy is terminated.
[0032] Combining the task health score with the short-term change trend to obtain a strategy execution effect evaluation index; constructing a recovery progress curve based on the strategy execution effect evaluation index includes:
[0033] Setting a combined weight coefficient for the task health score and the short-term change trend, and weighting the task health score and the short-term change trend according to the combined weight coefficient to obtain an initial execution effect evaluation index;
[0034] Calculate a time decay factor based on the duration of policy execution, multiply the time decay factor by the initial execution effect evaluation index to obtain a time-series weighted evaluation index; calculate the difference between the current task health score and the task health score at the initial moment of policy execution, and use the ratio of the difference to the preset target improvement as the recovery progress factor;
[0035] The time series weighted evaluation index is combined with the recovery progress factor to obtain a strategy execution effect evaluation index, and a recovery progress curve is constructed based on the strategy execution effect evaluation index. The recovery progress curve uses an exponential form to describe the change pattern of the strategy execution effect over time.
[0036] A second aspect of an embodiment of the present invention provides an intelligent monitoring system based on task execution history information, including:
[0037] The first unit is configured to collect operation data of a target task, store the operation data in a historical information database, and construct a task operation profile based on the historical information database;
[0038] The second unit is configured to monitor the target task in real time based on the task operation profile to obtain real-time operation data, convert the real-time operation data according to the characteristic dimensions of the task operation profile, and calculate the task health score in real time;
[0039] The third unit is configured to trigger an intelligent diagnosis mechanism when detecting that the task health score is lower than a preset health threshold, use a pattern matching algorithm to identify the current abnormality type, match the corresponding abnormality handling solution according to the current abnormality type, and generate a targeted handling strategy;
[0040] A fourth unit is configured to execute the processing strategy, continuously monitor the changing trend of the task health score, and terminate the execution of the processing strategy when the task health score recovers to above the preset health threshold;
[0041] The fifth unit is used to store the abnormal scene characteristics, the task health score change curve, the execution process and the processing effect of the processing strategy as a new abnormal record in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
[0042] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0043] processor;
[0044] a memory for storing processor-executable instructions;
[0045] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0046] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0047] The beneficial effects of this application are as follows:
[0048] By constructing a task operation profile based on historical operation data, the present invention realizes all-round real-time monitoring of the target task, can quickly identify abnormal conditions and accurately calculate the health of the task, effectively reducing the workload of manual monitoring and improving the system operation and maintenance efficiency.
[0049] The present invention adopts pattern matching algorithm to automatically identify the exception type and match the corresponding processing scheme, forming a targeted processing strategy, realizing the automated processing of exception situations, shortening the problem solving time, reducing the impact of system exceptions on business, and improving the stability and availability of the system.
[0050] The present invention feeds back information such as scene characteristics, health changes, processing strategies and effects during the exception handling process to the historical information database, forming a closed-loop optimization mechanism, enabling the intelligent diagnosis system to continuously learn and evolve, continuously improving the accuracy and processing efficiency of exception diagnosis, reducing system maintenance costs, and extending the system life cycle. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Schematic diagram of the process of an intelligent monitoring method based on task execution history information according to an embodiment of the present invention;
[0052] Figure 2 This is a bar chart comparing the scores of various dimensions of task health in different test scenarios according to an embodiment of the present invention;
[0053] Figure 3 This is a flowchart of an embodiment of the present invention for calculating root cause probability and processing strategy based on time series characteristics;
[0054] Figure 4 Schematic diagram of the intelligent monitoring process based on task execution history information according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0056] The technical solution of the present invention is described in detail below with reference to specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0057] Figure 1 FIG. 1 is a flow chart of an intelligent monitoring method based on task operation history information according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0058] Collecting the target task's operating data, storing the operating data in a historical information database, and constructing a task operation profile based on the historical information database;
[0059] Based on the task operation profile, the target task is monitored in real time to obtain real-time operation data, the real-time operation data is converted according to the characteristic dimensions of the task operation profile, and the task health score is calculated in real time;
[0060] When it is detected that the task health score is lower than the preset health threshold, the intelligent diagnosis mechanism is triggered, and the pattern matching algorithm is used to identify the current abnormality type. The corresponding abnormality handling solution is matched according to the current abnormality type, and a targeted handling strategy is generated;
[0061] Executing the processing strategy, continuously monitoring the changing trend of the task health score, and terminating the execution of the processing strategy when the task health score recovers to above the preset health threshold;
[0062] The abnormal scene characteristics, task health score change curve, execution process of the processing strategy and processing effect are recorded as new abnormality records and stored in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
[0063] In an optional embodiment, converting the real-time operation data according to the characteristic dimensions of the task operation profile and calculating the task health score in real time includes:
[0064] Extracting a feature dimension template from the task operation profile, the feature dimension template including a time dimension benchmark value, a resource dimension benchmark value, and a performance dimension benchmark value; comparing the real-time operation data with the benchmark values of each dimension in the feature dimension template, and calculating the deviation value of the real-time operation data relative to the benchmark values of each dimension;
[0065] Based on a preset scoring rule, the deviation value is mapped to a preset scoring interval to obtain a score for each dimension, where the upper limit of the preset scoring interval corresponds to a no-deviation state, and the lower limit of the preset scoring interval corresponds to a maximum deviation state;
[0066] Based on the real-time fluctuations of each dimension, the weight coefficient of the deviation value is dynamically calculated. The weight coefficient changes nonlinearly with the degree of deviation. The scores of each dimension and the corresponding weight coefficient are weighted to obtain a weighted dimension score. The weighted dimension scores are combined and calculated. When a significant deviation is detected in the score of any dimension, the influence of the dimension in the final score is magnified to generate a task health score.
[0067] Obtain the characteristic dimension template from the task operation profile. This template includes baseline values for time, resources, and performance. Time baseline values may include time metrics such as average task execution time and data processing latency; resource baseline values may include resource consumption metrics such as CPU utilization, memory usage, and storage space usage; and performance baseline values include performance-related metrics such as throughput, error rate, and response time. These baseline values are derived through statistical analysis of historical task operation data and represent the expected performance of the task under normal operating conditions.
[0068] Real-time task execution data is collected, including real-time metrics such as execution duration, CPU utilization, and memory usage at the current point in time. The system converts the collected real-time execution data into a format corresponding to the metrics in the feature dimension template. For example, raw CPU usage data is converted from time series data to percentages, and memory usage is uniformly converted from bytes to megabytes, ensuring consistency and comparability of data formats.
[0069] Compare the converted real-time running data with the baseline value in the feature dimension template to calculate the deviation. For example, if the baseline value for CPU usage is 40% and the current actual value is 60%, the relative deviation is (60-40) / 40 = 50%, indicating that resource usage exceeds the baseline value by 50%. The system uses different deviation calculation methods for different metrics. For metrics with lower expected values (such as execution time and error rate), the deviation is (actual value - baseline value) / baseline value. For metrics with expected values within a certain range (such as memory usage), the deviation is |actual value - baseline value| / baseline value. For metrics with higher expected values (such as throughput), the deviation is (baseline value - actual value) / baseline value.
[0070] According to the preset scoring rules, the system maps the deviation value to the preset scoring range. The scoring range is set from 0 to 100 points, where 100 points corresponds to a zero-deviation state (deviation value of 0) and 0 points corresponds to a maximum-deviation state. The mapping uses a nonlinear function. When the deviation is small, the score decreases slowly; when the deviation exceeds a certain threshold, the score decreases rapidly. Specifically, the mapping function can be designed as follows: if the deviation value is less than 10%, the score is between 90-100 points; if the deviation value is between 10%-30%, the score is between 70-90 points; if the deviation value is between 30%-60%, the score is between 40-70 points; if the deviation value is between 60%-100%, the score is between 10-40 points; if the deviation value exceeds 100%, the score is between 0-10 points.
[0071] Based on the real-time fluctuations of each dimension, a weight coefficient for the deviation value is dynamically calculated. The weight coefficient changes nonlinearly with the degree of deviation: the larger the deviation, the higher the weight coefficient, thus reflecting the importance of the abnormal indicator in the final score. Weight calculation uses an exponential function mechanism, and when the deviation value exceeds the preset threshold, the weight coefficient increases exponentially. For example, for CPU usage, the initial weight is 0.3. When the deviation exceeds 50%, the weight is adjusted to 0.3 × (1 + 0.5) = 0.45; when the deviation exceeds 100%, the weight is adjusted to 0.3 × (1 + 1) = 0.6.
[0072] The weighted dimension scores are calculated by adding the scores of each dimension to the corresponding weight coefficients. Assume that the time dimension score is 85 with a weight of 0.35, the resource dimension score is 70 with a weight of 0.45, and the performance dimension score is 90 with a weight of 0.2. The weighted calculation result is: 85 × 0.35 + 70 × 0.45 + 90 × 0.2 = 79.75 points.
[0073] The weighted dimension scores are combined to generate the final task health score. This combined calculation not only considers the weighted average but also pays special attention to any dimensions with significant deviations. If any dimension score deviates significantly (below a preset threshold, such as 50 points), the system amplifies that dimension's influence in the final score. Specifically, the system sets an "abnormality penalty factor"; if any dimension score falls below the threshold, the final score is multiplied by this penalty factor. For example, if the resource dimension score is severely low at 40 points, below the 50-point threshold, the penalty factor is set to 0.8, resulting in a final health score of 79.75 × 0.8 = 63.8 points.
[0074] The calculated task health score is used to classify and address the task's operational status. Scores are typically categorized into five levels: 90-100 indicates a "healthy" status, requiring no intervention; 80-90 indicates a "good" status, with regular inspections recommended; 65-80 indicates a "fair" status, requiring attention for potential issues; 50-65 indicates an "abnormal" status, requiring optimization and adjustment; and 0-50 indicates a "critical" status, requiring immediate intervention or task restart.
[0075] In practice, the system compares health scores against preset thresholds and automatically triggers alerts when the score falls below them. For example, when the health score of a batch processing task suddenly dropped from a normal 85 to 62, the system detected an abnormal memory usage in the resource dimension (the score was only 45 points). It immediately issued an alert and provided a detailed dimension analysis report, indicating that memory usage had seriously deviated by 95% of the baseline value. Based on this information, operations personnel quickly located and resolved the memory leak, and the task health score subsequently returned to normal.
[0076] Through this implementation, the system can comprehensively assess the health of computing tasks in real time, promptly identify potential issues, and improve task performance and system stability. This multi-dimensional, dynamically weighted health scoring mechanism more accurately reflects the actual operation of tasks than traditional single-metric monitoring methods, providing strong support for intelligent operation and maintenance decision-making.
[0077] Figure 2This is a bar chart comparing the scores of each dimension of task health under different test scenarios of the embodiment of the present invention. The figure shows the performance evaluation results of the system under four different operating scenarios. The time dimension score uses a time series feature extraction algorithm, the resource dimension score is based on a dynamic resource utilization monitoring algorithm, and the performance dimension score uses a comprehensive service quality evaluation algorithm. In the normal operating scenario, the time dimension score reaches 95.8%, the resource dimension score is 93.2%, and the performance dimension score reaches a maximum of 97.6%. As the complexity of the scenario increases, the scores of each dimension show a decreasing trend, falling to 87.5%, 85.7%, and 82.4% respectively in the light load scenario. When the system enters the high load scenario, the scores of the three dimensions further drop to 72.3%, 68.4%, and 65.9%. In the abnormal fluctuation scenario, the scores of each dimension of the system reach the lowest point, at 58.6%, 52.9%, and 47.3% respectively. The scoring trend shows that as the complexity of the scenario increases, the performance dimension based on service quality assessment decreases the most, indicating that the overall system performance is most sensitive to scenario changes. The time dimension score based on timing analysis is relatively stable, indicating that the system has good robustness in time response. This multi-dimensional scoring system comprehensively reflects the system's adaptability and performance in various scenarios through a combination of different technical algorithms, providing an accurate quantitative basis for system optimization.
[0078] In an optional embodiment, when it is detected that the task health score is lower than a preset health threshold, the intelligent diagnosis mechanism is triggered, a pattern matching algorithm is used to identify the current abnormality type, and a corresponding abnormality handling solution is matched according to the current abnormality type to generate a targeted handling strategy including:
[0079] Obtain multi-dimensional monitoring indicator data of the target task, calculate the health mean and health standard deviation based on the historical health data of the target task, and periodically correct the health mean and health standard deviation in combination with the task execution cycle and the periodic fluctuation coefficient to obtain a dynamic health threshold;
[0080] When the task health score is lower than the dynamic health threshold, constructing a time series feature sequence based on the multi-dimensional monitoring indicator data;
[0081] Obtaining a historical anomaly feature vector from an anomaly pattern library, the historical anomaly feature vector having a corresponding anomaly type identifier, using a pattern matching algorithm to calculate a similarity score between the temporal feature sequence and the historical anomaly feature vector, and determining a current anomaly type based on the similarity score; using the pattern matching algorithm to calculate a matching degree between the temporal feature sequence and a historical processing solution, and generating an emergency processing strategy set based on the matching degree and the current anomaly type;
[0082] Calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm;
[0083] The emergency treatment strategy set and the root cause treatment strategy set are combined to form a candidate treatment solution set, the candidate treatment solution set is prioritized, and the treatment solution with the highest priority is selected as the targeted treatment strategy.
[0084] The task monitoring platform collects multi-dimensional monitoring indicator data for the target task, including CPU usage, memory usage, response time, and error rate. Taking the data processing task as an example, the system collects health data for the task over the past 30 days, including multi-dimensional indicator values at hourly sampling points. The system calculates the average health score for the data processing task to be 85 points, with a standard deviation of 5 points. Considering the task's significant daily execution cycle (increased processing volume in the early morning hours), the system introduces a periodic fluctuation coefficient of 1.2 to adjust the health indicator.
[0085] Divide 24 hours into six time periods, and set different dynamic health thresholds for different time periods: the dynamic health threshold during non-peak hours is 75 points (mean minus two standard deviations), and the dynamic health threshold during peak hours (0:00 to 3:00 in the morning) is 72 points (mean minus two standard deviations combined with the periodic fluctuation coefficient).
[0086] When monitoring detected that the data processing task's health score was 70 at 2:00 AM on a certain day, below the dynamic health threshold of 72 for the corresponding time period, the system immediately constructed a time series feature sequence. This sequence included data on changes in multiple monitoring indicators in the 10 minutes before the task triggered the anomaly, including key indicators such as a sharp increase in CPU utilization from 45% to 92%, memory usage from 60% to 85%, error rate from 0.1% to 2.5%, and a task backlog increase from 100 to 850.
[0087] The system retrieves historical anomaly feature vectors from the anomaly pattern library, which stores feature vectors of various previously occurring anomalies and their corresponding anomaly type identifiers. Using the DTW (Dynamic Time Warping) algorithm, the system calculates the similarity scores between the constructed time series feature sequence and each of the historical anomaly feature vectors in the library. The results show that this sequence has the highest similarity score of 0.89 with the historical anomaly feature vector for the "Insufficient processing resources during peak data input" type. Therefore, the system determines the current anomaly type as "Insufficient processing resources during peak data input."
[0088] After determining the anomaly type, the system uses the same pattern matching algorithm to calculate the degree of match between the current time series feature sequence and historical response plans. For the current anomaly type, the system retrieves three historical emergency response strategies: dynamically expanding processing resources, reducing data input frequency, and temporarily increasing task priority. The calculated matching degrees are 0.85, 0.72, and 0.63, respectively. Based on these matching degrees, the system generates an emergency response strategy set that includes these three strategies, preserving their priority order.
[0089] The system calculated the root cause probability distribution based on the collected time series feature sequences. By analyzing indicator trends and correlations, the system identified the following root causes and their probabilities: a sudden surge in data from the upstream data source (0.75), insufficient processing unit configuration (0.65), and inefficient processing due to insufficient code optimization (0.45). Based on the identified anomaly type, "Insufficient processing resources during peak data input," the system generated three root cause resolution strategies: optimizing the upstream data source's throttling mechanism, increasing the basic configuration of processing units, and restructuring the data processing algorithm to improve efficiency.
[0090] These three root cause treatment strategies were optimized and screened using a pattern matching algorithm, analyzing each strategy's historical effectiveness, implementation costs, and expected returns. The overall score for "Optimizing the upstream data source throttling mechanism" was 0.85, "Increasing the basic configuration of processing units" was 0.78, and "Restructuring the data processing algorithm" was 0.62.
[0091] After combining the emergency processing strategy set and the root cause processing strategy set, the system forms the following candidate processing solution set: dynamically expand processing resources (emergency, matching degree 0.85), reduce data input frequency (emergency, matching degree 0.72), temporarily increase task priority (emergency, matching degree 0.63), optimize upstream data source flow control mechanism (root cause, score 0.85), increase the basic configuration of processing units (root cause, score 0.78), and reconstruct data processing algorithms (root cause, score 0.62).
[0092] These candidate solutions were prioritized, taking into account their compatibility / score, implementation difficulty, response time, and expected results. The results were: dynamically expanding processing resources (overall priority 0.92), optimizing upstream data source throttling mechanisms (overall priority 0.87), reducing data input frequency (overall priority 0.80), increasing basic processing unit configuration (overall priority 0.75), temporarily increasing task priorities (overall priority 0.65), and refactoring data processing algorithms (overall priority 0.55).
[0093] The system ultimately selected the highest-priority "dynamic expansion of processing resources" as a targeted processing strategy. It automatically issued instructions to the resource management system to temporarily add two processing units to the original four and adjust task scheduling parameters. This reduced CPU utilization to 65% within five minutes, restored the task backlog to normal levels within 15 minutes, and restored the health score to 83. The system also submitted the "optimization of upstream data source throttling mechanisms" as a follow-up root cause solution to the task management platform, which the development team is following up on.
[0094] This solution achieves rapid response and intelligent processing of task anomalies, which not only solves urgent problems but also provides long-term suggestions to ensure the long-term stable and efficient operation of the system.
[0095] In an optional embodiment, calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm includes:
[0096] Calculating the time series correlation of each monitoring indicator in the time series feature sequence, wherein the time series correlation is obtained by calculating the standardized covariance of the monitoring indicator and its historical value under different time delays, wherein the standardized covariance is calculated based on the mean and standard deviation of the current moment and the historical moment;
[0097] Calculating a root cause posterior probability distribution based on the time series correlation and a preset root cause prior probability, wherein the root cause posterior probability distribution represents the distribution probability of different root cause types; calculating the information entropy and sensitivity of each monitoring indicator in the time series feature sequence, calculating the feature importance based on the information entropy and the sensitivity, and performing time series smoothing on the feature importance to obtain a time-varying feature weight;
[0098] Calculating the strength of the mapping relationship between the root cause type and the processing strategy based on the time-varying feature weight; generating a candidate strategy set based on the root cause posterior probability distribution and the mapping relationship strength, and calculating the degree of resource competition between the strategies in the candidate strategy set;
[0099] A strategy combination score is calculated based on the root cause posterior probability distribution, the resource competition degree, and the overlap between strategies; and according to the strategy combination score, the root cause posterior probability distribution is optimized and screened under the condition that a resource budget constraint and a conflict threshold constraint are satisfied.
[0100] like Figure 3 As shown, the method further includes:
[0101] Calculate the time series correlation of each monitoring indicator in the time series feature sequence. The time series correlation is obtained by calculating the standardized covariance of the monitoring indicator and its historical value under different time delays. Specifically, for the monitoring indicator x(t), the time series correlation under time delay d can be measured by the correlation between the indicator value x(t) at the current time t and the indicator value x(td) at the historical time td. When calculating the standardized covariance, first calculate the mean μ of the monitoring indicator in the current time window current and standard deviation σ current , and the mean μ of the monitoring indicators in the historical time window history and standard deviation σ history For example, when the current mean of system CPU usage is 75% with a standard deviation of 8%, and the historical mean is 60% with a standard deviation of 5%, a time series correlation calculation reveals a high correlation of 0.85 between the CPU usage and the historical value from 5 minutes ago, indicating that this anomaly has a clear time series correlation pattern.
[0102] The root cause posterior probability distribution is calculated based on time series correlation and preset root cause prior probabilities. The root cause prior probabilities are derived from historical anomaly statistics. For example, in server systems, the prior probability of a memory leak is 0.15, the prior probability of network congestion is 0.25, and the prior probability of database lock contention is 0.2. Combined with the calculated time series correlation, the root cause posterior probability is calculated using probability update rules. For example, when a continuous increase in memory usage is observed with a correlation of 0.92 with historical values, the posterior probability of a memory leak increases from the prior 0.15 to 0.68, indicating that the memory leak is the root cause of the current anomaly.
[0103] Calculate the information entropy and sensitivity of each monitoring indicator in the time series feature sequence. Information entropy measures the uncertainty of the monitoring indicator and is obtained by discretizing the indicator value and calculating the probability distribution. Sensitivity measures the indicator's responsiveness to system anomalies and is determined by calculating the rate of change in the indicator value before and after the anomaly. Taking the network traffic indicator as an example, the number of packets per second during normal periods is 1000 ± 200, which rises to 5000 ± 500 when an anomaly occurs. The calculated sensitivity value is 4.0, indicating that the indicator is highly sensitive to the current anomaly.
[0104] Feature importance is calculated based on information entropy and sensitivity. Feature importance comprehensively considers the information content of a metric and its sensitivity to anomalies. Features with higher importance are more likely to reflect system anomalies. For example, when the database connection pool is exhausted, the feature importance of the number of active connections is 0.95, while the feature importance of disk I / O is only 0.12, indicating that the number of connections is more important for identifying such anomalies.
[0105] Time-varying feature weights are obtained by smoothing the feature importance over time. Using a sliding window weighted average, feature weights are smoothly varied over time, reducing the impact of noise. For example, using a 5-minute window to smooth the CPU usage metric, the original weight sequence [0.82, 0.79, 0.85, 0.75, 0.88] is smoothed to 0.818, improving weight stability.
[0106] The strength of the mapping relationship between root cause types and handling strategies is calculated based on time-varying feature weights. For each root cause type, the system pre-sets multiple handling strategies, and uses feature weights to assess the suitability of each strategy for that specific root cause. For example, for performance degradation caused by CPU-intensive tasks, the mapping strength of the "add service instances" strategy is calculated to be 0.92, the "adjust thread pool size" strategy is 0.75, and the "limit requests" strategy is 0.83.
[0107] A set of candidate strategies is generated based on the root cause posterior probability distribution and the strength of the mapping relationship. Using the mapping strength as a weighting factor for strategy selection, a preliminary set of candidate strategies is generated based on the root cause probability distribution. For the aforementioned CPU-intensive task, the generated candidate strategies include adding four service instances, adjusting the maximum number of threads in the thread pool to 200, and setting the per-second rate limit to 500.
[0108] Calculate the degree of resource contention between the policies in the candidate policy set. This degree of resource contention is determined by evaluating the contention for system resources when multiple policies are executed simultaneously. For example, when the "Add service instances" and "Expand cache capacity" policies are executed simultaneously, the calculated degree of contention for memory resources is 0.77, indicating a high risk of resource conflict.
[0109] The strategy combination score is calculated based on the posterior probability distribution of the root cause, the degree of resource contention, and the overlap between policies. Policy overlap measures the degree of overlap in the scope of multiple policies and is calculated by comparing the similarity of policy target indicators. The combination score comprehensively considers the policy's targeting of the root cause, resource contention, and functional overlap. For example, the strategy combination of "Adjusting the Database Connection Pool Size" and "Optimizing SQL Queries" scored 0.89, indicating a strong combination.
[0110] Based on the strategy combination scores, the posterior probability distribution of the root cause is optimized and screened, subject to the resource budget and conflict threshold constraints. The resource budget constraint ensures that the total resource consumption of the selected strategy combination does not exceed the system's available resources, and the conflict threshold constraint ensures that the degree of resource contention between strategies is below a safety threshold. For example, given 8GB of available memory, the "Optimize Memory Caching Mechanism" and "Adjust Garbage Collection Parameters" strategies are selected, with a total memory requirement of 5GB and a resource contention level of 0.3, meeting the system constraints.
[0111] In an optional embodiment, executing the processing strategy and continuously monitoring the changing trend of the task health score, and terminating the execution of the processing strategy when the task health score recovers to above the preset health threshold, includes:
[0112] Construct a time sliding window, construct a health time series based on the task health score, calculate the mean and variance of the health time series, and use the mean and variance to represent the overall health level of the task; calculate the health difference between adjacent moments in the health time series, and calculate the short-term change trend based on the health difference. The short-term change trend is used to predict the development direction of the task health;
[0113] Recording the task health score after executing the processing strategy, combining the task health score with the short-term change trend to obtain a strategy execution effect evaluation index; constructing a recovery progress curve based on the strategy execution effect evaluation index, and continuing to execute the processing strategy when the slope of the recovery progress curve is positive;
[0114] Continuously comparing the task health score with a preset health threshold, and when the task health score exceeds the preset health threshold, starting a stability check process; calculating the health fluctuation range within the time sliding window in the stability check process, and confirming that the task health has returned to stability when the health fluctuation range is lower than the preset fluctuation threshold;
[0115] When the task health returns to stability and the task health score is continuously higher than the preset health threshold, the execution process of the processing strategy is terminated.
[0116] Construct a time sliding window to monitor changes in task health status in real time. The time sliding window can be set to 10 minutes, with task health score data collected every 30 seconds. For example, a data processing task running on a cloud computing platform can record health scores calculated from metrics such as CPU utilization, memory usage, and response time every 30 seconds, saving data from the 20 sampling points in the last 10 minutes. Based on this sampling data, a health time series sequence can be constructed, such as [85, 83, 80, 75, 71, 68, 65, 62, 58, 55, 53, 50, 48, 45, 43, 42, 40, 41, 43, 45], indicating that the health score gradually decreases from 85 to 40 and then begins to recover.
[0117] For the constructed health time series, calculate its mean and variance. For the above series, the mean is 57.6 and the variance is 255.44. The mean reflects the overall health of the task within the time window, while the variance indicates the stability of the health state. A higher mean and lower variance indicate an overall healthy and stable task.
[0118] To understand health trends, the system calculates the difference between health values at adjacent moments. For the above sequence, the resulting difference sequence is [-2, -3, -5, -4, -3, -3, -3, -4, -3, -2, -3, -2, -3, -2, -1, -2, 1, 2, 2]. By weighting these differences, with more recent differences receiving higher weights, we can derive an indicator of short-term trends. For example, taking the weighted average of the last five differences, with weights of [0.1, 0.15, 0.2, 0.25, 0.3], we calculate a short-term trend of 0.85. Positive values indicate an upward trend in health, while negative values indicate a downward trend. The magnitude of the value reflects the strength of the trend.
[0119] When the system detects that the task health is below the preset health threshold (such as 60), it initiates a processing strategy, such as resource expansion, load balancing, or error retry. After executing the processing strategy, the system continuously records the task health score, forming a new health sequence, such as [45, 48, 52, 56, 59, 62, 65, 67, 69, 71]. At the same time, the system combines the current task health score with the short-term change trend indicator to construct a policy execution effect evaluation index. The specific method is to multiply the current health score by (1 + short-term change trend). For example, when the health is 62 and the short-term change trend is 0.05, the policy execution effect evaluation index is 62×(1+0.05)=65.1.
[0120] Based on the time series of policy execution effectiveness evaluation indicators, the system constructs a recovery progress curve. This curve is formed by calculating the recovery progress indicators at each time point and connecting them, for example, [47.25, 50.88, 54.86, 58.8, 62.36, 65.1, 67.6, 69.35, 71.07, 72.84]. The system calculates the slope of the recovery progress curve. A positive slope indicates that the treatment strategy is effective and the system continues to execute it. If the slope is negative for three consecutive sampling points, the treatment strategy needs to be adjusted or replaced.
[0121] When a task's health score exceeds a preset health threshold (e.g., 60), the system doesn't immediately terminate the processing strategy. Instead, it initiates a stability check. During this stability check, the system calculates the health fluctuation range within the time sliding window—the difference between the maximum and minimum values. For example, the health sequence in the most recent window is [59, 62, 65, 67, 69, 71, 72, 73, 74, 75], with a fluctuation range of 75 - 59 = 16. If the fluctuation range falls below the preset fluctuation threshold (e.g., 20), the task's health is considered stable.
[0122] The system also verifies how long the task's health score remains above the preset health threshold. Typically, the health score must remain above the threshold for 5 consecutive minutes (10 sampling points). If the health score fluctuates during this period but remains above the threshold, and the fluctuation range is within an acceptable range, the system confirms that the task has returned to a healthy state.
[0123] When the system confirms that the task's health has stabilized and remains above the preset health threshold, the execution of the treatment strategy is terminated. Before terminating the treatment strategy, the system records all recovery process data, including the initial health score, treatment strategy type, execution time, and health changes during the recovery process, for subsequent analysis and optimization of the treatment strategy.
[0124] Through this technical solution, the system can accurately monitor changes in task health, promptly detect anomalies and initiate remediation strategies, and then terminate the remediation strategies appropriately once the task has recovered, effectively improving the system's self-healing capabilities and resource utilization efficiency. This solution is suitable for all types of computing tasks that require continuous monitoring and dynamic adjustments, and is particularly well-suited for scenarios such as cloud computing and big data processing.
[0125] In an optional embodiment, the task health score is combined with the short-term change trend to obtain a strategy execution effect evaluation index; and constructing a recovery progress curve based on the strategy execution effect evaluation index includes:
[0126] Setting a combined weight coefficient for the task health score and the short-term change trend, and weighting the task health score and the short-term change trend according to the combined weight coefficient to obtain an initial execution effect evaluation index;
[0127] Calculate a time decay factor based on the duration of policy execution, multiply the time decay factor by the initial execution effect evaluation index to obtain a time-series weighted evaluation index; calculate the difference between the current task health score and the task health score at the initial moment of policy execution, and use the ratio of the difference to the preset target improvement as the recovery progress factor;
[0128] The time series weighted evaluation index is combined with the recovery progress factor to obtain a strategy execution effect evaluation index, and a recovery progress curve is constructed based on the strategy execution effect evaluation index. The recovery progress curve uses an exponential form to describe the change pattern of the strategy execution effect over time.
[0129] The process of combining the task health score and short-term trend to derive the policy execution effectiveness evaluation metric involves multiple technical steps. The system first sets a combined weighting coefficient to adjust the importance of the task health score and short-term trend in the final evaluation metric. In practice, different weightings can be set based on specific business scenarios. For example, the weight of the task health score could be set to 0.7, while the weight of the short-term trend could be set to 0.3. This weighting configuration reflects the business judgment that the current state is more important than the trend.
[0130] The initial execution performance evaluation index is calculated by multiplying the task health score by the corresponding weight and the sum of the short-term trend multiplied by the corresponding weight. For example, if the task health score is 85 and the short-term trend is positive at 2.5 points per day, the initial execution performance evaluation index can be calculated as 85 × 0.7 + 2.5 × 0.3 = 60.25 points.
[0131] To account for the impact of policy execution time, the system calculates a time decay factor. The longer the policy execution period, the smaller the decay factor, indicating that the policy execution effect decreases marginally over time. The time decay factor can be implemented using an inverse proportional function. That is, when the policy execution time is T days, the time decay factor can be set to 1 / (1 + 0.05 × T). Here, 0.05 is the decay rate parameter, which can be adjusted based on the business scenario. For example, when the policy has been executed for 10 days, the time decay factor is 1 / (1 + 0.05 × 10) = 0.67. The system multiplies this decay factor by the initial execution effect evaluation index described above to obtain the time-weighted evaluation index. In the above example, the calculated result is 60.25 × 0.67 = 40.37 points.
[0132] The system also calculates a recovery progress factor to assess the extent to which policy execution has improved task health. First, the difference between the current task health score and the task health score at the start of policy execution is calculated. The recovery progress factor is then calculated as the ratio of this difference to the target improvement. For example, if the task health score was 70 before policy execution and is currently 85, with a target improvement of 20, the recovery progress factor is (85-70) / 20 = 0.75, indicating that 75% of the target improvement has been achieved.
[0133] To derive the final policy execution effectiveness evaluation index, the system combines the time-weighted evaluation index with the recovery progress factor. This combination can be a weighted average, for example, setting the time-weighted evaluation index to a weight of 0.6 and the recovery progress factor to a weight of 0.4. Using these weights, the final policy execution effectiveness evaluation index is 40.37 × 0.6 + 0.75 × 0.4 = 24.52. This index comprehensively reflects multiple factors, including current health status, change trends, time decay, and recovery progress.
[0134] Based on the strategy execution effectiveness evaluation indicators calculated above, the system constructs a recovery progress curve. This curve uses an exponential form to describe how strategy execution effectiveness changes over time. The exponential form effectively simulates the characteristic of strategy execution, where initial effectiveness is significant and marginal benefits diminish in later stages. The recovery progress curve is expressed as follows: the recovery progress value at time point t is equal to the initial recovery value plus the target improvement multiplied by the negative exponent of one minus the natural base (the exponent is negative K multiplied by time t). K is a parameter that represents the recovery rate and can be set based on historical data fitting or business experience.
[0135] This example illustrates the implementation of a series of recovery strategies after an e-commerce platform's order processing system experienced a failure. The initial task health score was 60, and the expected goal was to increase the health score to 95, with a target improvement of 35 points. On the fifth day of strategy execution, the health score was 78, with a short-term trend of 2.8 points per day. By setting the combined weighting coefficients to 0.7 for the health score and 0.3 for the trend, the calculated initial execution performance evaluation index was 78 × 0.7 + 2.8 × 0.3 = 55.44 points.
[0136] The time decay factor is 1 / (1 + 0.05 × 5) = 0.8, resulting in a time-weighted evaluation index of 55.44 × 0.8 = 44.35. The recovery progress factor is (78 - 60) / 35 = 0.51, indicating that 51% of the target improvement has been achieved. Combining the time-weighted evaluation index and the recovery progress factor with a weighting of 0.6:0.4 yields a policy execution effectiveness evaluation index of 44.35 × 0.6 + 0.51 × 0.4 = 26.82.
[0137] As time passes, the system updates and calculates the policy execution effectiveness evaluation index daily, forming an exponential recovery progress curve. On the 10th day, the health score rises to 85 points, with the short-term trend slowing to 1.5 points per day. Using the same method, the policy execution effectiveness evaluation index is 22.35. On the 15th day, the health score is 90 points, with a short-term trend of 0.8 points per day, resulting in a policy execution effectiveness evaluation index of 18.41. The curve shows a gradual decrease in the evaluation index over time, indicating that while health is still improving, the marginal benefit is decreasing, which is consistent with the execution pattern of most recovery strategies.
[0138] The recovery progress curve constructed in this way can intuitively demonstrate the effectiveness of policy execution, help decision makers determine whether the current policy needs to be adjusted or enhanced, and provide data support for system fault recovery management.
[0139] like Figure 4 As shown, the method further includes:
[0140] 1. Baseline rule definition
[0141] Baseline rule definition, including basic information: baseline name, baseline cycle type, responsible person, guarantee tasks, committed completion time for each cycle, priority increase, warning margin duration, alarm type: type (error, slowdown), slowdown threshold (percentage), notification configuration: notification method, number of notifications, do not disturb period, and notification recipients.
[0142] 2. Average running time of pre-computation tasks in each scheduling cycle:
[0143] Before the daily baseline is converted to an instance, the average running time of the task in each scheduling period is pre-calculated based on the historical instance data of the task execution in the past N days.
[0144] 3. Converting Baseline Rules to Instances
[0145] The baseline instance for the next day is pre-transferred on a scheduled basis, converting the baseline rules into corresponding instances. All instances of the assurance task and all its upstream tasks on the current business date are counted, generating cycle records for each scheduling cycle of the task instances corresponding to the baseline instance. The current cycle is initially recorded as the first cycle in the baseline rule instance. All paths of the task instances in this cycle in the DAG workflow are calculated during the baseline rule instance cycle.
[0146] 4. Monitor and prioritize baseline instances
[0147] After the baseline rules for the day are successfully transferred to the instance, a monitoring job is created for each baseline instance starting at midnight the next day.
[0148] To monitor a baseline instance, first obtain the current cycle of the baseline instance (i.e., the earliest unfinished cycle record). If all cycles are completed, the baseline instance is completed, and the monitoring job for the baseline instance is removed. Otherwise, the current cycle status needs to be monitored cyclically and stored in the database.
[0149] Baseline instance cycle monitoring, first obtain all task instances of the current cycle: if all leaf node task instances are frozen, the completion status of the current cycle is set to all completed and the baseline status is set to other; if all leaf nodes are successful, the completion status is set to completed and the baseline status is judged based on the actual completion time (when the actual completion time exceeds the breaking time, it is broken, when the actual completion time exceeds the warning time but does not exceed the breaking time, there is no warning, and if it does not exceed the warning time, it is safe); otherwise, it is considered to be running in the current cycle, and all current task instances in the current cycle (that is, running or failed task instances) are analyzed, and compared based on the running time, current time, and historical average running time of the current task instance: 1. The baseline status of the cycle with a broken task instance is set to broken, and the task instance that broke the line the earliest is recorded as a critical task instance. All task instance paths in this cycle contain critical task instances. The first task instance that has the latest expected completion time is the critical path, and the last task instance in the critical path is the latest expected task instance. Then the task instances in the critical path are raised to an unpreset priority. 2. The baseline status of a task instance with warnings but no broken-line task instances is set to warning. The earliest warning task instance is recorded as the critical task instance. The path of all task instances in this cycle contains a critical task instance and has the latest expected completion time. The last task instance in the critical path is the latest task instance. Then the task instances in the critical path are raised to an unpreset priority. 3. The baseline status of a cycle without expected task instances is set to safe. The earliest task instance that may break the line is recorded as the critical task instance. The path of all task instances in this cycle contains a critical task instance and has the latest expected completion time. The last task instance in the critical path is the latest task instance.
[0150] 5. Abnormal event monitoring
[0151] 5.1 Abnormal failure event monitoring
[0152] Monitor the running status of the instance of the guarantee task and all its upstream tasks on the day. When monitoring an abnormally failed task instance, a newly discovered failure event is generated. For newly discovered failure events, continue to monitor and alert until the event is marked as recovered (the status is restored to success) or ignored, and then stop monitoring. If it is manually marked as being processed, no monitoring and alerting will be performed during the processing period. After the processing period, the task instance will continue to be monitored and alerted until the alarm limit is reached. No alarm notification will be issued during the do not disturb period.
[0153] 5.2 Abnormal slowdown event monitoring
[0154] Monitor all running and successful instances of the assurance task and all its upstream tasks on the same day. Failed instances are monitored through abnormal failure events and are therefore no longer monitored here. For running instances, their current runtime (current time - start time) is compared with their historical average runtime for the period. If the duration exceeds the slowdown threshold (percentage), they are considered to have slowed down. For successfully running instances, their actual runtime (completion time - start time) is compared with their historical average runtime for the period. If the duration exceeds the slowdown threshold (percentage), they are considered to have slowed down. For slowed task instances, new slowdown events are generated. For newly discovered failure events, monitoring and alarming are continued until the event is marked as ignored, at which point monitoring is stopped. If the event is manually marked as being processed, monitoring and alarming are stopped during the processing period. After the processing period, the task instance will continue to be monitored and alarmed until the alarm limit is reached. No alarm notifications are issued during the do not disturb period.
[0155] 6. Baseline and event alerts
[0156] 6.1 Baseline Alarm
[0157] Scan the baseline instance cycle and analyze all instances that meet the baseline status of warning or breach, have enabled alarm notifications, are not within the processing time, have not exceeded the number of alarms, have not exceeded the latest alarm time, and are not in the do not disturb period. If they meet the conditions, send alarm notifications and record the cumulative number of alarm notifications.
[0158] The baseline alarm notification includes: baseline name, baseline status (warning, breach), margin, key task instance, key task, and occurrence time;
[0159] 6.2 Event Alarm
[0160] Scan baseline abnormal events and analyze all events that meet the conditions of newly discovered or being processed, alarm notification enabled, not within the processing time, not exceeding the number of alarms, not exceeding the latest alarm time, and not within the do not disturb period. If they meet the conditions, send alarm notifications and record the cumulative number of alarm notifications;
[0161] The event alarm notification includes: event type (error / slowdown), event task instance, event task, occurrence time, and slowdown percentage (only for slowdown events).
[0162] A second aspect of an embodiment of the present invention provides an intelligent monitoring system based on task execution history information, including:
[0163] The first unit is configured to collect operation data of a target task, store the operation data in a historical information database, and construct a task operation profile based on the historical information database;
[0164] The second unit is configured to monitor the target task in real time based on the task operation profile to obtain real-time operation data, convert the real-time operation data according to the characteristic dimensions of the task operation profile, and calculate the task health score in real time;
[0165] The third unit is configured to trigger an intelligent diagnosis mechanism when detecting that the task health score is lower than a preset health threshold, use a pattern matching algorithm to identify the current abnormality type, match the corresponding abnormality handling solution according to the current abnormality type, and generate a targeted handling strategy;
[0166] A fourth unit is configured to execute the processing strategy, continuously monitor the changing trend of the task health score, and terminate the execution of the processing strategy when the task health score recovers to above the preset health threshold;
[0167] The fifth unit is used to store the abnormal scene characteristics, the task health score change curve, the execution process and the processing effect of the processing strategy as a new abnormal record in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
[0168] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0169] processor;
[0170] a memory for storing processor-executable instructions;
[0171] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0172] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0173] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent monitoring method based on task operation history information, characterized in that: include: Collecting the target task's operating data, storing the operating data in a historical information database, and constructing a task operation profile based on the historical information database; Based on the task operation profile, the target task is monitored in real time to obtain real-time operation data, the real-time operation data is converted according to the characteristic dimensions of the task operation profile, and the task health score is calculated in real time; When it is detected that the task health score is lower than the preset health threshold, the intelligent diagnosis mechanism is triggered, and the pattern matching algorithm is used to identify the current abnormality type. The corresponding abnormality handling solution is matched according to the current abnormality type, and a targeted handling strategy is generated; Executing the processing strategy, continuously monitoring the changing trend of the task health score, and terminating the execution of the processing strategy when the task health score recovers to above the preset health threshold, including: Construct a time sliding window, construct a health time series based on the task health score, calculate the mean and variance of the health time series, and use the mean and variance to represent the overall health level of the task; calculate the health difference between adjacent moments in the health time series, and calculate the short-term change trend based on the health difference. The short-term change trend is used to predict the development direction of the task health; Recording the task health score after executing the processing strategy, combining the task health score with the short-term change trend to obtain a strategy execution effect evaluation index; constructing a recovery progress curve based on the strategy execution effect evaluation index, and continuing to execute the processing strategy when the slope of the recovery progress curve is positive; Continuously comparing the task health score with a preset health threshold, and when the task health score exceeds the preset health threshold, starting a stability check process; calculating the health fluctuation range within the time sliding window in the stability check process, and confirming that the task health has returned to stability when the health fluctuation range is lower than the preset fluctuation threshold; When the task health returns to stability and the task health score is continuously higher than the preset health threshold, terminating the execution process of the processing strategy; The abnormal scene characteristics, task health score change curve, execution process of the processing strategy and processing effect are recorded as new abnormality records and stored in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
2. The method according to claim 1, characterized in that Converting the real-time operation data according to the characteristic dimensions of the task operation profile and calculating the task health score in real time includes: Extracting a feature dimension template from the task operation profile, the feature dimension template including a time dimension benchmark value, a resource dimension benchmark value, and a performance dimension benchmark value; comparing the real-time operation data with the benchmark values of each dimension in the feature dimension template, and calculating the deviation value of the real-time operation data relative to the benchmark values of each dimension; Based on a preset scoring rule, the deviation value is mapped to a preset scoring interval to obtain a score for each dimension, where the upper limit of the preset scoring interval corresponds to a no-deviation state, and the lower limit of the preset scoring interval corresponds to a maximum deviation state; Based on the real-time fluctuations of each dimension, the weight coefficient of the deviation value is dynamically calculated. The weight coefficient changes nonlinearly with the degree of deviation. The scores of each dimension and the corresponding weight coefficient are weighted to obtain a weighted dimension score. The weighted dimension scores are combined and calculated. When a significant deviation is detected in the score of any dimension, the influence of the dimension in the final score is magnified to generate a task health score.
3. The method according to claim 1, characterized in that When it is detected that the task health score is lower than the preset health threshold, the intelligent diagnosis mechanism is triggered, and the pattern matching algorithm is used to identify the current abnormality type. According to the current abnormality type, the corresponding abnormality handling solution is matched, and a targeted handling strategy is generated, including: Obtain multi-dimensional monitoring indicator data of the target task, calculate the health mean and health standard deviation based on the historical health data of the target task, and periodically correct the health mean and health standard deviation in combination with the task execution cycle and the periodic fluctuation coefficient to obtain a dynamic health threshold; When the task health score is lower than the dynamic health threshold, constructing a time series feature sequence based on the multi-dimensional monitoring indicator data; Obtaining a historical anomaly feature vector from an anomaly pattern library, the historical anomaly feature vector having a corresponding anomaly type identifier, using a pattern matching algorithm to calculate a similarity score between the temporal feature sequence and the historical anomaly feature vector, and determining a current anomaly type based on the similarity score; using the pattern matching algorithm to calculate a matching degree between the temporal feature sequence and a historical processing solution, and generating an emergency processing strategy set based on the matching degree and the current anomaly type; Calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm; The emergency treatment strategy set and the root cause treatment strategy set are combined to form a candidate treatment solution set, the candidate treatment solution set is prioritized, and the treatment solution with the highest priority is selected as the targeted treatment strategy.
4. The method according to claim 3, characterized in that Calculating a root cause probability distribution based on the time series feature sequence, generating a root cause processing strategy set according to the root cause probability distribution and the current abnormality type, and optimizing and screening the root cause processing strategy set using the pattern matching algorithm include: Calculating the time series correlation of each monitoring indicator in the time series feature sequence, wherein the time series correlation is obtained by calculating the standardized covariance of the monitoring indicator and its historical value under different time delays, wherein the standardized covariance is calculated based on the mean and standard deviation of the current moment and the historical moment; Calculating a root cause posterior probability distribution based on the time series correlation and a preset root cause prior probability, wherein the root cause posterior probability distribution represents the distribution probability of different root cause types; calculating the information entropy and sensitivity of each monitoring indicator in the time series feature sequence, calculating the feature importance based on the information entropy and the sensitivity, and performing time series smoothing on the feature importance to obtain a time-varying feature weight; Calculating the strength of the mapping relationship between the root cause type and the processing strategy based on the time-varying feature weight; generating a candidate strategy set based on the root cause posterior probability distribution and the mapping relationship strength, and calculating the degree of resource competition between the strategies in the candidate strategy set; A strategy combination score is calculated based on the root cause posterior probability distribution, the resource competition degree, and the overlap between strategies; and according to the strategy combination score, the root cause posterior probability distribution is optimized and screened under the condition that a resource budget constraint and a conflict threshold constraint are satisfied.
5. The method according to claim 1, wherein Combining the task health score with the short-term change trend to obtain a strategy execution effect evaluation index; Constructing a recovery progress curve based on the strategy execution effect evaluation indicators includes: Setting a combined weight coefficient for the task health score and the short-term change trend, and weighting the task health score and the short-term change trend according to the combined weight coefficient to obtain an initial execution effect evaluation index; Calculate a time decay factor based on the duration of policy execution, multiply the time decay factor by the initial execution effect evaluation index to obtain a time-series weighted evaluation index; calculate the difference between the current task health score and the task health score at the initial moment of policy execution, and use the ratio of the difference to the preset target improvement as the recovery progress factor; The time series weighted evaluation index is combined with the recovery progress factor to obtain a strategy execution effect evaluation index, and a recovery progress curve is constructed based on the strategy execution effect evaluation index. The recovery progress curve uses an exponential form to describe the change pattern of the strategy execution effect over time.
6. An intelligent monitoring system based on task operation history information, used to implement the method according to any one of claims 1 to 5, characterized in that: include: The first unit is configured to collect operation data of a target task, store the operation data in a historical information database, and construct a task operation profile based on the historical information database; The second unit is configured to monitor the target task in real time based on the task operation profile to obtain real-time operation data, convert the real-time operation data according to the characteristic dimensions of the task operation profile, and calculate the task health score in real time; The third unit is configured to trigger an intelligent diagnosis mechanism when detecting that the task health score is lower than a preset health threshold, use a pattern matching algorithm to identify the current abnormality type, match the corresponding abnormality handling solution according to the current abnormality type, and generate a targeted handling strategy; A fourth unit is configured to execute the processing strategy, continuously monitor the changing trend of the task health score, and terminate the execution of the processing strategy when the task health score recovers to above the preset health threshold; The fifth unit is used to store the abnormal scene characteristics, the task health score change curve, the execution process and the processing effect of the processing strategy as a new abnormal record in the historical information database to optimize the diagnostic accuracy and processing effect of the intelligent diagnosis mechanism.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Intelligent task alarm rule self-learning method and system based on support priority
CN119441832A
Python monitoring task resource use method and system
CN119512883A
Server health state diagnosis method based on GAT-LP algorithm
CN120086105A