Intelligent operation and maintenance method of streaming computing tasks based on real-time monitoring
By monitoring multiple performance indicators of the streaming computing task in real time, identifying high load, attention and abnormal nodes, and dynamically adjusting the standard consistency and deviation index, the problems of excessive system load and response time delay in the streaming computing task of e-commerce platforms are solved, and efficient early warning and precise intervention are achieved.
Patent Information
- Application Number
- CN202510628947.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When facing the high concurrency and real-time streaming computing tasks of e-commerce platforms, the existing operation and maintenance system cannot adapt to business needs, resulting in excessive system load and delayed response time, and inflexible resource scheduling, and unable to respond quickly to system load changes.
By monitoring multiple performance indicators in the streaming calculation task in real time, including order reception rate, success rate, task delay, data processing rate, resource utilization rate and checkpoint completion time, the high load, attention and abnormal nodes are identified using layer-by-layer progressive judgment logic, dynamically adjust the preset standard consistency and deviation index, and issue operation and maintenance alerts.
It improves the accuracy and sensitivity of abnormal identification, can respond to node load evolution trends in real time, improves the ability to identify potential performance bottlenecks and fault precursors, achieves efficient early warning and precise intervention, and solves the problems of excessive system load and delayed response time.
Smart Images

Figure CN120179506B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent operation and maintenance technology, and in particular to an intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring. Background Art
[0002] The rapid growth of the e-commerce industry, coupled with the surge in user numbers and diversification of business needs, has forced e-commerce platforms to face extremely complex computing tasks and massive data processing. These tasks often involve real-time data processing and complex operations and maintenance management. Especially in high-concurrency scenarios like promotions and flash sales, system stability and efficient resource scheduling are key factors in determining the platform's success. However, existing operations and maintenance systems often struggle to provide sufficient real-time and flexibility for massive streaming computing tasks, leading to problems such as excessive system load and delayed response times.
[0003] The patent document with publication number CN117389837A discloses a method and system for operating operation and maintenance tasks for supercomputing and intelligent computing environments. The method includes: S1: loading the operation and maintenance tasks in the original supercomputing and intelligent computing environments; S2: aligning and synchronizing the operation and maintenance tasks in the original supercomputing and intelligent computing environments to generate operation and maintenance tasks related to the joint computing environment; S3: obtaining resource requirements in the operation and maintenance tasks and performing resource identification; S4: processing the data obtained from resource identification; S5: referencing the processed data and completing authentication adaptation; S6: performing dynamic unified orchestration after authentication adaptation; S7: performing task operations according to the orchestration results.
[0004] It can be seen that the operation method for operation and maintenance tasks for supercomputing and intelligent computing environments has the following problems: First, the method focuses on the unified operation and maintenance of supercomputing and intelligent computing environments, but when processing the high concurrency and real-time streaming computing tasks of e-commerce platforms, it cannot adapt to the business needs of e-commerce platforms. Especially when faced with sudden traffic and large-scale promotional activities, the synchronization and orchestration process of operation and maintenance tasks in the method causes delays and cannot quickly respond to changes in system load; it is not flexible enough in resource scheduling and dynamic adjustment, and cannot be accurately and efficiently allocated according to the real-time load and task requirements in e-commerce scenarios, which ultimately affects the overall performance of the system and user experience. Summary of the Invention
[0005] To this end, the present invention provides an intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring, which is used to overcome the problems of excessive system load and delayed response time in the existing technology due to overly independent resource scheduling and low synchronization of operation and maintenance tasks through real-time data analysis and dynamic threshold adjustment.
[0006] To achieve the above objectives, the present invention provides an intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring, comprising:
[0007] Obtain in real time the order receiving rate, order success rate, task execution delay value, data processing rate, state accumulation degree, resource utilization rate, and checkpoint completion time of each monitoring node in the streaming computing task of the order platform;
[0008] Determine a number of high-load nodes based on the order receiving rate, order success rate, and consistency with preset standards;
[0009] Determine a number of focus nodes according to the data processing rate, the state accumulation degree, and a preset standard deviation index of each of the high-load nodes;
[0010] Determine a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes;
[0011] Determine a number of fault risk nodes according to the task execution delay values of any two of the abnormal nodes;
[0012] Adjusting the preset standard consistency according to the number of the fault risk nodes within the preset adjustment period to form an adjusted standard consistency, or adjusting the preset standard deviation index to form an adjusted standard deviation index;
[0013] An operation and maintenance alarm is issued for all the fault risk nodes that are re-determined based on the adjustment standard consistency or the adjustment standard deviation index.
[0014] Furthermore, the process of determining a number of high-load nodes according to the order receiving rate, the order success rate, and the preset consistency threshold includes:
[0015] Calculate the standard deviation of the order receiving rate of all the orders within the preset load determination time period to obtain a receiving rate fluctuation value;
[0016] Calculate the standard deviation of the success rate of all the orders within the preset load determination time to obtain a success rate fluctuation value;
[0017] Normalizing the receiving rate fluctuation value to obtain a receiving rate normalized value, and normalizing the success rate fluctuation value to obtain a success rate normalized value;
[0018] Calculating a correlation coefficient between the normalized value of the receiving rate and the normalized value of the success rate to obtain a fluctuation consistency;
[0019] A number of high-load nodes are determined according to the fluctuation consistency and the preset consistency threshold.
[0020] Furthermore, the process of determining a number of high-load nodes according to the fluctuation consistency and the preset consistency threshold includes:
[0021] When the fluctuation consistency is greater than the preset consistency threshold, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
[0022] Furthermore, the process of determining a number of high-load nodes according to the fluctuation consistency and the preset consistency threshold further includes:
[0023] When the fluctuation consistency is less than or equal to the preset consistency threshold, the standard deviation of the order receiving rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary receiving fluctuation values, and the standard deviation of the order success rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary success fluctuation values;
[0024] Drawing a curve showing how all the temporary reception fluctuation values change over time within the preset load determination time period to obtain a reception change curve;
[0025] Plotting a curve showing changes of all the temporary success fluctuation values over time within the preset load determination time period to obtain a success change curve;
[0026] Calculating the cosine similarity between the reception change curve and the success change curve to obtain the fluctuation change similarity;
[0027] When the fluctuation change similarity is greater than a preset standard change similarity, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
[0028] Furthermore, the process of determining a number of focus nodes according to the data processing rate, the state accumulation degree, and the preset standard deviation index of each of the high-load nodes includes:
[0029] Calculating the standard deviation of all data processing rates within a predetermined period of time to obtain a processing rate fluctuation value;
[0030] Calculating the standard deviation of the accumulation degree of all the states within the preset attention determination time length to obtain an accumulation fluctuation value;
[0031] Calculating a relative deviation between the processing rate fluctuation value and a preset processing rate fluctuation threshold to obtain a processing fluctuation deviation;
[0032] Calculating a relative deviation between the accumulation fluctuation value and a preset accumulation fluctuation threshold value to obtain an accumulation fluctuation deviation;
[0033] A number of focus nodes are determined according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index.
[0034] Furthermore, the process of determining a number of focus nodes according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index includes:
[0035] Calculating the product of the processing fluctuation deviation and a preset processing fluctuation weight to obtain a processing fluctuation factor;
[0036] Calculating the product of the stacking fluctuation deviation and a preset stacking fluctuation weight to obtain a stacking fluctuation factor;
[0037] Calculating the sum of the processing fluctuation factor and the stacking fluctuation factor to obtain a fluctuation deviation index;
[0038] When the fluctuation deviation index is greater than the preset standard deviation index, the high-load node is determined to be the focus node, so as to determine a plurality of focus nodes.
[0039] Furthermore, the process of determining a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes includes:
[0040] Calculate the average of all resource usage rates within the preset abnormality judgment time to obtain the average resource usage rate;
[0041] When the average resource usage rate is greater than a preset resource threshold, marking the focus node as a resource node;
[0042] Calculate the average completion time of all the checkpoints within the preset abnormality judgment time to obtain the average completion time;
[0043] When the average completion time is greater than a preset completion threshold, marking the node of interest as a delay node;
[0044] When the concerned node is marked as the resource node and the delay node at the same time, determining the corresponding concerned node as the abnormal node to determine a number of abnormal nodes;
[0045] When the focus node is only marked as the resource node, calculating the standard deviation of all resource usage rates to obtain a resource fluctuation value;
[0046] When the resource fluctuation value is greater than a preset resource fluctuation threshold, determining the focus node as the abnormal node to determine a number of abnormal nodes;
[0047] When the focus node is only marked as the delay node, calculating the standard deviation of the completion time of all the checkpoints to obtain a completion time fluctuation value;
[0048] When the resource fluctuation value is greater than a preset completion time fluctuation threshold, the focus node is determined to be the abnormal node, so as to determine a number of abnormal nodes.
[0049] Furthermore, the process of determining a number of fault risk nodes according to the task execution delay values of any two of the abnormal nodes includes:
[0050] Constructing a knowledge graph with each of the abnormal nodes as a graph node;
[0051] Calculating the standard deviation of the task execution delay value of each abnormal node within a preset risk determination time period to obtain a plurality of delay fluctuation values;
[0052] Calculating the correlation coefficient of the delay fluctuation values of any two abnormal nodes to obtain a delay correlation;
[0053] When the delay correlation is greater than a preset standard correlation, connecting the corresponding two abnormal nodes to obtain a plurality of graph edges;
[0054] When the number of graph edges corresponding to each abnormal node is greater than a preset standard number of edges, the abnormal node is determined to be a fault risk node.
[0055] Furthermore, the preset standard consistency is adjusted according to the number of the fault risk nodes within the preset adjustment period to form an adjusted standard consistency, or the preset standard deviation index is adjusted. The process of forming the adjusted standard deviation index includes:
[0056] Obtaining the number of the fault risk nodes at each moment within the preset adjustment time to obtain a number of risk node numbers;
[0057] When the number of risk nodes is greater than a preset node number threshold, calculating the inverse of the standard deviation of all risk nodes to obtain the node number concentration;
[0058] The preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted to form the adjusted standard deviation index.
[0059] Furthermore, the preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted. The process of forming the adjusted standard deviation index includes:
[0060] When the node number concentration is less than the minimum value of the preset standard concentration range, reducing the preset standard consistency according to the relative deviation between the minimum value of the preset standard concentration range and the node number concentration and a preset adjustment coefficient to form the adjusted standard consistency;
[0061] When the node quantity concentration is greater than the maximum value of the preset standard concentration range, the preset standard deviation index is increased according to the relative deviation between the node quantity concentration and the maximum value of the preset standard concentration range and the adjustment coefficient to form the adjusted standard deviation index.
[0062] Compared with the prior art, the beneficial effect of the present invention is that, by converting the intrinsic relationship between different performance indicators into a step-by-step judgment logic, the accuracy and sensitivity of anomaly identification are enhanced: in the early stage, the order receiving rate and success rate are used to identify nodes that may have traffic pressure, and the data processing rate and state accumulation are combined to further distinguish whether it is a real lack of processing capacity or just a short-term fluctuation. Subsequently, the resource utilization rate and checkpoint completion time are used to cross-verify whether the node load has affected the stable operation of the system, and finally the fault propagation risk is identified based on the change in task execution delay. According to the fluctuation of the fault risk node within a certain period of time, the consistency threshold and fluctuation sensitivity required for judgment can be dynamically adjusted, so that the operation and maintenance strategy can respond to the node load evolution trend in real time, and improve the ability to identify potential performance bottlenecks and fault precursors, thereby achieving efficient early warning and precise intervention of the running status of streaming tasks without increasing additional resource overhead, effectively solving the problems of excessive system load and delayed response time due to overly independent resource scheduling and low synchronization of operation and maintenance tasks.
[0063] Furthermore, by introducing a matching analysis of the fluctuation trends of the reception rate and success rate, we effectively avoid misjudgments caused by fluctuations in a single indicator and strengthen the ability to identify true sources of high-load stress. Normalization and correlation calculation not only improve the computational adaptability between indicators, but also can keenly capture the "synchronous rise or fall" of node states during changes in business pressure, thereby revealing hidden processing bottlenecks and providing a reliable structural foundation for subsequent hierarchical identification and dynamic adjustment.
[0064] Furthermore, by calculating the correlation between the order receiving rate and the normalized fluctuation trend of the order success rate, the dynamic connection between the node processing capacity and the task completion quality is indirectly reflected, and then a unified fluctuation consistency index is used for judgment, so that the identification of high-load nodes is based on the coordinated fluctuation of the two core business parameters, enhancing the pertinence and accuracy of the judgment.
[0065] Furthermore, when the initial correlation is not obvious, temporal fluctuation analysis and curve similarity measurement are introduced, and the historical fluctuation evolution trend at each moment in the time window is used as a supplementary criterion to further reveal the possible deep fluctuation synergy between reception and processing efficiency, providing a more temporal-sensitive and detailed judgment basis for the identification of high-load nodes, thereby improving the fine-grained perception ability and intelligent judgment level of the monitoring system.
[0066] Furthermore, by introducing the relative deviation between the fluctuation amplitude and the threshold, the dynamic quantification of the degree of indicator abnormality is strengthened, and on the basis of high load, the system is guided to focus on those nodes whose internal processing capacity and accumulation status have changed significantly and are unstable, thereby achieving early identification of potential problem nodes, effectively enhancing the overall operation and maintenance strategy's responsiveness to the evolution of load behavior and the control accuracy.
[0067] Furthermore, by combining processing fluctuation deviation and accumulation fluctuation deviation and assigning different weights to them respectively, the impact of processing rate and accumulation degree on system status can be flexibly highlighted, thereby forming a comprehensive fluctuation deviation index. By setting the weights, the importance of each factor can be adjusted according to the actual application situation, ensuring that nodes that require special attention can be accurately identified when the load is abnormal, thereby achieving precise resource scheduling and maintenance optimization.
[0068] Furthermore, by comprehensively analyzing fluctuations in resource utilization and checkpoint completion times, we can accurately identify abnormal nodes and dynamically adjust based on specific resource and latency performance to avoid misjudgments or missed detections. The average and fluctuation values of resource utilization and completion times provide a detailed quantitative basis for node health. Furthermore, through labeling and threshold setting, we provide a systematic and hierarchical identification mechanism for anomaly detection, effectively improving monitoring accuracy and responsiveness.
[0069] Furthermore, by comprehensively considering the correlation and connectivity of latency fluctuations between nodes, we can effectively identify nodes with potential failure risks. Calculating the standard deviation and correlation of latency fluctuations helps capture the coordinated fluctuation relationships between nodes. This detection of fluctuation patterns not only improves early warning capabilities for potential failures, but also ensures that abnormal behavior across multiple nodes can be promptly linked and managed in a clustered manner.
[0070] Furthermore, by monitoring the number and concentration of nodes at risk of failure, fluctuations in system load and abnormal conditions can be accurately determined. When the number of risky nodes is high and their concentration is high, adjustments to the consistency or standard deviation index allow for timely response to potential high risks and optimize system load distribution and task execution efficiency. This process ensures dynamic adjustments based on actual conditions when faced with failure risks, enhancing the flexibility and intelligence of operations and maintenance management.
[0071] Furthermore, by precisely controlling the preset standard consistency and standard deviation index, dynamic adjustments can be made based on changes in node concentration. By adopting appropriate adjustment strategies at different concentration levels, the system's ability to cope with load fluctuations is optimized. When concentration is low, reducing consistency can enhance sensitivity; while when concentration is high, increasing the deviation index can strengthen the system's fault tolerance, effectively preventing misjudgments and ensuring the system maintains high efficiency in the face of fluctuating load environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is a flow chart of the intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring in this embodiment;
[0073] Figure 2 This is a logic decision diagram for determining high-load nodes in this embodiment;
[0074] Figure 3 This is a decision logic diagram for determining the focus node in this embodiment;
[0075] Figure 4 This is a decision logic diagram for determining fault risk nodes in this embodiment. DETAILED DESCRIPTION
[0076] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0077] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0078] See also Figure 1 As shown, it is a flow chart of the intelligent operation and maintenance method of streaming computing tasks based on real-time monitoring in this embodiment;
[0079] This embodiment provides an intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring, including:
[0080] Obtain in real time the order receiving rate, order success rate, task execution delay value, data processing rate, state accumulation degree, resource utilization rate, and checkpoint completion time of each monitoring node in the streaming computing task of the order platform;
[0081] Determine a number of high-load nodes based on the order receiving rate, order success rate, and consistency with preset standards;
[0082] Determine a number of focus nodes according to the data processing rate, the state accumulation degree, and a preset standard deviation index of each of the high-load nodes;
[0083] Determine a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes;
[0084] Determine a number of fault risk nodes according to the task execution delay values of any two of the abnormal nodes;
[0085] Adjusting the preset standard consistency according to the number of the fault risk nodes within the preset adjustment period to form an adjusted standard consistency, or adjusting the preset standard deviation index to form an adjusted standard deviation index;
[0086] An operation and maintenance alarm is issued for all the fault risk nodes that are re-determined based on the adjustment standard consistency or the adjustment standard deviation index.
[0087] The order receiving rate, order success rate, task execution delay, data processing rate, state accumulation, resource utilization, and checkpoint completion time of each monitoring node in the order platform's streaming computing tasks are collected in real time through Flink's native metrics system. This system supports custom indicator items and dynamic threshold judgment logic in the form of expressions. During task operation, PushGatewayReporter is configured to push all node metrics to Prometheus, achieving unified data collection and real-time monitoring.
[0088] All preset standard values in this embodiment support dynamic configuration in the form of expressions, such as based on the ratio of FlinkTaskManager JVM heap memory usage. When a threshold is triggered, an alarm is generated, including the triggering cause and recommended parameter adjustment plan.
[0089] The operation and maintenance alert in this embodiment is an alert service integrated with Prometheus, combined with communication interfaces such as SMS, email, or enterprise WeChat, to push task indicator abnormality information and optimization suggestions to task development and operation and maintenance personnel, realizing a closed-loop operation and maintenance response.
[0090] In this embodiment, based on the streaming computing task environment of the e-commerce platform, the following parameters are obtained in real time:
[0091] The order acceptance rate refers to the number of orders received per second in a streaming computing system. On e-commerce platforms, the order acceptance rate directly reflects the platform's current business traffic and is a key indicator for determining traffic peaks and system load.
[0092] The order success rate refers to the proportion of orders that are successfully completed during the order processing process (e.g., written to the database, payment verification completed, etc.). This metric is used to measure the overall processing stability and accuracy of the system.
[0093] The task execution latency value represents the time delay between order data entering the task node and completion of processing. On e-commerce platforms, low latency is essential for ensuring real-time responses (such as recommendations, price calculation, and inventory verification).
[0094] The data processing rate refers to the number of orders actually processed per second by a streaming task. It reflects the system's throughput and is an important metric compared to the order intake rate.
[0095] State accumulation measures the degree of accumulation of state data within a node (such as keyed state and window cache). In e-commerce stream processing, operations such as user behavior analysis and product click window statistics rely on state storage.
[0096] Resource utilization is a CPU usage indicator that reflects the actual resource usage of a node when executing tasks. Under high concurrency conditions, resource utilization helps operations and maintenance identify performance bottlenecks and determine whether resources need to be expanded.
[0097] Checkpoint completion time refers to the time it takes for a streaming task to complete its state snapshot (checkpoint). In e-commerce platforms, checkpoints ensure fault tolerance for tasks. If checkpoints take too long or fail frequently, recovery capabilities will be compromised, increasing the risk of task interruption.
[0098] The preset standard consistency is a standard value used to measure the consistency between the current collected data and the historical or standard data. It depends on the monitoring system's requirements for recognition accuracy and the fluctuation range of on-site collection conditions. It is usually set between 85% and 95%. In this embodiment, it is set to 90%, which can effectively identify the stability deviation of the collected data and improve the accuracy of the system's judgment.
[0099] The preset standard deviation index is a standard value that reflects whether the fluctuation degree of the current acquisition parameters is within a reasonable range. It is based on the mean square error level of historical data in the same type of scenario. The value range is usually between 0.05 and 0.2. In this embodiment, it is set to 0.1. It can promptly identify abnormal changes caused by external disturbances and provide a basis for judgment for subsequent adjustments to the system.
[0100] The preset adjustment time refers to the minimum response time required for the system to perform automatic adjustment when the consistency or deviation index does not meet the standard. It depends on the device response speed, the field data refresh frequency and the stability of the control module. It is generally set between 3 seconds and 30 seconds. In this embodiment, it is set to 10 seconds. While ensuring the adjustment effect, it can avoid system shock caused by frequent adjustments.
[0101] By collecting multiple key performance indicators (KPIs) from the order platform's streaming computing tasks in real time, such as order receipt rate, success rate, execution latency, and processing rate, we hierarchically screen out high-load nodes, nodes of concern, abnormal nodes, and nodes at risk of failure. During the node status identification process, we dynamically determine the node status based on the preset consistency and standard deviation index. Based on the number of nodes at risk of failure, we adaptively adjust the core judgment parameters within a preset adjustment period, enabling early warning of potential risks and intelligent operation and maintenance decision-making.
[0102] By converting the intrinsic relationship between different performance indicators into a step-by-step judgment logic, the accuracy and sensitivity of anomaly identification are enhanced: in the early stage, the order reception rate and success rate are used to identify nodes that may be under traffic pressure. The data processing rate and state accumulation degree are combined to further distinguish whether it is a real lack of processing capacity or just a short-term fluctuation. Subsequently, the resource utilization rate and checkpoint completion time are used to cross-verify whether the node load has affected the stable operation of the system. Finally, the fault propagation risk is identified based on the change in task execution delay. According to the fluctuation of the fault risk node within a certain period of time, the consistency threshold and fluctuation sensitivity required for judgment can be dynamically adjusted, so that the operation and maintenance strategy can respond to the node load evolution trend in real time, and improve the ability to identify potential performance bottlenecks and fault precursors. Therefore, without increasing additional resource overhead, efficient early warning and precise intervention of the running status of streaming tasks can be achieved, effectively solving the problems of excessive system load and delayed response time due to overly independent resource scheduling and low synchronization of operation and maintenance tasks.
[0103] Specifically, the process of determining a number of high-load nodes based on the order receiving rate, the order success rate, and a preset consistency threshold includes:
[0104] Calculate the standard deviation of the order receiving rate of all the orders within the preset load determination time period to obtain a receiving rate fluctuation value;
[0105] Calculate the standard deviation of the success rate of all the orders within the preset load determination time to obtain a success rate fluctuation value;
[0106] Normalizing the receiving rate fluctuation value to obtain a receiving rate normalized value, and normalizing the success rate fluctuation value to obtain a success rate normalized value;
[0107] Calculating a correlation coefficient between the normalized value of the receiving rate and the normalized value of the success rate to obtain a fluctuation consistency;
[0108] A number of high-load nodes are determined according to the fluctuation consistency and the preset consistency threshold.
[0109] The preset load determination duration is a time window used to measure fluctuations in order acceptance rates and order success rates. It depends on the real-time nature of business processing, the frequency of data fluctuations, and the system's tolerance for these fluctuations. It is typically set between 5 and 300 seconds to ensure both the capture of fluctuation trends and real-time analysis. In this embodiment, it is set to 60 seconds. This allows for the full extraction of the periodic characteristics of node load fluctuations while maintaining real-time processing, providing a stable and representative data foundation for subsequent fluctuation consistency calculations and the identification of high-load nodes.
[0110] By calculating the standard deviation of the order receiving rate and the order success rate within the preset load determination time, the degree of fluctuation of the business load of each node is quantified, and then the fluctuation values of the two are normalized so that indicators of different dimensions can be compared on the same scale; then the correlation coefficient between the two is calculated to extract the consistency of their fluctuation trends, that is, the fluctuation consistency; finally, the consistency is compared with the preset threshold, and the nodes whose fluctuations converge to high-intensity business changes are screened out as high-load nodes, laying the foundation for subsequent node screening.
[0111] By introducing a matching analysis of the fluctuation trends of the reception rate and success rate, we effectively avoid misjudgments caused by fluctuations in a single indicator and strengthen the ability to identify true sources of high-load stress. Normalization and correlation calculation not only improve the computational adaptability between indicators, but also can keenly capture node states that "rise or fall synchronously" with changes in business pressure, thereby revealing hidden processing bottlenecks and providing a reliable structural foundation for subsequent hierarchical identification and dynamic adjustment.
[0112] Please continue reading Figure 2 As shown, it is a logical decision diagram for determining a high-load node in this embodiment;
[0113] The process of determining a number of high-load nodes according to the fluctuation consistency and the preset consistency threshold includes:
[0114] When the fluctuation consistency is greater than the preset consistency threshold, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
[0115] By comparing the fluctuation consistency of each monitoring node with the preset consistency threshold, when the fluctuation consistency is greater than the threshold, it is determined that the order reception rate and order success rate of the node have highly synchronized fluctuations in a short period of time, thereby identifying it as a high-load node and serving as the starting point for subsequent operation and maintenance attention.
[0116] By calculating the correlation between the normalized fluctuation trend of the order receiving rate and the order success rate, the dynamic connection between the node processing capacity and the task completion quality is indirectly reflected, and then a unified fluctuation consistency index is used for judgment, so that the identification of high-load nodes is based on the coordinated fluctuation of the two core business parameters, which enhances the pertinence and accuracy of the judgment.
[0117] Specifically, the process of determining the plurality of high-load nodes according to the fluctuation consistency and the preset consistency threshold further includes:
[0118] When the fluctuation consistency is less than or equal to the preset consistency threshold, the standard deviation of the order receiving rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary receiving fluctuation values, and the standard deviation of the order success rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary success fluctuation values;
[0119] Drawing a curve showing how all the temporary reception fluctuation values change over time within the preset load determination time period to obtain a reception change curve;
[0120] Plotting a curve showing changes of all the temporary success fluctuation values over time within the preset load determination time period to obtain a success change curve;
[0121] Calculating the cosine similarity between the reception change curve and the success change curve to obtain the fluctuation change similarity;
[0122] When the fluctuation change similarity is greater than a preset standard change similarity, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
[0123] If the fluctuation consistency fails to exceed the preset consistency threshold, to further determine whether there is a potential high-load risk, the standard deviation of the order acceptance rate and order success rate from each starting point to the current moment within the preset load determination period is calculated. These fluctuation change curves are then generated, and the cosine similarity of the two curves is calculated to obtain the fluctuation change similarity. When this similarity exceeds the preset standard change similarity, it can be confirmed that the node's fluctuation trend during processing is highly synchronized, and the node is identified as a high-load node.
[0124] When the initial correlation is not obvious, temporal fluctuation analysis and curve similarity measurement are introduced, and the historical fluctuation evolution trend at each moment in the time window is used as a supplementary criterion to further reveal the possible deep fluctuation synergy between reception and processing efficiency, providing a more temporal-sensitive and detail-detailed judgment basis for the identification of high-load nodes, thereby improving the fine-grained perception capability and intelligent judgment level of the monitoring system.
[0125] Specifically, the process of determining a number of focus nodes according to the data processing rate, the state accumulation degree, and the preset standard deviation index of each high-load node includes:
[0126] Calculating the standard deviation of all data processing rates within a predetermined period of time to obtain a processing rate fluctuation value;
[0127] Calculating the standard deviation of the accumulation degree of all the states within the preset attention determination time length to obtain an accumulation fluctuation value;
[0128] Calculating a relative deviation between the processing rate fluctuation value and a preset processing rate fluctuation threshold to obtain a processing fluctuation deviation;
[0129] Calculating a relative deviation between the accumulation fluctuation value and a preset accumulation fluctuation threshold value to obtain an accumulation fluctuation deviation;
[0130] A number of focus nodes are determined according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index.
[0131] The preset attention determination time is used to determine the time range for analyzing node fluctuations. It depends on the response time requirements of the task, the data processing cycle, and the operation and maintenance requirements of the system. It is usually set between 40 seconds and 5 minutes to ensure that the changing trend of fluctuations can be captured. In this embodiment, the preset attention determination time is set to 3 minutes, which can accurately capture the fluctuation characteristics of high-load nodes without affecting the system response, and make timely operation and maintenance adjustments.
[0132] The preset processing rate fluctuation threshold is a standard for determining whether the data processing rate exceeds the normal fluctuation range. It depends on the system's processing capacity, historical processing rate, and fluctuations in task load. It is usually set between 10% and 50% to ensure that the system can sensitively detect abnormal processing rates. In this embodiment, it is set to 30%, which can effectively distinguish larger processing fluctuations and help identify nodes where bottlenecks may occur.
[0133] The preset accumulation fluctuation threshold is used to determine whether fluctuations in the accumulation state are abnormal. When setting this threshold, it is necessary to consider the system's processing capacity, the complexity of task execution, and the risk of task accumulation. It is typically set between 10% and 40% to effectively identify potential accumulation risks. In this embodiment, it is set to 20%, which can promptly identify abnormal changes in the accumulation state and effectively control the stability of data stream processing.
[0134] For identified high-load nodes, we further analyze the fluctuations of two key metrics, data processing rate and state accumulation, over a predetermined timeframe. We calculate their standard deviations and compare their relative deviations against their respective fluctuation thresholds to obtain the processing fluctuation deviation and accumulation fluctuation deviation. Combined with the preset standard deviation index, these deviations are comprehensively assessed to identify nodes of concern that may be experiencing resource bottlenecks or abnormal data processing trends.
[0135] By introducing the relative deviation between the fluctuation amplitude and the threshold, the dynamic quantification of the degree of indicator abnormality is strengthened. On the basis of high load, the system is guided to focus on those nodes whose internal processing capacity and accumulation status have changed significantly and are unstable, thereby achieving early identification of potential problem nodes and effectively enhancing the overall operation and maintenance strategy's responsiveness and control accuracy to the evolution of load behavior.
[0136] Please continue reading Figure 3 As shown, it is a decision logic diagram for determining the focus node in this embodiment;
[0137] The process of determining a number of focus nodes according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index includes:
[0138] Calculating the product of the processing fluctuation deviation and a preset processing fluctuation weight to obtain a processing fluctuation factor;
[0139] Calculating the product of the stacking fluctuation deviation and a preset stacking fluctuation weight to obtain a stacking fluctuation factor;
[0140] Calculating the sum of the processing fluctuation factor and the stacking fluctuation factor to obtain a fluctuation deviation index;
[0141] When the fluctuation deviation index is greater than the preset standard deviation index, the high-load node is determined to be the focus node, so as to determine a plurality of focus nodes.
[0142] The preset processing fluctuation weight is set according to the degree of impact of data processing rate fluctuations on system performance, which depends on the system's sensitivity and importance to the data processing rate. It is usually set between 0.2 and 0.5. In this embodiment, it is set to 0.3, which can ensure that under high load conditions, the fluctuation of data processing rate can fully reflect the potential performance problems of the system and provide an effective basis for subsequent adjustments.
[0143] The preset accumulation fluctuation weight is set based on the impact of state accumulation fluctuations on system health. Accumulation fluctuations typically directly impact task execution efficiency and responsiveness, so the weight should reflect the impact of accumulation changes on the overall system and is typically set between 0.5 and 0.8. In this example, a value of 0.7 is used to reasonably assess accumulation changes while ensuring their importance to system stability is not underestimated, thereby helping the system promptly identify risks that may arise from accumulation fluctuations.
[0144] By multiplying the processed and accumulated fluctuation deviations with their corresponding weights, we obtain the respective fluctuation factors. These fluctuation factors are then summed to form the fluctuation deviation index. By comparing this with the preset standard deviation index, if the fluctuation deviation index is greater than the preset standard deviation index, these high-load nodes are identified as nodes of concern for further monitoring and optimization.
[0145] By combining processing fluctuation deviation and accumulation fluctuation deviation and assigning different weights to each, we can flexibly highlight the impact of processing rate and accumulation degree on system status, thereby forming a comprehensive fluctuation deviation index. By setting weights, we can adjust the importance of each factor according to the actual application situation, ensuring that nodes that require special attention can be accurately identified when the load is abnormal, thereby achieving precise resource scheduling and maintenance optimization.
[0146] Specifically, the process of determining a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes includes:
[0147] Calculate the average of all resource usage rates within the preset abnormality judgment time to obtain the average resource usage rate;
[0148] When the average resource usage rate is greater than a preset resource threshold, marking the focus node as a resource node;
[0149] Calculate the average completion time of all the checkpoints within the preset abnormality judgment time to obtain the average completion time;
[0150] When the average completion time is greater than a preset completion threshold, marking the node of interest as a delay node;
[0151] When the concerned node is marked as the resource node and the delay node at the same time, determining the corresponding concerned node as the abnormal node to determine a number of abnormal nodes;
[0152] When the focus node is only marked as the resource node, calculating the standard deviation of all resource usage rates to obtain a resource fluctuation value;
[0153] When the resource fluctuation value is greater than a preset resource fluctuation threshold, determining the focus node as the abnormal node to determine a number of abnormal nodes;
[0154] When the focus node is only marked as the delay node, calculating the standard deviation of the completion time of all the checkpoints to obtain a completion time fluctuation value;
[0155] When the resource fluctuation value is greater than a preset completion time fluctuation threshold, the focus node is determined to be the abnormal node, so as to determine a number of abnormal nodes.
[0156] The preset abnormality judgment duration is used to determine the time window within which the average value of resource utilization and checkpoint completion time is calculated. It depends on the task processing cycle and its change frequency. It is usually set between 2 hours and 1 day. In this embodiment, it is set to 4 hours. It can balance the system's response speed and the stability of long-term trends, and help to capture load changes and abnormal fluctuations in system performance in a timely manner.
[0157] The preset resource threshold is the upper limit of resource utilization set when determining whether a node is resource overloaded. It depends on the system's hardware configuration and resource allocation strategy and is typically set to 80% to 90% of the node's maximum resource capacity. In this embodiment, it is set to 85% to ensure timely response when resource usage approaches the limit, avoiding performance degradation caused by system resource overload.
[0158] The preset completion threshold is a standard time used to determine whether a node is experiencing latency issues. This refers to the maximum time a node requires to complete a task. This threshold is determined by the task's normal execution time and the expected response time, and is typically set at 1.5 to 2 times the standard execution time. In this embodiment, the threshold is set at 1.5 times the standard execution time to promptly identify nodes with response times exceeding the normal range, thereby preventing latency issues from impacting overall task performance.
[0159] The preset resource fluctuation threshold is used to determine whether the node resource usage fluctuates too much. It is usually reflected in the standard deviation of the node's resource utilization rate. It depends on the load fluctuation characteristics of the system and is usually set to a fluctuation range of 5% to 15%. In this embodiment, it is set to 10%. It can promptly identify nodes with unstable resource usage so that further optimization measures can be taken.
[0160] The preset completion time fluctuation threshold refers to whether the fluctuation range of the completion time of the node exceeds the set range during the task execution process. It is set according to the complexity of the task and the stability of the execution environment. It is generally set between 5% and 20%. In this embodiment, it is set to 15%. It can effectively distinguish the execution delay fluctuations caused by external disturbances or insufficient resources, and ensure the stable operation of system performance.
[0161] First, by calculating the average resource utilization rate and checkpoint completion time of all focus nodes within the preset abnormality judgment time, determine whether it exceeds the set threshold. If the resource utilization rate exceeds the preset resource threshold, the node is marked as a resource node; if the checkpoint completion time exceeds the preset completion threshold, it is marked as a delay node. For focus nodes marked as both resource nodes and delay nodes, they are directly determined to be abnormal nodes. If a focus node is only marked as a resource node, the standard deviation of the resource utilization rate is calculated to determine whether its fluctuation exceeds the preset resource fluctuation threshold; if it is only marked as a delay node, the standard deviation of the checkpoint completion time is calculated to determine whether its fluctuation exceeds the preset completion time fluctuation threshold, thereby determining whether it is an abnormal node.
[0162] By comprehensively analyzing fluctuations in resource utilization and checkpoint completion times, we can accurately identify abnormal nodes and dynamically adjust based on specific resource and latency performance to avoid misjudgments or missed detections. The average and fluctuation values of resource utilization and completion times provide a detailed quantitative basis for node health. Furthermore, through labeling and threshold setting, we provide a systematic and hierarchical identification mechanism for anomaly detection, effectively improving monitoring accuracy and responsiveness.
[0163] Please continue reading Figure 4 As shown, it is a decision logic diagram for determining a fault risk node in this embodiment;
[0164] The process of determining a number of fault risk nodes according to the task execution delay values of any two abnormal nodes includes:
[0165] Constructing a knowledge graph with each of the abnormal nodes as a graph node;
[0166] Calculating the standard deviation of the task execution delay value of each abnormal node within a preset risk determination time period to obtain a plurality of delay fluctuation values;
[0167] Calculating the correlation coefficient of the delay fluctuation values of any two abnormal nodes to obtain a delay correlation;
[0168] When the delay correlation is greater than a preset standard correlation, connecting the corresponding two abnormal nodes to obtain a plurality of graph edges;
[0169] When the number of graph edges corresponding to each abnormal node is greater than a preset standard number of edges, the abnormal node is determined to be a fault risk node.
[0170] The preset risk determination period is typically determined based on the periodicity of task execution and the system's response speed. It depends on the task completion cycle and the timeframe to be monitored, and is typically set between one hour and one day to ensure that latency fluctuations are captured. In this example, a six-hour period is used to accurately reflect short- to medium-term latency fluctuations and promptly identify potential risks.
[0171] The preset standard correlation depends primarily on the expected correlation of task delays between nodes and the fault tolerance capability. It is typically set between 0.6 and 0.9 to ensure that the correlation calculation can identify true correlations. In this embodiment, the preset standard correlation is set to 0.75, which strikes a balance between accuracy and sensitivity, ensuring inter-node collaboration while avoiding misjudgments caused by minor fluctuations.
[0172] The preset standard number of edges depends on the number of nodes in the system and the complexity of the task, reflecting the connection density between nodes in the knowledge graph, and is usually set between 2 and 5. In this embodiment, the preset standard number of edges is set to 3, which can effectively control the complexity and computational cost of the graph while ensuring accurate recognition.
[0173] First, a knowledge graph is constructed based on each abnormal node. The latency fluctuation value is obtained by calculating the standard deviation of the task execution latency. Next, the correlation coefficient of the latency fluctuation values of any two abnormal nodes is calculated, and a connection is determined based on a preset standard correlation. If the number of edges in the connected graph exceeds the preset standard number of edges, the abnormal nodes are determined to be fault-risk nodes.
[0174] By comprehensively considering the correlation and connectivity of latency fluctuations between nodes, we can effectively identify nodes with potential failure risks. Calculating the standard deviation and correlation of latency fluctuations helps capture the coordinated fluctuation relationships between nodes. This detection of fluctuation patterns not only improves early warning capabilities for potential failures but also ensures that abnormal behavior across multiple nodes can be promptly linked and managed in a clustered manner.
[0175] Specifically, the preset standard consistency is adjusted according to the number of the fault risk nodes within the preset adjustment period to form the adjusted standard consistency, or the preset standard deviation index is adjusted to form the adjusted standard deviation index. The process includes:
[0176] Obtaining the number of the fault risk nodes at each moment within the preset adjustment time to obtain a number of risk node numbers;
[0177] When the number of risk nodes is greater than a preset node number threshold, calculating the inverse of the standard deviation of all risk nodes to obtain the node number concentration;
[0178] The preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted to form the adjusted standard deviation index.
[0179] The preset node number threshold is a standard value used to determine whether the number of nodes at risk of failure is abnormal. It usually depends on the capacity of the system, the scale of the task, and the distribution of the number of nodes during normal operation. The purpose of setting this threshold is to ensure that when the number of nodes fluctuates beyond the normal range, potential problems can be discovered in a timely manner and corresponding measures can be taken. It is usually set between 10 and 100, and the specific value will depend on the scale of the system and the accuracy requirements of monitoring. In this embodiment, the preset node number threshold is set to 50, which can effectively capture abnormal changes in the number of nodes while avoiding over-response to occasional small fluctuations.
[0180] The preset standard concentration range is used to measure the normal fluctuation range of node concentration. It determines whether the concentration of fault-risk nodes exceeds expectations within a certain period of time. It depends on the system's stability requirements and response capabilities and is typically set between 0.5 and 2. In this embodiment, the preset standard concentration range is set to 1.5. This ensures that while maintaining normal load, it can flexibly respond to relatively concentrated fault-risk nodes, thereby improving the system's sensitivity and resilience.
[0181] During the preset adjustment period, the system first monitors and obtains the number of nodes at risk of failure at each moment, generating a series of risk node count data. If this number exceeds the preset node threshold, the inverse standard deviation of the entire risk node count is calculated to determine the node concentration. Based on the relationship between this concentration and the preset standard concentration range, the system adjusts the preset standard consistency or standard deviation index to dynamically optimize monitoring parameters and improve system responsiveness and accuracy.
[0182] By monitoring the number and concentration of nodes at risk of failure, we can accurately assess fluctuations in system load and abnormal conditions. When the number of risky nodes is high and their concentration is high, adjustments to the consistency or standard deviation index allow for timely response to potential high risks and optimize system load distribution and task execution efficiency. This process ensures dynamic adjustments based on actual conditions when faced with failure risks, enhancing the flexibility and intelligence of operations and maintenance management.
[0183] Specifically, the preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted. The process of forming the adjusted standard deviation index includes:
[0184] When the node number concentration is less than the minimum value of the preset standard concentration range, reducing the preset standard consistency according to the relative deviation between the minimum value of the preset standard concentration range and the node number concentration and a preset adjustment coefficient to form the adjusted standard consistency;
[0185] When the node quantity concentration is greater than the maximum value of the preset standard concentration range, the preset standard deviation index is increased according to the relative deviation between the node quantity concentration and the maximum value of the preset standard concentration range and the adjustment coefficient to form the adjusted standard deviation index.
[0186] The preset standard concentration range is determined based on system requirements and design goals, reflecting the need to maintain an acceptable degree of node concentration. It is typically set between 5% and 95% to avoid over-concentration or over-dispersion that could affect system stability. This range depends on the load level and system sensitivity requirements of the actual application scenario. In this embodiment, it is set between 10% and 90% to ensure stable operation under varying load conditions.
[0187] The preset adjustment coefficient is a parameter for adjusting the standard consistency or standard deviation index. It is usually set according to the response requirements and flexibility of the system, depending on the adaptability to changes and the tolerance to load fluctuations. It is usually set between 0.5-2. In this embodiment, it is set to 1.5. It can provide a reasonable adjustment range when there is a large change in node concentration, so as to ensure that the system can be adaptively adjusted under different working conditions.
[0188] When the node concentration deviates from the preset standard concentration range, the preset standard consistency or preset standard deviation index is adjusted appropriately based on its relative deviation from the minimum or maximum value and the preset adjustment coefficient. When the node concentration is below the minimum value of the standard range, the preset standard consistency is reduced to increase the system's sensitivity to changes in a small concentration of nodes. When the node concentration is above the maximum value of the standard range, the preset standard deviation index is increased to improve the system's adaptability to changes in a large concentration of nodes. This adjustment mechanism ensures that the system can flexibly respond to different node fluctuations and optimize overall performance.
[0189] By precisely controlling the preset standard consistency and standard deviation index, dynamic adjustments can be made based on changes in node concentration. By adopting appropriate adjustment strategies at different concentration levels, the system's ability to cope with load fluctuations is optimized. When concentration is low, reducing consistency can enhance sensitivity; while when concentration is high, increasing the deviation index can strengthen the system's fault tolerance, effectively preventing misjudgments and ensuring the system maintains high efficiency in the face of changing load environments.
[0190] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. An intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring, characterized in that: include: Obtain in real time the order receiving rate, order success rate, task execution delay value, data processing rate, state accumulation degree, resource utilization rate, and checkpoint completion time of each monitoring node in the streaming computing task of the order platform; Determine a number of high-load nodes based on the order receiving rate, order success rate, and consistency with preset standards; Determine a number of focus nodes according to the data processing rate, the state accumulation degree, and a preset standard deviation index of each of the high-load nodes; Determine a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes; Determine a number of fault risk nodes according to the task execution delay values of any two of the abnormal nodes; Adjusting the preset standard consistency according to the number of the fault risk nodes within the preset adjustment period to form an adjusted standard consistency, or adjusting the preset standard deviation index to form an adjusted standard deviation index; issuing an operation and maintenance alert for all the fault risk nodes re-determined based on the adjustment standard consistency or the adjustment standard deviation index; The preset standard consistency is a standard value used to measure the consistency between the current collected data and the standard data; The preset standard deviation index is a standard value that reflects whether the fluctuation degree of the current acquisition parameters is within a reasonable range, based on the mean square deviation level of historical data in the same type of scenario; the preset adjustment time refers to the minimum response time required for the system to perform automatic adjustment when the consistency or deviation index does not meet the standard.
2. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 1 is characterized in that: The process of determining a number of high-load nodes according to the order receiving rate, the order success rate, and the preset consistency threshold includes: Calculate the standard deviation of the order receiving rate of all the orders within the preset load determination time period to obtain a receiving rate fluctuation value; Calculate the standard deviation of the success rate of all the orders within the preset load determination time to obtain a success rate fluctuation value; Normalizing the receiving rate fluctuation value to obtain a receiving rate normalized value, and normalizing the success rate fluctuation value to obtain a success rate normalized value; Calculating a correlation coefficient between the normalized value of the receiving rate and the normalized value of the success rate to obtain a fluctuation consistency; A number of high-load nodes are determined according to the fluctuation consistency and the preset consistency threshold.
3. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 2 is characterized in that: The process of determining a number of high-load nodes according to the fluctuation consistency and the preset consistency threshold includes: When the fluctuation consistency is greater than the preset consistency threshold, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
4. The intelligent operation and maintenance method for streaming computing tasks based on real-time monitoring according to claim 3 is characterized in that: The process of determining a number of high-load nodes according to the fluctuation consistency and the preset consistency threshold further includes: When the fluctuation consistency is less than or equal to the preset consistency threshold, the standard deviation of the order receiving rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary receiving fluctuation values, and the standard deviation of the order success rate from the initial moment to each moment within the preset load determination time period is calculated to obtain a number of temporary success fluctuation values; Drawing a curve showing how all the temporary reception fluctuation values change over time within the preset load determination time period to obtain a reception change curve; Plotting a curve showing changes of all the temporary success fluctuation values over time within the preset load determination time period to obtain a success change curve; Calculating the cosine similarity between the reception change curve and the success change curve to obtain the fluctuation change similarity; When the fluctuation change similarity is greater than a preset standard change similarity, the monitoring node is determined to be the high-load node, so as to determine a number of the high-load nodes.
5. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 4 is characterized in that: The process of determining a number of focus nodes according to the data processing rate, the state accumulation degree, and the preset standard deviation index of each high-load node includes: Calculating the standard deviation of all data processing rates within a predetermined period of time to obtain a processing rate fluctuation value; Calculating the standard deviation of the accumulation degree of all the states within the preset attention determination time length to obtain an accumulation fluctuation value; Calculating a relative deviation between the processing rate fluctuation value and a preset processing rate fluctuation threshold to obtain a processing fluctuation deviation; Calculating a relative deviation between the accumulation fluctuation value and a preset accumulation fluctuation threshold value to obtain an accumulation fluctuation deviation; A number of focus nodes are determined according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index.
6. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 5 is characterized in that: The process of determining a number of focus nodes according to the processing fluctuation deviation, the accumulation fluctuation deviation, and the preset standard deviation index includes: Calculating the product of the processing fluctuation deviation and a preset processing fluctuation weight to obtain a processing fluctuation factor; Calculating the product of the stacking fluctuation deviation and a preset stacking fluctuation weight to obtain a stacking fluctuation factor; Calculating the sum of the processing fluctuation factor and the stacking fluctuation factor to obtain a fluctuation deviation index; When the fluctuation deviation index is greater than the preset standard deviation index, the high-load node is determined to be the focus node, so as to determine a plurality of focus nodes.
7. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 6 is characterized in that: The process of determining a number of abnormal nodes according to the resource usage rate and the checkpoint completion time of each of the focus nodes includes: Calculate the average of all resource usage rates within the preset abnormality judgment time to obtain the average resource usage rate; When the average resource usage rate is greater than a preset resource threshold, marking the focus node as a resource node; Calculate the average completion time of all the checkpoints within the preset abnormality judgment time to obtain the average completion time; When the average completion time is greater than a preset completion threshold, marking the node of interest as a delay node; When the concerned node is marked as the resource node and the delay node at the same time, determining the corresponding concerned node as the abnormal node to determine a number of abnormal nodes; When the focus node is only marked as the resource node, calculating the standard deviation of all resource usage rates to obtain a resource fluctuation value; When the resource fluctuation value is greater than a preset resource fluctuation threshold, determining the focus node as the abnormal node to determine a number of abnormal nodes; When the focus node is only marked as the delay node, calculating the standard deviation of the completion time of all the checkpoints to obtain a completion time fluctuation value; When the resource fluctuation value is greater than a preset completion time fluctuation threshold, the focus node is determined to be the abnormal node, so as to determine a number of abnormal nodes.
8. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 7 is characterized in that: The process of determining a number of fault risk nodes according to the task execution delay values of any two abnormal nodes includes: Constructing a knowledge graph with each of the abnormal nodes as a graph node; Calculating the standard deviation of the task execution delay value of each abnormal node within a preset risk determination time period to obtain a plurality of delay fluctuation values; Calculating the correlation coefficient of the delay fluctuation values of any two abnormal nodes to obtain a delay correlation; When the delay correlation is greater than a preset standard correlation, connecting the corresponding two abnormal nodes to obtain a plurality of graph edges; When the number of graph edges corresponding to each abnormal node is greater than a preset standard number of edges, the abnormal node is determined to be a fault risk node.
9. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 8 is characterized in that: The process of adjusting the preset standard consistency according to the number of the fault risk nodes within the preset adjustment period to form the adjusted standard consistency, or adjusting the preset standard deviation index to form the adjusted standard deviation index includes: Obtaining the number of the fault risk nodes at each moment within the preset adjustment time to obtain a number of risk node numbers; When the number of risk nodes is greater than a preset node number threshold, calculating the inverse of the standard deviation of all risk nodes to obtain the node number concentration; The preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted to form the adjusted standard deviation index.
10. The intelligent operation and maintenance method for stream computing tasks based on real-time monitoring according to claim 9 is characterized in that: The preset standard consistency is adjusted according to the node quantity concentration and the preset standard concentration range to form the adjusted standard consistency, or the preset standard deviation index is adjusted. The process of forming the adjusted standard deviation index includes: When the node number concentration is less than the minimum value of the preset standard concentration range, reducing the preset standard consistency according to the relative deviation between the minimum value of the preset standard concentration range and the node number concentration and a preset adjustment coefficient to form the adjusted standard consistency; When the node quantity concentration is greater than the maximum value of the preset standard concentration range, the preset standard deviation index is increased according to the relative deviation between the node quantity concentration and the maximum value of the preset standard concentration range and the adjustment coefficient to form the adjusted standard deviation index.
Citation Information
Patent Citations
Cluster multi-dimensional anomaly monitoring method and device, equipment and storage medium
CN113220534A
Operation and maintenance task operation method and system oriented to supercomputing and intelligent computing environments
CN117389837A