Server fault monitoring method and system

By performing threshold configuration, windowed aggregation and multi-scale analysis on the server's key indicator data, resource pressure factor points are constructed to realize real-time monitoring and response in high concurrency scenarios, solving the problem of difficulty in timely identification and prevention of peak occupation of server resources in the existing technology, and improving prevention and control capabilities and system stability.

CN120144409AInactive Publication Date: 2025-06-13CHANGSHA SKERUI ELECTRONIC TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510443599.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the existing technology to identify and prevent extreme peak occupation of server resources in high concurrency scenarios in time, resulting in potential risks of server overload or downtime.

Method used

By performing threshold configuration, windowed aggregation and multi-scale analysis on the server's key indicator data, resource pressure factor integration is constructed, and capacity expansion, current limiting, and fuse are automatically or semi-automatically selected to achieve real-time monitoring and response to the peak load interval.

Benefits of technology

It realizes accurate monitoring and prediction of resource abnormalities in a very short time, improves the ability to identify and prevent and control high concurrent loads in the microwave band, avoids excessive capacity expansion or frequent scheduling, and reduces operational risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144409A_ABST
    Figure CN120144409A_ABST
Patent Text Reader

Abstract

The invention discloses a server fault monitoring method and system, and relates to the technical field of server fault monitoring, and the method comprises the steps: carrying out the threshold configuration of key index data, and making a response when the key index data meets the threshold configuration; windowing aggregation and integration are carried out on the key index data, multi-scale energy changes are compared among different windows, and if the amplification ratio of window energy reaches or exceeds an amplification ratio threshold value, abnormal feature information is output; obtaining a resource pressure accumulation value through the integration of the constructed resource pressure factors, carrying out strategy selection according to the size of the resource pressure accumulation value and the abnormal feature type, and integrating a key index curve, a scheduling log and an alarm event generated in a load peak interval. And comparing the key index distribution in the abnormal peak value interval with the original baseline distribution and making a response. The operation of capacity expansion, current limiting, fusing and the like is automatically or semi-automatically selected, and automatic strategy selection is performed according to scenes, so that the service-level flexibility and protection mechanism is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server fault monitoring, and specifically to a server fault monitoring method and system. Background Technique

[0002] With the continuous expansion of the scale of online services, high-concurrency scenarios such as e-commerce promotions, live broadcasts, and rapid transaction matching frequently occur, and the server may experience an instantaneous surge in the number of accesses within an extremely short period of time. Conventional low-frequency monitoring usually samples at a minute or longer cycle, easily misses millisecond-to-second fluctuations, and is difficult to capture extreme peak occupations of thread pools, CPUs, and network resources in a timely manner, thus laying a hidden danger of server overload or downtime. To improve the ability to identify and prevent high-concurrency loads in the microwave band, it is necessary to combine technical means such as high-frequency acquisition, windowed aggregation, multi-scale analysis, and dynamic threshold configuration to create a server fault monitoring system that can accurately monitor and predict resource anomalies within an extremely short period.

[0003] In the invention patent with the application publication number CN112860527A, a fault monitoring method and device for an application server are provided. The fault monitoring method includes: obtaining a log file corresponding to call information on the application server and an operating system log; extracting transaction data and elapsed time data on the application server according to the log file, and extracting hardware operation data of the application server according to the operating system log; determining a fault monitoring result of the application server based on the transaction data, the elapsed time data, the hardware operation data, and a preset decision tree model; wherein, the decision tree model is used to determine whether the current state of the application server is faulty according to the transaction data, elapsed time data, and hardware operation data of the application server. The present invention can effectively improve the accuracy of application server fault monitoring and the operation and maintenance efficiency.

[0004] Combined with the above application and the content in the prior art:

[0005] In actual deployment, the traditional low-granularity monitoring system faces the following core problems: When facing the microwave-band traffic impact triggered by "seckill rush purchase" or "peak live broadcast", key indicators such as CPU, thread pool, and network connection often fluctuate significantly within seconds or even hundreds of milliseconds. However, conventional monitoring cannot identify them in time due to too low sampling frequency or lack of multi-scale analysis. In addition, there is no flexible early warning and severe threshold setting mechanism for key indicators such as thread pool utilization rate and sudden increase rate of hot caches. As a result, if there is a local overlimit, the system often does not have time to trigger capacity expansion or flow limiting, which in turn leads to cascading failures. Especially in high-concurrency scenarios, if the persistence and peak intensity of resource pressure cannot be quickly determined, there are often two polarization problems of either excessive capacity expansion or missed fault reporting, which not only wastes computing resources but also increases operation risks. Therefore, there is an urgent need for a server fault monitoring and prevention method based on high-frequency monitoring, multi-dimensional aggregation, and adaptive threshold configuration to cope with the sudden increase of resources from milliseconds to seconds, and combined with real-time data analysis and automated scheduling strategies, to identify and handle potential overload risks in time.

[0006] For this reason, the present invention provides a server fault monitoring method and system. Summary of the Invention

[0007] (1) Technical Problems to be Solved

[0008] Aiming at the deficiencies of the prior art, the present invention provides a server fault monitoring method and system. By configuring thresholds for key indicator data, a response is made when the key indicator data meets the threshold configuration; the key indicator data is windowed and aggregated and integrated, and the multi-scale energy changes are compared between different windows. If the increase ratio of window energy reaches or exceeds the increase ratio threshold, abnormal feature information is output; the resource pressure cumulative value is obtained by integrating the constructed resource pressure factor, and strategies are selected according to the size of the resource pressure cumulative value and the type of abnormal features. The key indicator curves, scheduling logs, and alarm events generated during the load peak interval are integrated, and the key indicator distribution in the abnormal peak interval is compared with the original baseline distribution and a response is made. Automatically or semi-automatically select operations such as capacity expansion, flow limiting, and fusing, and select automated strategies for different scenarios to achieve business-level flexibility and protection mechanisms, solving the technical problems described in the background art.

[0009] (2) Technical Solutions

[0010] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0011] A server fault monitoring method includes periodically reading key indicator data of a server and its core components and sending it to a monitoring center. After constructing an aggregation structure and setting identification rules in the monitoring center, threshold configuration is performed on the key indicator data, and a response is made when the key indicator data meets the threshold configuration;

[0012] Windowize and integrate the key indicator data. After performing multi-scale analysis on the current window using wavelet transform, compare the multi-scale energy changes between different windows. If the energy increase ratio Γ(w t ) reaches or exceeds the increase ratio threshold γ, generate and record abnormal event information, and output abnormal feature information; t )

[0013] Integrate the resource pressure cumulative value Ψ obtained by integrating the constructed resource pressure factor Φ(t). If the resource pressure cumulative value Ψ exceeds the threshold Γ, select a strategy based on the size of the resource pressure cumulative value Ψ and the type of abnormal feature, and perform one or more of elastic expansion, flow limiting, and fusing on the server;

[0014] Integrate the key indicator curves, scheduling logs, and alarm events generated during the load peak interval, and compare the key indicator distribution within the abnormal peak interval with the original baseline distribution. Generate the divergence coefficient JS(p||q) to quantify the difference between the two and make a response.

[0015] Furthermore, deploy high-frequency acquisition agents on the server and core components to periodically read the following key indicator data and send it to the monitoring center, including: CPU occupancy rate, thread pool load, number of network connections, cache hit rate. Save node information at the agent side; merge the key indicator data based on the node information and timestamp, and integrate it into a comprehensive monitoring record; add a "micro-window mark" to each record in the monitoring center, with w t representing the current window number;

[0016] Perform threshold configuration, including: setting the thread pool warning threshold W and the thread pool critical threshold C, the sudden increase rate of hot keys, and the peak offset factor R(t). When any threshold configuration is met, mark the corresponding high-frequency key indicator data as potentially abnormal and cache it in the monitoring center; when the marked entries reach a predetermined number or there is a centralized spike, trigger the windowing analysis and abnormal identification process.

[0017] Furthermore, windowize and aggregate the key indicator data with a duration of Δt, integrate it into a status record, and mark the window number w t and the timestamp interval [t, t + Δt); let X(τ) represent the composite sequence within the window, and the wavelet coefficient W X (a, b) at scale a and time shift b can be expressed as: Where: X(τ) is the signal function obtained by aggregating the current window, ψ(·) is the wavelet basis function, a is the scale parameter, and b is the time shift parameter.

[0018] Furthermore, after obtaining the wavelet coefficient W XAfter (a, b), calculate the multi-scale energy Ω(b) at a time shift of b to measure the intensity of resource fluctuations: where A is the set of selected scales, β is the exponent used to control non-linear amplification, and the larger it is, the more sensitive it is to local spikes;

[0019] To compare the local energy changes between different windows, if Γ(w t ) exceeds the predetermined increase ratio threshold γ, it is considered that there is an abnormal increase in the corresponding time period; define the increase ratio Γ(w t ) as follows: where Ω(w t ) is the multi-scale energy of this window w t , Ω(w t-1 ) is the multi-scale energy of the previous window, and ε is a small constant to prevent the denominator from being zero; where each window w t corresponds to the timestamp interval [t, t + Δt); Γ(w t ) is the increase ratio of the energy of window w t .

[0020] Furthermore, in the window w t that satisfies the increase ratio determination condition Г(w t ) ≥ γ or has a serious threshold overrun, an abnormal event will be generated and recorded, and the abnormal event information will be saved in a unified format or pushed to the alarm message queue; including, the trigger time w t , the abnormal indicator and its peak value, the wavelet multi-scale energy Ω(w t ) and the increase ratio Γ(w t ); cluster the combination of the wavelet feature vector and other key indicators to output abnormal feature information.

[0021] Furthermore, introduce a resource pressure factor Φ(t) to measure the real-time pressure level; where, at time t, the difference between key indicators such as the thread pool utilization rate and the CPU utilization rate and their baseline values or severe lines is denoted as ΔR i (t), and ΔR i (t) is synthesized into a comprehensive pressure factor Φ(t), and the formula is as follows:

[0022]

[0023] Parameter meaning: R(t) is the resource usage vector, usually of dimension n×1, where r i (t) represents the i-th resource usage degree corresponding to time t; W is an n×n weight matrix; α is a non-linear amplification coefficient (α ≥ 1);

[0024] Integrate the resource pressure factor Φ(t) within the corresponding abnormal window over the duration Δt to obtain the resource pressure accumulation value Ψ:

[0025] where: t 0 is the abnormal window start time, τ is the integration variable; if the resource pressure cumulative value Ψ exceeds the threshold Γ, an automated scheduling instruction is triggered.

[0026] Furthermore, policy selection is performed based on the magnitude of the resource pressure cumulative value Ψ and the abnormal feature type. Specifically, if the resource pressure cumulative value Ψ is mainly caused by continuous over-limit of the thread pool or CPU, additional instances or container nodes are preferentially started to expand the overall computing power; if high concurrent access occurs in the downstream database or cache layer, the access volume of hot keys is simultaneously evaluated, and traffic limiting is implemented on the API gateway or proxy layer to reduce the instantaneous request peak; if it is found that the downstream service has a risk of response collapse, a circuit breaker mechanism is applied to temporarily block requests to the downstream service.

[0027] Furthermore, the integrated content includes the evolution of key metrics during peak periods, the execution time of elastic scaling, the number of scaled nodes, or the start and stop times of traffic limiting / circuit breaking, the abnormal types and severity levels marked within each time window;

[0028] Align the load curve with the timeline of the disposal operations, and analyze whether metrics such as CPU usage rate and the length of the thread pool waiting queue show a significant decline before and after operations such as scaling and traffic limiting are executed. If certain metrics still remain high after the disposal measures, an analysis instruction is sent to the outside.

[0029] Furthermore, after receiving the analysis instruction, let p(x) represent the metric distribution of the original baseline (obtained by statistics in normal or previously known scenarios), and q(x) represent the metric distribution observed during this peak period. The Jensen-Shannon divergence can be used to generate the divergence coefficient JS(p||q) to quantify the difference between the two:

[0030]

[0031] where; is the intermediate distribution, KL represents the Kullback-Leibler divergence; p(x) is the original baseline distribution, and q(x) is the metric distribution monitored in the new scenario;

[0032] Set the distribution threshold JS th to determine whether the distribution difference is sufficient to indicate that the original baseline is outdated; when JS(p||q) > JS th , revise or reconstruct the corresponding threshold configuration and recalibrate the threshold configuration of the key metrics.

[0033] A server fault monitoring system, including,

[0034] The threshold configuration unit periodically reads the key metric data of the server and core components and sends it to the monitoring center. After constructing an aggregation structure and setting identification rules in the monitoring center, it configures the thresholds for the key metric data and responds when the key metric data meets the threshold configuration.

[0035] The abnormal event recording unit performs windowed aggregation and integration on the key metric data. After performing multi-scale analysis on the current window using wavelet transform, it compares the multi-scale energy changes between different windows. If the amplitude increase ratio Γ(w t of the energy reaches or exceeds the amplitude increase ratio threshold γ, it generates and records abnormal event information and outputs abnormal feature information. t ) reaches or exceeds the amplitude increase ratio threshold γ, generates abnormal event information and records it, and outputs abnormal feature information.

[0036] The policy matching unit integrates the constructed resource pressure factor Φ(t) to obtain the resource pressure accumulation value Ψ. If the resource pressure accumulation value Ψ exceeds the threshold Γ, it selects a policy based on the size of the resource pressure accumulation value Ψ and the type of abnormal feature, and performs one or more of elastic expansion, flow limiting, and fusing on the server.

[0037] The difference response unit integrates the key metric curves, scheduling logs, and alarm events generated during the peak load period, and compares the key metric distribution within the abnormal peak interval with the original baseline distribution, generates the divergence coefficient JS(p||q) to quantify the difference between the two and makes a response.

[0038] (III) Beneficial effects

[0039] The present invention provides a server fault monitoring method and system, having the following beneficial effects:

[0040] 1. By setting the warning threshold and severe threshold of the thread pool, the sudden increase factor of hot keys, and the peak shift factor, etc., different limits can be defined for different metrics, realizing a more refined alarm and detection strategy, and ensuring the tight connection and automation of the entire monitoring and detection process.

[0041] 2. Aggregate the high-frequency monitoring data with Δt as the window to ensure that the resource usage can be statistically analyzed even on an extremely short time scale, and quickly locate possible instantaneous peaks or fluctuations. The sequence of a single metric or a combination of multiple metrics can be input into subsequent wavelet transform or other algorithms to further enrich the dimension of abnormal identification.

[0042] 3. By examining the signal at different frequencies (scales) through wavelet transform, the instantaneous spikes of resources such as thread pools / CPUs in the microwave band can be highlighted, avoiding the problem that traditional time-domain analysis cannot capture "extremely short-term peaks"; calculating the multi-scale energy Ω(b) concentrates the wavelet coefficients at each scale into a scalar, helping to quickly determine whether there are significant abnormal peaks in the corresponding window and allowing for comparison with the context; comparing Ω(b) between adjacent windows, and only marking it as a sudden anomaly if the increase ratio exceeds the limit, avoiding repeated triggering of alarms during the stable or slowly rising stage and improving the accuracy of alarms.

[0043] 4. When the increase ratio of the window energy or the severe threshold of the indicator is exceeded, an abnormal event is automatically generated in the monitoring center, with information such as timestamp, indicator peak value, and multi-scale energy, providing sufficient support for operation and maintenance decision-making.

[0044] 5. By integrating (or discretely summing) the comprehensive stress factors within the abnormal window, the cumulative value of resource stress is obtained; if this value continuously exceeds the business-set threshold, it indicates that the overload state is not only a short-term spike but rather relatively continuous and severe, helping to accurately distinguish between self-healing small fluctuations and resource crises that require immediate intervention, facilitating the system to achieve a balance between stability and cost, and reducing over-expansion or frequent scheduling.

[0045] 6. According to the cumulative value of resource stress and the type of abnormal characteristics, operations such as expansion, flow limiting, and fusing are automatically or semi-automatically selected, and scenario-based automated strategy selection is implemented to achieve business-level flexibility and protection mechanisms. After the scheduling is completed, the monitoring center continues to observe the trend of the comprehensive stress factors. If they are still at a high level, the strategy can be repeated or upgraded; if they decline, it indicates that the disposal is effective, avoiding waste caused by over-expansion or frequent flow limiting.

[0046] 7. Through an advanced distribution comparison method, the deviation degree of the current peak interval indicator distribution from the original baseline distribution in terms of shape is quantitatively measured, avoiding reliance on subjective judgment. If the divergence coefficient is higher than the threshold, the thread pool warning line, cache hot spot threshold, etc. are dynamically revised, which is beneficial to improving the accuracy of subsequent monitoring; if it is confirmed that the baseline is outdated, new thresholds and detection strategies are automatically loaded; if there is only a small deviation, only fine-tuning of the alarm rules is required, saving operation and maintenance costs and maintaining stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic flow chart of the server fault monitoring method of the present invention;

[0048] Figure 2 is a schematic structural diagram of the server fault monitoring system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Please refer to Figure 1 , the present invention provides a server fault monitoring method, including,

[0051] Step 1: Periodically read the key indicator data of the server and core components and send it to the monitoring center. After constructing an aggregation structure and setting identification rules in the monitoring center, perform threshold configuration on the key indicator data, and make a response when the key indicator data meets the threshold configuration;

[0052] The said Step 1 includes the following contents:

[0053] Step 101: Deploy high-frequency collection agents in the server and core components (including API gateways, application processes, database / cache layers, etc.) to periodically read the following key indicator data and send it to the monitoring center, including:

[0054] CPU occupancy rate (which can distinguish the load of each core), thread pool load (number of active threads, maximum number of threads, length of waiting queue), number of network connections (active connections and queued connections); cache hit rate (including access statistics of hot keys);

[0055] Each collection agent performs sampling operations in milliseconds or seconds-level cycles, and adopts an asynchronous or batch sending mode at the transport layer to reduce the performance impact on business applications; by saving node information (such as node ID, container ID, label of the business module where it is located) at the agent side, it is ensured that the monitoring data has a traceable context;

[0056] Periodically sample key indicators (CPU occupancy rate, thread pool load, number of network connections, cache hit rate) at milliseconds or seconds-level intervals, which can accurately capture resource fluctuations in the microwave band, thereby avoiding missing short-term peaks in conventional low-frequency monitoring; adopt an asynchronous or batch sending mode to reduce the real-time pressure on the running core business and ensure the minimum performance overhead of the probe itself;

[0057] By saving node information (node ID, container ID, business module to which it belongs, etc.) at the collection agent side, the "source location" of the monitoring data is realized, which is convenient for quickly locating the specific server or service component when an abnormality occurs subsequently.

[0058] Step 102: To facilitate the unified management of monitoring data within a micro-window, it is necessary to construct an aggregation structure and set identification rules in the monitoring center, including:

[0059] Merge the key indicator data based on node information and timestamp, and integrate multi-dimensional status data such as CPU, thread pool, network connection count, and cache hotspot statistics within the same time slice into a comprehensive monitoring record; add a "micro-window mark" to each record in the monitoring center, where w t represents the current window number, ensuring the coherence of real-time comparison between adjacent windows and subsequent feature calculations; push the original record to the message queue, and use distributed cache or time series database for hierarchical storage;

[0060] In the monitoring center, merge the key indicators of different dimensions based on node information and timestamp to form a comprehensive monitoring record, enabling operation and maintenance personnel to observe data such as CPU, threads, network, and cache simultaneously in a single view for easy overall observation; add a "micro-window mark" to each record to make subsequent sliding window or time period analysis more intuitive. Each window can quickly compare resource usage differences, support second-level or even sub-second anomaly detection, and improve the efficiency of real-time comparison;

[0061] Hierarchical storage ensures scalability. Combining the message queue, distributed cache, or time series database can maintain high-throughput data writing and query speeds in high-concurrency scenarios, providing highly available storage support for subsequent multi-scale analysis and anomaly marking.

[0062] Step 103: Perform threshold configuration, including: set the thread pool warning threshold W and the thread pool critical threshold C, such as W = 80% and C = 90% for the thread pool utilization rate;

[0063] Sudden increase rate of hot Key: Define the sudden increase factor T. If the access volume of a certain Key in the current micro-window surges more than T times compared to the previous window, it is regarded as an abnormal fluctuation;

[0064] Peak offset factor R(t): If R(t) is greater than a predefined threshold (such as 15 or 20), it is initially judged that a significant peak offset occurs in the corresponding window; measure the strong increase in CPU occupancy within the micro-window through the following formula:

[0065]

[0066] where v(t) represents the average value of CPU occupancy in the current micro-window, m(t - 1) represents the CPU baseline value of the previous window period (the median of the previous window can be taken), and ∈ is a small constant used to avoid division-by-zero errors;

[0067] When any threshold configuration is met, the corresponding high-frequency key indicator data is marked as potentially abnormal and cached in the monitoring center; when the number of marked entries reaches a predetermined quantity or there is a centralized spike, the windowed analysis and anomaly identification process is triggered;

[0068] When in use, combine the content in steps 101 to 103:

[0069] By setting the warning threshold, severe threshold, sudden increase factor of hot key, and peak offset factor, etc. for the thread pool, different limit values can be defined for different indicators, achieving a more refined alarm and detection strategy.

[0070] If a certain indicator (such as thread pool utilization rate) only slightly exceeds the warning threshold, it can be first marked as "potentially abnormal"; only when more over-limit conditions are met (such as a significant increase in the peak offset factor) is it determined to be a severe anomaly, greatly reducing the false alarm rate and reducing misreports and missed reports. When the number of marked entries accumulates to a certain quantity or there is a concentrated explosion, the subsequent steps (windowed aggregation and anomaly detection) can be automatically invoked to ensure the tight connection and automation of the entire monitoring and detection process.

[0071] Step 2: Window and aggregate the key indicator data and integrate it. After performing multi-scale analysis on the current window using wavelet transform, compare the multi-scale energy changes between different windows. If the increase ratio Γ(w t of the energy of ) reaches or exceeds the increase ratio threshold γ, generate and record the abnormal event information, and output the abnormal feature information; t ) reaches or exceeds the increase ratio threshold γ, generate and record the abnormal event information, and output the abnormal feature information;

[0072] The said step 2 includes the following content:

[0073] Step 201: Window and aggregate the key indicator data with a duration of Δt. If a total of n monitoring data points are collected within the Δt period, then integrate the indicator values {x 1 ,x 2 ,…,x n} of these n points into a status record; the status record retains the CPU, thread pool, connection count, cache hit rate, etc., and marks the window number w t and the timestamp interval [t, t + Δt);

[0074] To capture potential anomalies at the micro time scale, wavelet transform is used for multi-scale analysis. Let X(τ) represent the synthesized sequence within the window (which can be the aggregation result of a single indicator or the synthesized quantity of a multi-dimensional vector). The wavelet coefficient W X (a, b) at scale a and time shift b can be expressed as:

[0075]

[0076] Where: X(τ) is the signal function aggregated by the current window (which can correspond to CPU occupancy, number of threads, etc.), ψ(·) is the wavelet basis function used to extract local features at different scales, a is the scale parameter, and the smaller the value, the more sensitive it is to high-frequency (fine-grained) changes; b is the time translation parameter for positioning the signal in terms of its location.

[0077] Aggregate the high-frequency monitoring data with Δt as the window to ensure that resource usage can be statistically analyzed even on extremely short time scales (second level or sub-second level), and quickly locate possible instantaneous peaks or fluctuations. A single metric (such as CPU occupancy rate) or a combined sequence of multiple metrics (such as CPU + threads + network connections) can be input into subsequent wavelet transforms or other algorithms to further enrich the dimensions of anomaly recognition; "condense" the monitoring data points within the window into a single status record, which not only retains the key peaks and average information but also reduces the interference or computational overhead of irrelevant data on subsequent analysis.

[0078] Step 202, after obtaining the wavelet coefficient W X (a, b), calculate the multi-scale energy Ω(b) at a certain time translation b to measure the intensity of resource fluctuations:

[0079]

[0080] where A is the selected scale set, such as {a 1 , a 2 ,...}, β is the exponent used to control non-linear amplification (such as 1 or 2), and the larger it is, the more sensitive it is to local spikes; Ω(b) is the multi-scale energy at time point b, used to measure the degree of anomaly at the corresponding moment;

[0081] To compare the local energy changes between different windows, if Γ(w t ) exceeds the predetermined increase ratio threshold γ, it is considered that there is an abnormal increase in the corresponding time period; define the increase ratio Γ(w t ) as follows:

[0082]

[0083] where Ω(w t ) is the multi-scale energy of this window w t , Ω(w t-1 ) is the multi-scale energy of the previous window, and ∈ is a small constant to prevent the denominator from being zero; where each window w t corresponds to the timestamp interval [t, t + Δt), and within this time period, the aggregated value of a number of wavelet coefficients W X (a, b) is calculated, and the corresponding aggregated value is named Ω(w t ), which is the multi-scale energy of this window w t ;

[0084] Among them, Γ(w t ) is the amplification ratio of the energy of window w t , which is used to measure the change rate in adjacent time periods. The amplification ratio threshold γ is determined based on historical data or stress test results to distinguish significant sudden increases from normal jitters;

[0085] When Γ(w t ) reaches or exceeds the amplification ratio threshold γ, and the corresponding CPU occupancy or thread pool usage rate within the corresponding window is close to the severe threshold C, it is preliminarily determined that there is a transient anomaly with a high risk in the corresponding window;

[0086] By examining the signal at different frequencies (scales) through wavelet transform, the instantaneous spikes of resources such as the thread pool / CPU in the microwave band can be highlighted, avoiding the problem that traditional time-domain analysis cannot capture "extremely short-term peaks";

[0087] Calculating the multi-scale energy Ω(b) concentrates the wavelet coefficients at each scale into a scalar, which helps to quickly determine whether there is a significant abnormal peak in the corresponding window and can be compared with the context; comparing Ω(b) between adjacent windows, and only marking it as a sudden anomaly if the amplification ratio exceeds the limit, avoiding repeated triggering of alarms during the stable or slowly rising stage and improving the accuracy of alarms.

[0088] Step 203, the monitoring center will generate and record abnormal events in the window w t that meets the amplification ratio determination condition Γ(w t ) ≥ γ or the severe threshold is exceeded. The abnormal event information is saved in a unified format or pushed to the alarm message queue; including, the trigger time w t , the main abnormal indicators (CPU, thread pool, etc.) and their peak values, the wavelet multi-scale energy Ω(w t ) and the amplification ratio Γ(w t );

[0089] Record the extreme points of each indicator within the abnormal window and the degree of association between the indicators within the same window. For example, whether the surge in the thread pool waiting queue occurs concomitantly with the sudden increase in the cache. Cluster the combination of the wavelet feature vector (such as including Ω, Γ, etc.) and other key indicators (hot key access volume, connection number, etc.) (such as Isolation Forest, One-Class SVM), and output abnormal feature information.

[0090] When in use, combine the content in steps 201 to 203:

[0091] When the energy amplification ratio of the determination window or the severe threshold of the index is exceeded, an abnormal event is automatically generated in the monitoring center, with information such as timestamp, index peak value, and multi-scale energy, providing sufficient support for operation and maintenance decision-making. Details such as wavelet feature vectors, hotspot key correlation, and thread queuing phenomena are recorded together, facilitating subsequent fault mode mining and comparison in more complex scenarios such as large promotions or live broadcasts. For the identified abnormal windows, the fault types can be further classified by clustering to achieve refined diagnosis, providing support for the next resource pressure measurement or scheduling strategy.

[0092] Step 3: Integrate the constructed resource pressure factor Φ(t) to obtain the resource pressure accumulation value Ψ. If the resource pressure accumulation value Ψ exceeds the threshold Γ, select a strategy based on the size of the resource pressure accumulation value Ψ and the type of abnormal characteristics, and perform one or more of elastic expansion, flow limiting, and fusing on the server;

[0093] The above Step 3 includes the following content:

[0094] Step 301: When determining whether to trigger expansion or flow limiting, in order to comprehensively consider abnormal characteristic information and threshold exceeding situations, introduce the resource pressure factor Φ(t) to measure the real-time pressure degree; among them, at time t, the difference between key indicators such as thread pool utilization rate and CPU utilization rate and their baseline values or severe lines is denoted as ΔR i (t), and ΔR i (t) is synthesized into a comprehensive pressure factor Φ(t), and the formula is as follows:

[0095]

[0096] Parameter meaning: R(t) is the resource usage vector, usually of n×1 dimension, where r i (t) represents the i-th resource usage degree at time t (such as CPU occupancy rate, thread pool utilization rate, network bandwidth occupancy, cache hit rate decline degree, etc.), usually normalized to the interval [0, 1], and exceeding 1 means exceeding the limit;

[0097] W is an n×n weight matrix used to describe the relative importance or mutual influence between various resource indicators; WR(t) is equivalent to performing a linear transformation on the resource vector, thereby highlighting some dimensions that are more critical to the business or have a higher coupling degree with other indicators; α is a non-linear amplification coefficient (α≥1), which determines the form of the l α -norm. When α = 2, it is the Euclidean norm; if α>2, it can more significantly amplify the influence of sharp peak values: α = 1 means the absolute value weighted sum;

[0098] Integrate the resource pressure factor Φ(t) within the corresponding abnormal window over the duration Δt to obtain the resource pressure accumulation value Ψ:

[0099]

[0100] where: t 0 is the start time of the abnormal window, and τ is the integration variable;

[0101] If the cumulative value Ψ of resource pressure exceeds the threshold Γ (pre-set by the business scenario, such as 200 or 500, etc., dimensionless numbers), it can be judged that the resource overload is continuous and severe, triggering an automated scheduling instruction; otherwise, only record the alarm in the monitoring center and maintain the observation;

[0102] By defining a "resource usage vector" and a "weight matrix", various metrics such as CPU occupancy, thread pool usage, network bandwidth, cache heat, etc. can be integrated into a comprehensive pressure factor, presenting complex resource data as a single value for easy real-time comparison and monitoring.

[0103] By integrating (or discretely summing) the comprehensive pressure factor within the abnormal window, the cumulative value of resource pressure is obtained; if this value is continuously higher than the business-set threshold, it indicates that the overload state is not only a short-term spike but rather relatively continuous and severe, helping to accurately distinguish between self-healing small fluctuations and resource crises that require immediate intervention. In the case of only an instantaneous spike but the cumulative value not exceeding the threshold, it is possible to choose not to perform large-scale capacity expansion or forced flow limiting, avoiding resource waste or service jitter caused by short-term pulses, which is beneficial to achieving a balance between system stability and cost and reducing excessive capacity expansion or frequent scheduling.

[0104] Step 302: Based on the size of the cumulative value Ψ of resource pressure and the type of abnormal characteristics (such as thread pool bottleneck, cache hot spot impact, network traffic explosion, etc.), make strategy selections, including

[0105] Elastic capacity expansion: If the cumulative value Ψ of resource pressure is mainly caused by continuous overlimit of the thread pool or CPU, preferentially start additional instances or container nodes to expand the overall computing power;

[0106] Flow limiting: When high-concurrency access is monitored in the downstream database or cache layer, and at the same time evaluate the sudden increase in the access volume of hot keys, implement flow limiting on the API gateway or proxy layer to reduce the instantaneous request peak;

[0107] Circuit breaker: If it is found that the downstream service has a risk of response collapse, apply the circuit breaker mechanism to temporarily block requests to the downstream service;

[0108] After the scheduling is completed, the monitoring center continuously monitors the trend of the comprehensive pressure factor Φ(t) in the subsequent time period. If it gradually falls back to the safe range or is between the safe and warning ranges, this scheduling strategy is considered effective; if the comprehensive pressure factor Φ(t) remains high or climbs again, the scheduling is repeated, or upgraded to a higher-level measure, such as multi-node capacity expansion combined with hierarchical flow limiting;

[0109] Record the current scheduling process and the change of the comprehensive pressure factor Φ(t), including the number of expanded nodes, the current limiting threshold, the fuse switch time, etc.;

[0110] When in use, combine the content in Steps 301 and 302:

[0111] Automatically or semi-automatically select operations such as expansion, current limiting, and fusing according to the resource pressure accumulation value and the type of abnormal characteristics (thread bottleneck, network explosion, cache impact, etc.), and select automated strategies for different scenarios to achieve business-level flexibility and protection mechanisms. After the scheduling is completed, the monitoring center continues to observe the trend of the comprehensive pressure factor. If it is still at a high level, the strategy can be repeated or upgraded; if it drops, it means that the disposal is effective, avoiding waste caused by excessive expansion or frequent current limiting.

[0112] Record the values of the comprehensive pressure factor, the number of expanded nodes, or the current limiting configuration, etc. before and after scheduling, providing detailed basis for analyzing the disposal timeliness, resource cost, and fault evolution process afterwards.

[0113] Step Four: Integrate the key indicator curves, scheduling logs, and alarm events generated during the load peak period, and compare the key indicator distribution in the abnormal peak period with the original baseline distribution, generate the divergence coefficient JS(p||q) to quantify the difference between the two and make a response;

[0114] The above Step Four includes the following content:

[0115] Step 401: The monitoring center integrates the key indicator curves, scheduling logs, and alarm events generated during the just-experienced load peak period (for example, from the abnormal start time ts to the load drop time te).

[0116] The integration content includes: the evolution of key indicators such as thread pool, CPU, and cache during the peak period, the execution time of elastic expansion, the number of expanded nodes, or the start and stop times of current limiting / fusing, the type and severity of the abnormal conditions marked in each time window; align the load curve with the timeline of the disposal operations, and analyze whether indicators such as CPU usage rate and the length of the thread pool waiting queue have dropped significantly before and after the operations of expansion, current limiting, etc. If some indicators still remain at a high level after the disposal measures, send an analysis instruction to the outside;

[0117] Collect information such as key indicators, scheduling operation logs, and alarm events during the load peak period. After comprehensive integration, the whole process of fault occurrence and disposal can be reproduced, realizing the collaborative playback of multiple data sources and providing support for subsequent optimization; by comparing the start and stop times of expansion and current limiting with the CPU / thread pool indicator decline curves, identify the response lag or target achievement degree of the disposal measures. If the resources do not drop within the expected time period, prompt that it may be necessary to upgrade the scheduling strategy or revise the execution mechanism;

[0118] Step 402: After receiving the analysis instruction, to examine the applicability of the current monitoring baseline and threshold configuration during this high-concurrency impact, compare the distribution of key indicators (such as CPU allocation, thread pool usage, cache hit rate) within the abnormal peak interval with the corresponding original baseline distribution.

[0119] Let p(x) represent the indicator distribution of the original baseline (statistically obtained under normal or previously known scenarios), and q(x) represent the indicator distribution observed during this peak period. The Jensen-Shannon divergence (JS) can be used to generate the divergence coefficient JS(p||q) to quantify the difference between the two:

[0120]

[0121] where; is the intermediate distribution, KL represents the Kullback-Leibler divergence; p(x) is the original baseline distribution (which can include the joint distribution of multi-dimensional features such as CPU, threads, and cache); q(x) is the indicator distribution monitored in the new scenario;

[0122] If the divergence coefficient JS(p||q) is too high, it indicates that there are significant differences in the resource usage patterns between the new scenario and the baseline. The monitoring center sets the distribution threshold JS th to determine whether the distribution difference is sufficient to indicate that the original baseline is outdated;

[0123] When JS(p||q) > JS th , it can be determined that this high-concurrency has exceeded the range describable by the original baseline, and the corresponding threshold configuration (such as the warning threshold and severe threshold of the thread pool, cache hotspot upper limit, etc.) needs to be revised or rebuilt, and a threshold update instruction is sent to the outside.

[0124] If the value of JS(p||q) is close to or lower than JS th , it can be considered that the original baseline is still applicable, and only minor parameter adjustments or small modifications to the alarm strategy are required, or no processing is needed.

[0125] After receiving the threshold update instruction, recalibrate the threshold configuration (such as warning thresholds and severe thresholds) of key indicators such as thread pool utilization rate, CPU peak value, and cache surge rate. For example: increase the severe threshold to adapt to a larger concurrent load; add new indicator thresholds, such as the limit value of the access frequency of specific hot key.

[0126] When in use, combine the content in Steps 401 and 402:

[0127] By means of an advanced distribution comparison method, quantitatively measure the deviation degree of the current peak interval index distribution from the original baseline distribution in terms of morphology, avoiding relying on subjective judgment. If the divergence coefficient is higher than the threshold, it indicates that the originally defined baseline cannot cover this abnormal load scenario, and it is necessary to dynamically revise the thread pool warning line, cache hot spot threshold, etc., which is beneficial to improving the subsequent monitoring accuracy; if it is confirmed that the baseline is outdated, automatically load new thresholds and detection strategies; if there is only a small deviation, only fine-tune the alarm rules, saving operation and maintenance costs and maintaining stability.

[0128] Please refer to Figure 2 , the present invention provides a server fault monitoring system, including,

[0129] A threshold configuration unit that periodically reads the key index data of the server and core components and sends it to the monitoring center. After constructing an aggregation structure and setting identification rules in the monitoring center, perform threshold configuration on the key index data and respond when the key index data meets the threshold configuration;

[0130] An abnormal event recording unit that performs windowed aggregation and integration on the key index data, uses wavelet transform to perform multi-scale analysis on the current window, and compares the multi-scale energy changes between different windows. If the increase ratio Γ(w t ) of the window wt energy reaches or exceeds the increase ratio threshold γ, generate and record abnormal event information and output abnormal feature information;

[0131] A strategy matching unit that integrates the constructed resource pressure factor Φ(t) to obtain the resource pressure cumulative value Ψ. If the resource pressure cumulative value Ψ exceeds the threshold Γ, select a strategy according to the size of the resource pressure cumulative value Ψ and the type of abnormal feature, and perform one or more of elastic expansion, flow limiting, and fusing on the server;

[0132] A difference response unit that integrates the key index curve, scheduling log, and alarm event generated during the load peak interval, and compares the key index distribution in the abnormal peak interval with the original baseline distribution, generates the divergence coefficient JS(p||q) to quantify the difference between the two and makes a response.

[0133] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0135] In several embodiments provided in the present application, it should be correspondingly understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0136] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0137] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application and should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A server failure monitoring method, characterized in that: include, Periodically read the key indicator data of the server and core components and send it to the monitoring center. After the monitoring center builds the aggregation structure and sets the identification rules, it configures the thresholds for the key indicator data and responds when the key indicator data meets the threshold configuration; The key indicator data is aggregated and integrated in a windowed manner. After multi-scale analysis of the current window using wavelet transform, the multi-scale energy changes are compared between different windows. If the window w t The energy increase ratio Γ(w t ) reaches or exceeds the increase ratio threshold γ, generates and records abnormal event information, and outputs abnormal feature information; The resource pressure cumulative value Ψ is obtained by integrating the constructed resource pressure factor Φ(t). If the resource pressure cumulative value Ψ exceeds the threshold Γ, a strategy is selected based on the size of the resource pressure cumulative value Ψ and the type of abnormal characteristics, and one or more of elastic expansion, current limiting and circuit breaking are performed on the server; The key indicator curves, scheduling logs and alarm events generated in the peak load interval are integrated, and the key indicator distribution in the abnormal peak interval is compared with the original baseline distribution. The divergence coefficient JS (p||q) is generated to quantify the difference between the two and respond.

2. A server failure monitoring method according to claim 1, characterized in that: Deploy high-frequency collection agents on servers and core components to periodically read the following key indicator data and send them to the monitoring center, including: CPU usage, thread pool load, number of network connections, and cache hit rate; By saving node information on the proxy side, merging key indicator data based on node information and timestamps, and integrating them into comprehensive monitoring records, adding a "micro-window mark" to each record in the monitoring center, t Indicates the current window number; Perform threshold configuration, including setting the thread pool warning threshold W and the thread pool severe threshold C, the hot key sudden increase rate, and the peak shift factor R(t). When any threshold configuration is met, the corresponding high-frequency key indicator data will be marked as potential anomalies and cached in the monitoring center; when the marked items reach the predetermined number or a concentrated surge occurs, the window analysis and anomaly identification process will be triggered.

3. A server failure monitoring method according to claim 2, characterized in that: The key indicator data is windowed and aggregated with a duration of Δt, integrated into a state record, and marked with the window number w t and timestamp interval [t,t+Δt); let X(τ) denote the synthetic sequence in the window, the wavelet coefficient W at scale a and time shift b X (a,b) can be expressed as: Where: X(τ) is the signal function obtained by aggregating the current window, ψ(·) is the wavelet basis function, a is the scale parameter, and b is the time shift parameter.

4. A server failure monitoring method according to claim 3, characterized in that: In obtaining the wavelet coefficient W X (a, b), calculate the multi-scale energy Ω(b) at a certain time shift b to measure the intensity of resource fluctuation: Ω(b) = ∑ a∈A |W X (a,b)| β ; Where A is the selected scale set, β is the index used to control nonlinear amplification, and the larger it is, the more sensitive it is to local spikes; If Γ(w t ) exceeds the predetermined increase ratio threshold γ, it is considered that an abnormal increase occurs in the corresponding period; the increase ratio Γ(w t )as follows: Among them, Ω(w t ) is the window w t The multi-scale energy, Ω(w t-1 ) is the multi-scale energy of the previous window, ∈ is a small constant to prevent the denominator from having zero value; where each window w t Corresponding timestamp interval [t, t+Δt); Γ(w t ) is the window w t Energy increase ratio.

5. A server failure monitoring method according to claim 4, characterized in that: When the amplification ratio determination condition Γ(w t )≥γ or severe threshold exceeded window w t In the process, abnormal events are generated and recorded, and the abnormal event information is saved in a unified format or pushed to the alarm message queue, including , trigger time w t , abnormal indicators and their peak values, wavelet multi-scale energy Ω(w t ) and the amplification ratio Γ(w t ) ; Cluster the combination of wavelet feature vectors and other key indicators and output abnormal feature information.

6. A server failure monitoring method according to claim 5, characterized in that: The resource pressure factor Φ(t) is introduced to measure the real-time pressure level; where, at time t, the difference between key indicators such as thread pool usage and CPU utilization and their baseline values ​​or severe lines is recorded as ΔR i (t), ΔR i (t) is synthesized into a comprehensive pressure factor Φ(t), the formula is as follows: Parameter meaning: R(t) is the resource usage vector, usually of n×1 dimension, R(t) = [r1(t), r2(t), …r n (t)] T , where r i (t) represents the resource usage of the i-th item corresponding to time t; W is an n×n weight matrix; α is a nonlinear amplification coefficient (α≥1); The resource pressure factor Φ(t) in the corresponding abnormal window is integrated over the duration Δt to obtain the cumulative value of resource pressure Ψ: Where: t0 is the start time of the abnormal window, τ is the integral variable; if the accumulated value of resource pressure Ψ exceeds the threshold Γ, the automatic scheduling instruction is triggered.

7. A server failure monitoring method according to claim 6, characterized in that: The strategy is selected based on the size of the accumulated resource pressure value Ψ and the type of abnormal characteristics. For example, if the accumulated resource pressure value Ψ is mainly caused by continuous overlimit of the thread pool or CPU, additional instances or container nodes are started first to expand the overall computing power. If high concurrent access is detected in the downstream database or cache layer, the sudden increase in the number of hot key accesses is evaluated at the same time, and the API gateway or proxy layer is limited to reduce the instantaneous request peak. If the downstream service is found to have a risk of response collapse, the circuit breaker mechanism is applied to temporarily block requests to the downstream service.

8. A server failure monitoring method according to claim 7, characterized in that: The integration includes: The evolution of key indicators during peak periods, the execution time of elastic expansion, the number of expansion nodes, or the start and stop times of current limiting / circuit breaking, and the type and severity of anomalies marked in each time window; Align the load curve with the timeline of the disposal operation, and analyze whether indicators such as CPU utilization and thread pool waiting queue length show a significant decline before and after operations such as capacity expansion and current limiting. If certain indicators remain high after the disposal measures, issue analysis instructions to the outside.

9. A server failure monitoring method according to claim 8, characterized in that: After receiving the analysis instruction, let p(x) represent the indicator distribution of the original baseline, which is statistically obtained under normal or previously known scenarios, and q(x) represent the indicator distribution observed during this peak period. The Jensen-Shannon divergence can be used to generate the divergence coefficient JS(p||q) to quantify the difference between the two: in; is the intermediate distribution, KL represents the Kullback-Leibler divergence; p(x) is the original baseline distribution, and q(x) is the indicator distribution monitored in the new scenario; Set distribution threshold JS th Used to determine whether the distribution difference is sufficient to indicate that the original baseline is outdated; when JS(p||q)>JS th , revise or rebuild the corresponding threshold configuration, and recalibrate the threshold configuration of key indicators.

10. A server fault monitoring system, characterized in that: include, The threshold configuration unit periodically reads the key indicator data of the server and core components and sends it to the monitoring center. After the monitoring center builds the aggregation structure and sets the identification rules, it configures the threshold of the key indicator data and responds when the key indicator data meets the threshold configuration; The abnormal event recording unit aggregates and integrates the key indicator data in a windowed manner, uses wavelet transform to perform multi-scale analysis on the current window, and compares the multi-scale energy changes between different windows. t The energy increase ratio Γ(w t ) reaches or exceeds the increase ratio threshold γ, generates and records abnormal event information, and outputs abnormal feature information; The strategy matching unit obtains the resource pressure accumulation value Ψ by integrating the constructed resource pressure factor Φ(t). If the resource pressure accumulation value Ψ exceeds the threshold Γ, a strategy is selected based on the size of the resource pressure accumulation value Ψ and the type of abnormal characteristics, and one or more of elastic expansion, current limiting and circuit breaking are performed on the server. The difference response unit integrates the key indicator curves, scheduling logs and alarm events generated in the load peak interval, and compares the key indicator distribution in the abnormal peak interval with the original baseline distribution, generates a divergence coefficient JS (p||q) to quantify the difference between the two and respond.

Citation Information

Patent Citations

  • Fault monitoring method and device for application server

    CN112860527A