Health evaluation method, electronic equipment, storage medium and product
Through the health evaluation method based on timing indicator data and log data, the recovery strategy is dynamically adjusted, and the false alarm and missed response problems caused by fixed threshold monitoring methods in the existing technology are solved, which significantly improves the system's recovery efficiency and fault tolerance capabilities.
Patent Information
- Application Number
- CN202510562353.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the prior art, fixed threshold monitoring method is used to detect calculation system failures, which lacks flexibility, resulting in false alarms and missed alarms, affecting system recovery efficiency.
Through the health evaluation method based on timing indicator data and log data, the deviation degree of monitoring indicators and log abnormality probability are determined respectively, combined with the health score function, the recovery strategy is dynamically adjusted to achieve accurate detection and evaluation of system failures.
It significantly improves the system's recovery efficiency and fault tolerance, avoids false alarms and missed reports, and ensures the stable and reliable operation of the system.
Smart Images

Figure CN120104456A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to health assessment methods, electronic devices, storage media and products. Background Art
[0002] In today's digital age, computing systems are widely used in various fields, such as finance, medical care, scientific research, and communications. Their stable operation is crucial to ensuring the normal operation of various industries. As the scale and complexity of computing systems continue to increase, the risk of failure faced by the system is also increasing. In order to ensure that computing systems can provide services continuously and stably, timely and accurate detection of system failures and taking corresponding recovery measures have become key requirements.
[0003] At present, related technologies use fixed threshold monitoring to detect computing system failures, that is, set a fixed threshold, and when the actual value of the monitoring indicator exceeds the corresponding threshold, the failure is determined and the recovery operation is triggered. However, since the threshold in the related technology is fixed and unchanging, it lacks flexibility and cannot be dynamically adjusted according to the severity of the failure and the actual situation of the system. It is difficult to adapt to the complex and changeable state of the computing system, and it is very easy to have false positives and false negatives, which affects the recovery efficiency of the computing system. Summary of the invention
[0004] The present application provides a health assessment method, an electronic device, a storage medium and a product, so as to at least solve the problem that a simple fixed threshold monitoring method used in related technologies lacks flexibility and is prone to false alarms and missed alarms, resulting in low recovery efficiency of the system to be evaluated.
[0005] This application provides a health assessment method, including: Based on the time series indicator data and log data in the system to be evaluated, the monitoring indicator deviation corresponding to the time series indicator data and the log anomaly probability corresponding to the log data are determined respectively; Based on the monitoring indicator deviation and the log abnormality probability, determine the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log abnormality probability; According to the monitoring indicator deviation, log abnormality probability, indicator fusion weight and log fusion weight, combined with the health score function, the health score of the system to be evaluated is determined, so as to perform recovery operations on the system to be evaluated according to the health score.
[0006] The present application also provides a health assessment device, comprising: A first determination unit is used to determine, based on the time series indicator data and the log data in the system to be evaluated, the monitoring indicator deviation corresponding to the time series indicator data and the log anomaly probability corresponding to the log data; A second determining unit is used to determine an indicator fusion weight corresponding to the monitoring indicator deviation and a log fusion weight corresponding to the log abnormality probability based on the monitoring indicator deviation and the log abnormality probability; The evaluation unit is used to determine the health score of the system to be evaluated based on the monitoring indicator deviation, log abnormality probability, indicator fusion weight and log fusion weight, combined with the health score function, so as to perform recovery operations on the system to be evaluated according to the health score.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned health assessment methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned health assessment methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned health assessment methods when executed by a processor.
[0010] Through this application, based on the time series indicator data and log data in the system to be evaluated, the deviation of the monitoring indicator corresponding to the time series indicator data and the log anomaly probability corresponding to the log data are determined respectively; based on the deviation of the monitoring indicator and the log anomaly probability, the indicator fusion weight corresponding to the deviation of the monitoring indicator and the log fusion weight corresponding to the log anomaly probability are determined; according to the deviation of the monitoring indicator, the log anomaly probability, the indicator fusion weight and the log fusion weight, combined with the health score function, the health score of the system to be evaluated is determined, so as to perform recovery operations on the system to be evaluated according to the health score. More accurate detection and evaluation of system faults is achieved, avoiding the problems of false alarms and missed alarms caused by the lack of flexibility of the fixed threshold monitoring method in related technologies, and being able to dynamically adjust the recovery strategy according to different fault severities and actual system conditions, significantly improving the system's recovery efficiency and fault tolerance, and ensuring stable and reliable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 A schematic diagram of the first health assessment method provided in the embodiment of the present application; Figure 2A schematic diagram of a second health assessment method provided in an embodiment of the present application; Figure 3 A schematic diagram of a third health assessment method provided in an embodiment of the present application; Figure 4 A flowchart of a fourth health assessment method provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a health assessment device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0014] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0015] Computing systems (such as intelligent computing systems) are the core infrastructure supporting applications such as artificial intelligence, big data analysis, and high-performance computing. Driven by the rapid development of technologies such as deep learning, natural language processing, and computer vision, their importance is becoming increasingly prominent. Intelligent computing systems need to process large-scale data and complex computing tasks. However, as the scale and complexity of the system continue to increase, the reliability and fault tolerance of the system face huge challenges. Hardware failures, software errors, network anomalies, and other problems frequently occur, which may lead to data loss, calculation errors, or service interruptions, seriously affecting the stability and availability of the system.
[0016] At present, related technologies use a simple threshold monitoring method to detect computing system failures. Specifically, related technologies set fixed thresholds for each monitoring indicator in the computing system, and continuously monitor the actual values of these monitoring indicators during operation. Once the actual value of a monitoring indicator exceeds the preset fixed threshold, the system determines that a failure has occurred and triggers the corresponding recovery operation.
[0017] However, this simple threshold monitoring method has obvious defects. On the one hand, because the threshold is fixed and lacks flexibility, it cannot be dynamically adjusted according to the severity of the fault and the actual situation of the system. In actual operation, the operating state of the system is complex and changeable. At different times and in different environments, the same monitoring indicator value may represent different fault conditions, and fixed thresholds are difficult to adapt to such changes. On the other hand, this method is prone to false alarms and missed alarms. False alarms will cause the system to frequently perform unnecessary recovery operations, waste system resources, and affect the normal operation efficiency of the system; missed alarms will prevent the fault from being discovered and handled in time, causing the fault to further deteriorate, reducing the system's fault tolerance, and seriously affecting the stability and reliability of the system.
[0018] In addition, there are also related technologies that only use log analysis to troubleshoot problems in computing systems. However, relying solely on log analysis, it is impossible to monitor the system status in real time and it is difficult to capture the dynamic changes in the system status in a timely manner; at the same time, the efficiency of processing and analyzing logs alone is relatively low, and it is difficult to respond quickly when the system fails. In addition, the fault detection method based only on log analysis is not accurate, cannot effectively integrate multi-source data, and is difficult to comprehensively and accurately evaluate the health status of the computing system.
[0019] In order to solve the problems existing in the relevant solutions, the embodiment of the present application provides a health evaluation method, which combines the mutation detection of time series indicator data with the semantic analysis of log text to respectively determine the deviation of the monitoring indicator corresponding to the time series indicator data and the log anomaly probability corresponding to the log data, thereby realizing the deep fusion of multi-source heterogeneous data. At the same time, an adaptive fusion algorithm based on confidence evaluation determines the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log anomaly probability in real time based on the monitoring indicator deviation and the log anomaly probability; according to the monitoring indicator deviation, the log anomaly probability, the indicator fusion weight and the log fusion weight, combined with the health scoring function, the optimal health score of the system to be evaluated is determined, which effectively solves the contradiction between the sensitivity and the false alarm rate of the traditional static threshold method, realizes more accurate detection and evaluation of system failures, significantly improves the recovery efficiency and fault tolerance of the system, and ensures stable and reliable operation of the system.
[0020] It should be noted that the health assessment method in this application is not only applicable to intelligent computing systems, but also to other distributed computing systems with high requirements for reliability and fault tolerance, such as cloud computing platforms, big data processing clusters, etc. In the cloud computing platform, time series indicator data such as resource utilization and network traffic of virtual machines, as well as log data of systems and applications can be collected. Through similar dual-channel anomaly detection and fault propagation map construction methods, fault detection, prediction and fault tolerance processing of cloud computing platforms can be achieved. In the big data processing cluster, the time series indicator data and log data of data storage and computing nodes are analyzed, and the method of this application is used to improve the reliability and stability of the cluster to ensure the normal progress of data processing tasks.
[0021] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0022] Figure 1 A flowchart of the first health assessment method provided in an embodiment of the present application.
[0023] like Figure 1 As shown, the method comprises the following steps: Step 101 : Based on the time series indicator data and log data in the system to be evaluated, the monitoring indicator deviation corresponding to the time series indicator data and the log abnormality probability corresponding to the log data are determined respectively.
[0024] In some embodiments, system monitoring and alarm tools are used to collect timing indicator data such as CPU / GPU utilization, memory pressure, network latency, storage system, etc.; data collectors (such as Fluentd) are used to collect various log data generated by the intelligent computing system, such as operating system logs, application logs, device logs, etc.
[0025] Time series indicator data refers to system performance data recorded over time, and is usually used to measure the operating status of the system to be evaluated at different time points. These data include but are not limited to GPU utilization (measurement of computing resource usage), memory pressure (reflection of whether memory usage is close to saturation), network latency (response time for data transmission), and storage system performance (such as read and write speed, I / O operations, etc.). This application can collect these time series indicator data according to the collection frequency through system monitoring and alarm tools (such as Prometheus, Zabbix, etc.) to monitor the system's operating status in real time.
[0026] The monitoring indicator deviation refers to the degree of difference between the current time series indicator data and the normal or expected value. Before determining the monitoring indicator deviation corresponding to the time series indicator data, this application needs to aggregate the time series indicator data of each data category to obtain multi-indicator deviation data for each data category. Based on the multi-indicator deviation data and the trained prediction model (such as LSTM, Long Short-Term Memory, long short-term memory network, used to predict the normal value range of the monitoring indicator), the monitoring indicator deviation corresponding to the multi-indicator deviation data is output. The greater the monitoring indicator deviation, the farther the current time series indicator data deviates from the normal state, there may be potential problems, and the higher the system health risk.
[0027] Log data refers to various recorded information generated by the system to be evaluated during operation, including operating system logs (recording operating system startup, shutdown, errors, etc.), application logs (recording application running status, abnormal conditions, etc.) and device logs (recording hardware device status, failures, etc.). This application can use data collectors (such as Fluentd, Logstash, etc.) to collect these log data, store them, and centrally manage them for subsequent analysis and processing.
[0028] The log anomaly probability refers to the probability of determining whether the log is in a preset anomaly category by analyzing the log data. This application can use log analysis tools (such as BERT, Bidirectional Encoder Representations from Transformers, pre-trained language models for parsing log text) and preset neural networks to mine and analyze log data and identify the log anomaly probability of the log being in a preset anomaly category.
[0029] Step 102: Based on the monitoring indicator deviation and the log abnormality probability, determine the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log abnormality probability.
[0030] In some embodiments, in the health evaluation of the system to be evaluated, the monitoring indicator deviation and the log abnormality probability are both important evaluation indicators. However, different indicators may have different degrees of influence on the health of the system. Therefore, it is necessary to assign a suitable weight to each indicator so that their importance can be accurately reflected in the comprehensive evaluation. The indicator fusion weight is used to measure the relative importance of the monitoring indicator deviation in the health evaluation, while the log fusion weight is used to measure the importance of the log abnormality probability.
[0031] This application can determine the standard deviation of the monitoring indicator deviation and the log anomaly probability in real time, and determine the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log anomaly probability according to the standard deviation. This is because the standard deviation is a statistic that measures the degree of discreteness of data distribution, reflecting the fluctuation range of the monitoring indicator deviation or the log anomaly probability.
[0032] Step 103, according to the monitoring indicator deviation, log abnormality probability, indicator fusion weight and log fusion weight, combined with the health score function, determine the health score of the system to be evaluated, so as to perform a recovery operation on the system to be evaluated according to the health score.
[0033] In some embodiments, the health score function is a model for comprehensively evaluating the health of the system to be evaluated. The health score function can use the monitoring indicator deviation, log abnormality probability, indicator fusion weight and log fusion weight as input parameters, and obtain the health score of the system through a certain calculation method. The health score can be a value between 0 and 100, and the higher the score, the better the health of the system, and the lower the score, the more serious the problem in the system.
[0034] Specifically, the present application may use a health score function to calculate the health score of the system based on the output results of LSTM and BERT (i.e., monitoring indicator deviation and log anomaly probability), indicator fusion weight, and log fusion weight.
[0035] The health score function is:
[0036] in, To monitor the deviation of indicators; is the log anomaly probability; is the indicator fusion weight, is the log fusion weight; k is the slope coefficient (for example, it can be 3).
[0037] Furthermore, the present application may further select a recovery operation corresponding to a preset interval according to the preset interval corresponding to the calculated health score, and perform a corresponding recovery operation on the system to be evaluated.
[0038] Through this application, based on the time series indicator data and log data in the system to be evaluated, the deviation of the monitoring indicator corresponding to the time series indicator data and the log anomaly probability corresponding to the log data are determined respectively; based on the deviation of the monitoring indicator and the log anomaly probability, the indicator fusion weight corresponding to the deviation of the monitoring indicator and the log fusion weight corresponding to the log anomaly probability are determined; according to the deviation of the monitoring indicator, the log anomaly probability, the indicator fusion weight and the log fusion weight, combined with the health score function, the health score of the system to be evaluated is determined, so as to perform recovery operations on the system to be evaluated according to the health score. More accurate detection and evaluation of system faults is achieved, avoiding the problems of false alarms and missed alarms caused by the lack of flexibility of the fixed threshold monitoring method in related technologies, and being able to dynamically adjust the recovery strategy according to different fault severities and actual system conditions, significantly improving the system's recovery efficiency and fault tolerance, and ensuring stable and reliable operation of the system.
[0039] Figure 2 The flowchart of the second health assessment method proposed in this application is further shown. Figure 1 The embodiment shown further explains step 101. Figure 2 The following steps may be included.
[0040] Step 201 , collecting time series indicator data according to a sampling frequency corresponding to the time series indicator data, and performing window normalization processing on the time series indicator data, so as to obtain multi-indicator deviation data according to the time series indicator data after the window normalization processing.
[0041] In some embodiments, the present application can collect timing indicator data according to an initial sampling frequency corresponding to the timing indicator data; if the timing indicator data is greater than or equal to a preset fault threshold corresponding to the timing indicator data, the initial sampling frequency is adjusted to collect the timing indicator data according to the adjusted sampling frequency; if the duration that the timing indicator data is less than the preset fault threshold is greater than or equal to a preset recovery period, the sampling frequency is restored to the initial sampling frequency to collect the timing indicator data using the initial sampling frequency.
[0042] The timing indicator data is collected according to the initial sampling frequency corresponding to the timing indicator data; if the load of the system to be evaluated is greater than or equal to a preset load threshold, the initial sampling frequency is reduced to collect the timing indicator data using the reduced sampling frequency.
[0043] After the above-mentioned sampling frequency adjustment, when the timing indicator data is obtained, the present application can standardize the timing indicator data for each sliding window corresponding to the sampling frequency to obtain standard timing indicator data; determine the indicator difference between the standard timing indicator data and the ideal data corresponding to the standard timing indicator data; based on the indicator difference and the preset indicator weight corresponding to the standard timing indicator data, perform weighted aggregation processing on the standard timing indicator data of each data category to obtain multi-indicator deviation data for each data category.
[0044] In an optional embodiment of the present application, the present application can deploy sampling frequencies according to different data categories corresponding to the time series indicator data. The time series indicator data is sampled at a sampling frequency of 1 minute. The time series indicator data refers to fault indicators, which are key quantitative indicators used to reflect the health status of the system and for fault analysis.
[0045] Among them, the timing indicator data of different data categories are shown in Tables 1 to 5. It should be noted that the timing indicator data in Tables 1 to 5 are only used as examples. The specific indicators can be adjusted according to actual needs and are not limited in the embodiments of the present application.
[0046] The timing indicator data for computing resource categories is shown in Table 1:
[0047] The timing indicator data for the memory system category is shown in Table 2:
[0048] The timing indicator data for the network communication category is shown in Table 3:
[0049] The timing indicator data for the storage system category is shown in Table 4:
[0050] The time series indicator data for the environmental support category is shown in Table 5:
[0051] After 3 complete business cycles (training / inference), this application enables the automatic parameter adjustment mode. First, set the sampling frequency to the initial sampling frequency, that is, set the initial sampling frequency of each timing indicator data to its upper limit sampling frequency for collection (that is, the initial sampling frequency of each timing indicator data in Tables 1 to 5 above is the maximum value of the sampling frequency range). When an abnormality in the timing indicator data is detected, that is, when the timing indicator data exceeds its corresponding preset fault threshold (such as each timing indicator data in Tables 1 to 5 above exceeds its corresponding threshold), adjust the initial sampling frequency and automatically associate the timing indicator data to enter the high-frequency sampling mode.
[0052] Among them, if the duration of the time series indicator data being less than the preset fault threshold is less than the preset recovery period, sampling is still performed according to the adjusted sampling frequency. For example, enhanced monitoring of the high-frequency sampling frequency of 1.5 times the benchmark is maintained within 30 minutes after the time series indicator data fault is recovered. If the duration of the time series indicator data being less than the preset fault threshold is greater than or equal to the preset recovery period, the sampling frequency is restored to the initial sampling frequency.
[0053] At the same time, the present application also takes into account the impact of sampling on system performance. If the load of the system to be evaluated is greater than or equal to the preset load threshold, the initial sampling frequency is reduced. For example, if the load of the system to be evaluated is ≥70%, the sampling frequency is dynamically reduced.
[0054] Furthermore, the present application can make the monitored time series indicator data more usable and representative through sliding window standardization and multi-indicator aggregation.
[0055] Specifically, the time series indicator data of the system to be evaluated will fluctuate over time, and sliding window standardization can normalize the data to a specific interval, making data of different times and magnitudes comparable. It can effectively remove noise and abnormal fluctuations in the data, highlight the core features of the data, and make subsequent model training and analysis more stable and accurate. For example, when monitoring GPU utilization, this application can use sliding window standardization to eliminate the impact of instantaneous fluctuations caused by task scheduling, allowing the model to focus on long-term trends and accurately identify anomalies. In addition, it can also improve the efficiency and convergence speed of model training, avoid gradient vanishing or explosion problems, and make model training more efficient and stable.
[0056] In the case of different sampling frequencies, the core goal of sliding window standardization is still to standardize the time series indicator data in each sliding window so that the data has zero mean and unit variance, while taking into account the impact of sampling frequency changes on the time series indicator data in the sliding window. This application can standardize the time series indicator data in each sliding window through the following formula to obtain standard time series indicator data.
[0057] For a sliding window X={x1, x2, …, xn} containing n time series indicator data, the standardized standard time series indicator data The calculation formula is: ; is the mean of the time series indicator data in the sliding window, is the standard deviation of the time series indicator data in the sliding window, is the i-th time series indicator data before normalization in the sliding window; in, and The calculation formula is: ; ; n is the number of time series indicator data in the sliding window.
[0058] Taking into account that the time series indicator data of different sampling frequencies, that is, the time series indicator data of time intervals corresponding to different sampling frequencies may contribute differently to the statistics, the present application can further use a time-weighted approach to calculate the mean and standard deviation within each sliding window.
[0059] Assume that each time series indicator data The corresponding time interval weight is ,and . Then the weighted mean and weighted standard deviation The calculation formula is: ; .
[0060] Normalized time series indicator data for: ; Among them, the weight It can be determined according to the time interval corresponding to the sampling frequency of the time series indicator data. For example, the closer the time interval, the higher the weight of the time series indicator data.
[0061] The weight of this application The calculation formula is: ; in, It is time series indicator data The corresponding timestamp; is the maximum timestamp in the sliding window.
[0062] Furthermore, the operation of the system to be evaluated involves multiple complex monitoring indicators, and it is difficult to fully grasp the system status by analyzing a single indicator alone. Therefore, after obtaining the standard time series indicator data, this application performs weighted aggregation processing on multiple standard time series indicator data according to preset data categories to obtain multi-indicator deviation data corresponding to each data category. For example, the CPU, GPU, memory, network, storage and other time series indicator data are integrated to construct a comprehensive feature vector, that is, multi-indicator deviation data, which fully reflects the system operation status.
[0063] Among them, the present application can use a multi-indicator aggregation formula to perform weighted aggregation processing on multiple time series indicator data of each data category.
[0064] The multi-indicator aggregation formula is: ; in, for (actual value), that is, the actual value of the standard timing indicator data in this application; for (ideal value), which is the ideal value of the ideal data in this application; The multi-indicator aggregation formula of the present application represents weighted aggregation of different standard time series indicator data of each data category, calculating the weighted sum of the difference between the actual value and the ideal value, and using this weighted sum as the multi-indicator deviation data in the present application.
[0065] Indicator weight represents the weight of the i-th standard time series indicator data in a certain data category, and . The indicator weight reflects the importance of each standard time series indicator data in the comprehensive evaluation. In the system to be evaluated, indicators such as CPU, GPU, memory, and network have different degrees of influence on the overall performance of the system. This application can highlight the role of key indicators by pre-setting different indicator weights. For example, in deep learning training tasks, GPU utilization has a greater impact on computing performance and can be assigned a higher indicator weight; while the disk I / O indicator has a relatively small impact and can be assigned a lower indicator weight; Index difference It is to calculate the absolute difference between the actual value and the ideal value of each standard time series indicator data in a certain data category. is the actual monitoring value of the i-th standard time series indicator data, It is the ideal value or normal range value of the standard time series indicator data. Taking CPU usage as an example, assuming that the ideal CPU usage is between 30% and 70%, when the actual CPU usage is 80%, the difference is |80%-70%|=10%. This difference reflects the degree to which the indicator deviates from the ideal state.
[0066] The ideal value of each standard timing indicator in this application (i.e., the ideal data in this application) is set according to the design specifications and performance goals of the system. For example, when designing network equipment, it is stipulated that the network delay should not exceed 50 milliseconds under normal load, so 50 milliseconds is the ideal value of network delay; the system memory design capacity is 32GB, and the memory usage rate should not exceed 80% during normal operation, that is, the ideal value of memory usage is 25.6GB (32GB×80%).
[0067] This application obtains multi-indicator deviation data, i.e., multi-indicator aggregation result M, by multiplying the indicator weight of each standard time series indicator data in each data category by the corresponding indicator difference and then summing them up. Multi-indicator deviation data comprehensively reflects the degree to which the overall indicator of each data category deviates from the ideal state. The larger the M value, the farther the system as a whole deviates from the ideal state, and the more serious the possible problems; conversely, the smaller the M value, the closer the system operation state is to the ideal state.
[0068] Step 202 , obtaining log data generated by the system to be evaluated, and merging the log data based on the similarity of the log data to obtain merged log data.
[0069] In some embodiments, log data generated by the system to be evaluated is obtained; the log data is parsed to extract keywords and parameter positions in the log data; based on the keywords and parameter positions, the similarity corresponding to the log data is determined, and the log data with similarity greater than or equal to a preset similarity threshold is merged to obtain merged log data.
[0070] In an optional embodiment of the present application, the present application may use a data collector (such as Fluentd) to collect various types of log data generated by the system to be evaluated, such as operating system logs, application logs, device logs, etc.
[0071] After that, a clustering-based log template extraction algorithm (such as Drain3) can be used. By defining the structure of the log template, similar logs are clustered into one category by clustering, so as to obtain each merged log data after merging in this application, and each category corresponds to a log template. The log template extraction algorithm is mainly based on the matching and clustering of the keywords and parameter positions of the log. First, the log data can be parsed to extract the keywords and parameter positions. Then, based on this information, the similarity between the log data is calculated, and the logs with high similarity are merged into the same cluster. The center of each cluster is a log template.
[0072] Step 203 , performing anomaly analysis on the multi-indicator deviation data and the merged log data to obtain the monitoring indicator deviation degree corresponding to the multi-indicator deviation data and the log anomaly probability corresponding to the merged log data.
[0073] In some embodiments, the present application can train the prediction model based on the training set and verification set corresponding to the original multi-indicator deviation data, and verify the trained prediction model using the test set of the original multi-indicator deviation data to obtain a verification result; if the verification result of the prediction model is normal, the multi-indicator deviation data is input into the trained prediction model to obtain the deviation degree of the proctoring indicator output by the trained prediction model.
[0074] Among them, the prediction model is trained based on the training set and validation set corresponding to the original multi-indicator deviation data, and the trained prediction model is verified using the test set of the original multi-indicator deviation data, and the verification result specifically includes: preprocessing the original multi-indicator deviation data to obtain the training set, validation set and test set corresponding to the multi-indicator deviation data; inputting the training set into the prediction model, and performing forward propagation to determine the loss value of the loss function of the prediction model; based on the loss value and the validation set, back propagation is performed to update the parameters of the prediction model until the loss function converges to obtain the trained prediction model; using the validation set data to verify the deviation of the monitoring indicators output by the trained prediction model to obtain the verification result; if the verification result is normal, the test set is input into the trained prediction model to obtain the deviation of the monitoring indicators output by the trained prediction model.
[0075] The original multi-indicator deviation data is preprocessed to obtain the training set, verification set and test set corresponding to the multi-indicator deviation data, which specifically includes: normalizing the original multi-indicator deviation data; dividing the normalized original multi-indicator deviation data into a training set, a verification set and a test set; and converting the data formats of the training set, the verification set and the test set into the target format corresponding to the prediction model through a sliding window.
[0076] In an optional embodiment of the present application, the prediction model of the present application may be an LSTM prediction engine, which uses the LSTM prediction engine to analyze the obtained multi-indicator deviation data and predict the corresponding monitoring indicator deviation. The LSTM (Long Short-Term Memory) prediction engine is a powerful recurrent neural network that can process time series indicator data with long-term dependencies and is suitable for analyzing and predicting the collected system time series indicators to be evaluated (such as CPU / GPU utilization, memory pressure, network latency, etc.).
[0077] Specifically, the present application can normalize the original multi-index deviation data obtained, for example, using the Min-Max normalization method to scale the data to the [0, 1] interval to eliminate the impact of the dimensions between different indicators and improve the training effect of the model. The normalization formula is: ; in, is the normalized original multi-index deviation data, x is the original multi-index deviation data currently being processed, is the maximum value of all original multi-index deviation data, It is the minimum value of all original multi-index deviation data.
[0078] Afterwards, the normalized original multi-index deviation data is divided into a training set, a validation set, and a test set according to a preset ratio. Among them, the preset ratio can be 70% for the training set, 15% for the validation set, and 15% for the test set, and can also be adjusted according to actual needs, which is not limited in the embodiments of the present application. The training set is used for model training, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model.
[0079] After the training set, validation set, and test set are obtained, the present application can further convert the original multi-index deviation data of each data set into a target format suitable for the input of the prediction model. The present application can adopt a sliding window method, taking the original multi-index deviation data of several consecutive time steps as input and the original multi-index deviation data of the next time step as output, and apply this method to the training set, validation set, and test set to construct their input and output pairs respectively. For example, if the sliding window size is 10, the input is the original multi-index deviation data from time t-9 to time t, and the output is the original multi-index deviation data from time t+1, that is, the input is , the output is ,in represents the original multi-index deviation data at the t-th time step, is the original multi-indicator deviation data for the next time step.
[0080] The target format of the converted raw multi-index data is a series of input and output pairs, each of which contains a fixed-length input vector and an output value.
[0081] Furthermore, after the above-mentioned preprocessing, the training set, validation set and test set corresponding to the multi-index deviation data are obtained, and the present application can train the prediction model.
[0082] Specifically, the present application can first construct an initial prediction model, that is, determine the number of neurons in the LSTM layer (this number will affect the expressiveness of the model. The more neurons, the stronger the expressiveness of the model, but the training time will also increase accordingly), add a fully connected layer after the LSTM layer, and map the output of the LSTM layer to the dimension of the predicted deviation of the monitoring indicator. After the fully connected layer, a suitable activation function can be selected, such as a linear activation function (for regression problems).
[0083] After that, the application trains the initial prediction model based on the training set. That is, for the prediction problem of multi-indicator deviation data, the mean square error (MSE) is usually used as the loss function to measure the error between the model prediction value and the true value. The formula of the loss function is: ; in, It is the monitoring indicator data predicted based on the original multi-indicator deviation data. is the predicted value corresponding to the original multi-indicator deviation data, and n is the number of multi-indicator deviation data.
[0084] The present application may also select a suitable optimizer, such as the Adam optimizer, to update the parameters of the model so as to minimize the value of the loss function.
[0085] The training set and verification machine are input into the model, forward propagation is performed to calculate the loss, and then the parameters of the model are updated through back propagation. This process is repeated until the loss function converges.
[0086] When the loss function converges, the trained prediction model is obtained, and then the test set is used to evaluate the performance of the trained prediction model. The evaluation indicators include mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), etc. That is, the test set is input into the trained prediction model to obtain the output result, and it is determined whether the evaluation index between the output monitoring indicator data and its predicted value is within the preset range to determine whether the model performance meets the requirements.
[0087] If the evaluation index meets the preset requirements, that is, the evaluation index is within the preset range, the verification result is determined to be normal, indicating that the model performs well on unseen data and has reliable generalization ability. If the evaluation index does not meet the preset requirements, that is, the evaluation index is not within the preset range, the verification result is determined to be abnormal, indicating that the model may have overfitting, underfitting or other problems, and further adjustment and optimization are required.
[0088] If the verification result of the trained prediction model is normal, the deviation data of multiple indicators are predicted to obtain the deviation degree of the monitoring indicator.
[0089] Based on the above content, this application uses the LSTM prediction engine to analyze the collected multi-indicator deviation data through a complete process, including data preprocessing, model construction, model training, model evaluation and prediction steps, so as to obtain accurate health indicator deviations and provide a basis for system fault detection and fault tolerance.
[0090] In this application, the monitoring indicator deviation Mt is the MAE (Mean Absolute Error, used to calculate the monitoring indicator deviation) loss, that is, this application uses MAE loss as a measure of the monitoring indicator deviation, which can intuitively reflect the degree of deviation between the monitoring indicator and the normal state. In the health score function H(t), the larger Mt (calculated by MAE loss), the farther the monitoring indicator deviates from the normal range, and the higher the system health risk; conversely, the smaller Mt, the closer the monitoring indicator is to the normal state, and the system is relatively healthier. For example, when the MAE loss of GPU utilization continues to increase, it means that the GPU operating state is unstable and the system may have potential faults. At this time, the Mt item in the health score function will increase, lowering the system health score H(t), indicating that the system is abnormal.
[0091] In some embodiments, the merged log data is segmented, and the segmented merged log data is converted into a target format corresponding to a semantic analysis model to obtain target log data; log features of the target log data are extracted through the semantic analysis model, and log features marked as target tags in the log features are used as semantic coding features of the target log data; based on a preset neural network, the semantic coding features and log features are classified as abnormalities to determine the abnormality probability of the merged log data being in different abnormal categories; based on the abnormality probability, the log abnormality probability corresponding to the merged log data is determined.
[0092] In an optional embodiment of the present application, the log analysis tool in the present application may be a BERT classification model. The present application parses the log text through the BERT classification model to identify error patterns.
[0093] Specifically, this application uses a pre-trained BERT classification model to process the log text of the merged log data. First, the log text of the merged log data is segmented, and then the segmented results are converted into an input format acceptable to the model (such as token_ids, attention_mask, etc.). Next, these input data are input into the BERT classification model, the model will extract features from the text, and finally take out the output vector corresponding to the [CLS] tag as the semantic encoding vector of the log text. This vector contains the semantic information of the log text and can be used for subsequent anomaly detection and classification tasks.
[0094] Furthermore, the present application can construct a neural network including two layers of fully connected layers (MLP), that is, a preset neural network. Then, the log features and semantic coding features after template extraction and semantic encoding are used as input, and the first fully connected layer transforms and combines the input features to increase the nonlinear expression ability of the model. Then, the output of the first layer is processed by an activation function (such as ReLU), and the result is input to the second fully connected layer. The output of the second fully connected layer is processed by a softmax function to obtain the abnormal probability that each merged log data belongs to different abnormal categories (such as normal, connection abnormality, and other abnormalities).
[0095] This application can determine whether the merged log data is abnormal and the type of abnormality based on the probability value of the abnormal probability. In order to convert the logarithmic probability vector z into a probability distribution, this application needs to use the softmax function. The softmax function is defined as follows: For a vector z = [z1,z2 , ..., zc] of length C, the softmax function converts it into a probability distribution y = [y1,y2 , ..., yc].
[0096] It should be noted that if there are multiple anomalies in the merged log data, the anomaly probabilities corresponding to multiple anomaly categories will be obtained. At this time, the application can use the anomaly probability with the highest probability value as the log anomaly probability. Or different probability fusion weights are assigned to different anomaly categories, and each anomaly probability is weighted and fused with its probability fusion weight to obtain the final log anomaly probability.
[0097] In summary, this application can combine real-time time series indicator data and log data to achieve cross-modal analysis of time series and text, mine the spatiotemporal correlation between monitoring data mutations and log error messages, and improve the accuracy of fault diagnosis.
[0098] Figure 3 The flowchart of the third health assessment method proposed in this application is further shown. Figure 1 The embodiment shown further explains step 102. Figure 3 The following steps may be included.
[0099] Step 301, determining a first standard deviation of a monitoring indicator deviation in a preset period and a second standard deviation of a log abnormality probability in a preset period.
[0100] Step 302: Determine the index fusion weight and the log fusion weight based on the first standard deviation and the second standard deviation.
[0101] In some embodiments, the present application may determine an indicator fusion weight based on the first standard deviation and the second standard deviation, and then determine a log fusion weight according to the indicator fusion weight.
[0102] Specifically, the present application can calculate the indicator fusion weight and the log fusion weight through a weight adaptive formula.
[0103] The weight adaptation formula is: , .
[0104] in, The first standard deviation of the monitoring indicator deviation in a preset period (for example, the last 2 minutes) is is the second standard deviation of the log anomaly probability in a preset period (for example, the last 2 minutes), is the inverse of the first standard deviation, is the inverse of the second standard deviation.
[0105] In summary, this application designs a weight adaptive formula for dynamic weight fusion, dynamically determines the indicator fusion weight and log fusion weight, and more accurately evaluates the system health status.
[0106] Figure 4 The flowchart of the fourth health assessment method proposed in this application is further shown. Figure 1 The embodiment shown further explains step 103. Figure 4 The following steps may be included.
[0107] Step 401, determining a preset interval corresponding to a health score.
[0108] Step 402: Execute a recovery operation corresponding to a preset interval.
[0109] In some embodiments, if the preset interval is the first preset interval, non-core functions of the system to be evaluated are shut down; if the preset interval is the second preset interval, the system to be evaluated is restarted; if the preset interval is the third preset interval, the backup system corresponding to the system to be evaluated is determined, and the system to be evaluated is switched to the backup system.
[0110] Specifically, the present application can execute a hierarchical recovery strategy based on the health score results. When H(t) is in different intervals, the system will perform corresponding recovery operations. For example, when the health score H(t) belongs to the first preset interval (0.6, 0.8), the system downgrades the service; when the health score H(t) belongs to the second preset interval (0.8, 1.0), the node (system) is restarted; when the health score H(t) belongs to the third preset interval [1.0, +∞), the cluster (system) is switched.
[0111] Among them, the determination of non-core functions can be based on the importance and frequency of use of system functions. Functions that have little impact on the core business logic of the system and will not cause the overall failure of the system even if they are temporarily unavailable are defined as non-core functions, such as some auxiliary data analysis and non-real-time report generation functions in the system to be evaluated. The determination of the backup system requires comprehensive consideration of factors such as the system architecture, data consistency requirements, and switching costs. This application can deploy in advance a set of independent systems that are similar to the main system in architecture and functions and have a data synchronization mechanism as a backup system. When the main system has a serious failure or the health score is in the third preset interval, it can quickly switch to the backup system to ensure business continuity.
[0112] In summary, this application takes different recovery operations according to the health score results, establishes a three-level recovery mechanism driven by health score (service degradation, node restart, cluster switching), and achieves accurate resource scheduling and fault isolation by quantitatively evaluating the severity of the fault, significantly reducing the mean time to repair (MTTR).
[0113] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0114] The embodiment of the present application also provides a health assessment device 500, Figure 5 A structural diagram of a health assessment device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, including: The first determination unit 510 is used to determine the monitoring indicator deviation corresponding to the time series indicator data and the log anomaly probability corresponding to the log data, respectively, based on the time series indicator data and the log data in the system to be evaluated; The second determining unit 520 is used to determine the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log abnormality probability based on the monitoring indicator deviation and the log abnormality probability; The evaluation unit 530 is used to determine the health score of the system to be evaluated based on the monitoring indicator deviation, log abnormality probability, indicator fusion weight and log fusion weight, combined with the health score function, so as to perform a recovery operation on the system to be evaluated according to the health score.
[0115] The health evaluation device of the present application determines the monitoring indicator deviation corresponding to the time series indicator data and the log anomaly probability corresponding to the log data based on the time series indicator data and log data in the system to be evaluated; determines the indicator fusion weight corresponding to the monitoring indicator deviation and the log anomaly probability based on the monitoring indicator deviation and the log anomaly probability; determines the health score of the system to be evaluated based on the monitoring indicator deviation, the log anomaly probability, the indicator fusion weight and the log fusion weight, combined with the health score function, so as to perform recovery operations on the system to be evaluated based on the health score. It realizes more accurate detection and evaluation of system faults, avoids the problems of false alarms and missed alarms caused by the lack of flexibility of the fixed threshold monitoring method in the related technology, and can dynamically adjust the recovery strategy according to different fault severity and the actual system situation, significantly improving the recovery efficiency and fault tolerance of the system, and ensuring the stable and reliable operation of the system to be evaluated.
[0116] Furthermore, in a possible implementation method of the embodiment of the present application, the first determination unit 510 is used to: collect the time series indicator data according to the sampling frequency corresponding to the time series indicator data, and perform window normalization processing on the time series indicator data to obtain multi-indicator deviation data based on the time series indicator data after window normalization processing; obtain the log data generated by the system to be evaluated, and merge the log data based on the similarity of the log data to obtain merged log data; perform anomaly analysis on the multi-indicator deviation data and the merged log data to obtain the monitoring indicator deviation corresponding to the multi-indicator deviation data and the log anomaly probability corresponding to the merged log data.
[0117] Furthermore, in a possible implementation of the embodiment of the present application, the first determination unit 510 is used to: collect the timing indicator data according to the initial sampling frequency corresponding to the timing indicator data; if the timing indicator data is greater than or equal to the preset fault threshold corresponding to the timing indicator data, adjust the initial sampling frequency to collect the timing indicator data according to the adjusted sampling frequency; if the duration that the timing indicator data is less than the preset fault threshold is greater than or equal to the preset recovery period, restore the sampling frequency to the initial sampling frequency to collect the timing indicator data using the initial sampling frequency.
[0118] Furthermore, in a possible implementation of the embodiment of the present application, the first determination unit 510 is used to: collect the timing indicator data according to the initial sampling frequency corresponding to the timing indicator data; if the load of the system to be evaluated is greater than or equal to a preset load threshold, reduce the initial sampling frequency to collect the timing indicator data using the reduced sampling frequency.
[0119] Furthermore, in a possible implementation method of the embodiment of the present application, the first determination unit 510 is used to: for each sliding window corresponding to the sampling frequency, perform standardization processing on the timing indicator data to obtain standard timing indicator data; determine the indicator difference between the standard timing indicator data and the ideal data corresponding to the standard timing indicator data; based on the indicator difference and the preset indicator weight corresponding to the standard timing indicator data, perform weighted aggregation processing on the standard timing indicator data of each data category to obtain multi-indicator deviation data for each data category.
[0120] Furthermore, in a possible implementation of the embodiment of the present application, the first determination unit 510 is used to: obtain the log data generated by the system to be evaluated; parse the log data to extract keywords and parameter positions in the log data; based on the keywords and the parameter positions, determine the similarity corresponding to the log data, and merge the log data whose similarity is greater than or equal to a preset similarity threshold to obtain merged log data.
[0121] Furthermore, in a possible implementation method of the embodiment of the present application, the first determination unit 510 is used to: train the prediction model based on the training set and the verification set corresponding to the original multi-indicator deviation data, and verify the trained prediction model using the test set of the original multi-indicator deviation data to obtain a verification result; if the verification result of the prediction model is normal, input the multi-indicator deviation data into the trained prediction model to obtain the deviation degree of the proctoring indicator output by the trained prediction model.
[0122] Furthermore, in a possible implementation method of the embodiment of the present application, the first determination unit 510 is used to: pre-process the original multi-indicator deviation data to obtain a training set, a validation set and a test set corresponding to the multi-indicator deviation data; input the training set into the prediction model, and perform forward propagation to determine the loss value of the loss function of the prediction model; based on the loss value and the validation set, back-propagate to update the parameters of the prediction model until the loss function converges to obtain the trained prediction model; use the validation set data to verify the deviation of the monitoring indicator output by the trained prediction model to obtain a verification result; if the verification result is normal, input the test set into the trained prediction model to obtain the deviation of the monitoring indicator output by the trained prediction model.
[0123] Furthermore, in a possible implementation method of the embodiment of the present application, the first determination unit 510 is used to: perform word segmentation on the merged log data, and convert the merged log data after word segmentation into a target format corresponding to a semantic analysis model to obtain target log data; extract log features of the target log data through the semantic analysis model, and use the log features marked as target tags in the log features as semantic coding features of the target log data; based on a preset neural network, perform abnormality classification on the semantic coding features and the log features to determine the abnormality probability of the merged log data being in different abnormal categories; based on the abnormality probability, determine the log abnormality probability corresponding to the merged log data.
[0124] Furthermore, in a possible implementation method of the embodiment of the present application, the second determination unit 520 is used to: determine the first standard deviation of the monitoring indicator deviation in the preset time period and the second standard deviation of the log anomaly probability in the preset time period; based on the first standard deviation and the second standard deviation, determine the indicator fusion weight and the log fusion weight.
[0125] Furthermore, in a possible implementation of the embodiment of the present application, the evaluation unit 530 is used to: determine a preset interval corresponding to the health score; and perform a recovery operation corresponding to the preset interval.
[0126] Furthermore, in a possible implementation of the embodiment of the present application, the evaluation unit 530 is used to: if the preset interval is a first preset interval, shut down non-core functions of the system to be evaluated; if the preset interval is a second preset interval, restart the system to be evaluated; if the preset interval is a third preset interval, determine a backup system corresponding to the system to be evaluated, and switch the system to be evaluated to the backup system.
[0127] For the description of the features in the embodiment corresponding to the health assessment device, reference can be made to the relevant description of the embodiment corresponding to the health assessment method, which will not be repeated here.
[0128] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned health assessment method embodiments.
[0129] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned health assessment method embodiments when run.
[0130] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0131] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above health assessment method embodiments are implemented.
[0132] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned health assessment method embodiments are implemented.
[0133] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0134] The above is a detailed introduction to a health assessment method, electronic device, storage medium and product provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A health assessment method, characterized in that: The method comprises: Based on the time series indicator data and log data in the system to be evaluated, respectively determine the monitoring indicator deviation corresponding to the time series indicator data and the log anomaly probability corresponding to the log data; Based on the monitoring indicator deviation and the log abnormality probability, determine the indicator fusion weight corresponding to the monitoring indicator deviation and the log fusion weight corresponding to the log abnormality probability; According to the monitoring indicator deviation, the log abnormality probability, the indicator fusion weight and the log fusion weight, combined with the health score function, the health score of the system to be evaluated is determined, so as to perform a recovery operation on the system to be evaluated according to the health score.
2. The method according to claim 1, characterized in that The determining, based on the time series indicator data and the log data in the system to be evaluated, respectively the monitoring indicator deviation corresponding to the time series indicator data and the log abnormality probability corresponding to the log data comprises: The time series indicator data is collected according to the sampling frequency corresponding to the time series indicator data, and the time series indicator data is subjected to window normalization processing, so as to obtain multi-indicator deviation data according to the time series indicator data after the window normalization processing; Acquire log data generated by the system to be evaluated, and merge the log data based on similarity of the log data to obtain merged log data; Anomaly analysis is performed on the multi-indicator deviation data and the merged log data to obtain the monitoring indicator deviation degree corresponding to the multi-indicator deviation data and the log anomaly probability corresponding to the merged log data.
3. The method according to claim 2, characterized in that The collecting the timing indicator data according to the sampling frequency corresponding to the timing indicator data comprises: Collect the timing indicator data according to the initial sampling frequency corresponding to the timing indicator data; If the timing indicator data is greater than or equal to a preset fault threshold corresponding to the timing indicator data, adjusting the initial sampling frequency to collect the timing indicator data according to the adjusted sampling frequency; If the duration that the timing indicator data is less than the preset fault threshold is greater than or equal to a preset recovery period, the sampling frequency is restored to the initial sampling frequency to collect the timing indicator data using the initial sampling frequency.
4. The method according to claim 2, characterized in that: The collecting the timing indicator data according to the sampling frequency corresponding to the timing indicator data comprises: Collect the timing indicator data according to the initial sampling frequency corresponding to the timing indicator data; If the load of the system to be evaluated is greater than or equal to a preset load threshold, the initial sampling frequency is reduced to collect the timing indicator data using the reduced sampling frequency.
5. The method according to claim 2, characterized in that: The performing window normalization processing on the time series indicator data to obtain multiple indicator deviation data according to the time series indicator data after the window normalization processing includes: For each sliding window corresponding to the sampling frequency, the time series indicator data is standardized to obtain standard time series indicator data; Determine an indicator difference between the standard timing indicator data and ideal data corresponding to the standard timing indicator data; Based on the indicator difference and the preset indicator weights corresponding to the standard time series indicator data, weighted aggregation processing is performed on the standard time series indicator data of each data category to obtain multi-indicator deviation data of each data category.
6. The method according to claim 2, characterized in that The acquiring the log data generated by the system to be evaluated, and merging the log data based on the similarity of the log data to obtain the merged log data includes: Obtaining log data generated by the system to be evaluated; Parsing the log data to extract keywords and parameter positions in the log data; Based on the keyword and the parameter position, the similarity corresponding to the log data is determined, and the log data with similarity greater than or equal to a preset similarity threshold is merged to obtain merged log data.
7. The method according to claim 2, characterized in that The performing anomaly analysis on the multi-indicator deviation data and the merged log data to obtain the monitoring indicator deviation degree corresponding to the multi-indicator deviation data and the log anomaly probability corresponding to the merged log data comprises: The prediction model is trained based on the training set and the validation set corresponding to the original multi-index deviation data, and the trained prediction model is validated using the test set of the original multi-index deviation data to obtain a validation result; If the verification result of the prediction model is normal, the multi-index deviation data is input into the trained prediction model to obtain the deviation of the proctoring index output by the trained prediction model.
8. The method according to claim 7, characterized in that The prediction model is trained based on the training set and the validation set corresponding to the original multi-index deviation data, and the trained prediction model is validated using the test set of the original multi-index deviation data, and the validation results obtained include: Preprocessing the original multi-index deviation data to obtain a training set, a validation set, and a test set corresponding to the multi-index deviation data; Inputting the training set into the prediction model, and performing forward propagation to determine the loss value of the loss function of the prediction model; Based on the loss value and the validation set, back-propagation updates the parameters of the prediction model until the loss function converges to obtain the trained prediction model; The verification set data is used to verify the deviation of the monitoring index output by the trained prediction model to obtain a verification result.
9. The method according to claim 2, characterized in that: The performing anomaly analysis on the multi-indicator deviation data and the merged log data to obtain the monitoring indicator deviation degree corresponding to the multi-indicator deviation data and the log anomaly probability corresponding to the merged log data comprises: Segmenting the merged log data, and converting the segmented merged log data into a target format corresponding to a semantic analysis model to obtain target log data; Extracting log features of the target log data through the semantic analysis model, and using log features marked as target marks in the log features as semantic coding features of the target log data; Based on a preset neural network, the semantic coding features and the log features are classified as abnormal, and the abnormal probability of the merged log data being in different abnormal categories is determined; Based on the abnormal probability, a log abnormality probability corresponding to the merged log data is determined.
10. The method according to claim 1, characterized in that The determining, based on the monitoring indicator deviation and the log abnormality probability, of an indicator fusion weight corresponding to the monitoring indicator deviation and a log fusion weight corresponding to the log abnormality probability comprises: Determine a first standard deviation of the monitoring indicator deviation in a preset period and a second standard deviation of the log abnormality probability in the preset period; The indicator fusion weight and the log fusion weight are determined based on the first standard deviation and the second standard deviation.
11. The method according to claim 1, characterized in that: The performing a recovery operation on the system to be evaluated according to the health score includes: Determine a preset interval corresponding to the health score; Execute the recovery operation corresponding to the preset interval.
12. The method according to claim 11, characterized in that The performing of the recovery operation corresponding to the preset interval includes: If the preset interval is the first preset interval, shutting down non-core functions of the system to be evaluated; If the preset interval is the second preset interval, restarting the system to be evaluated; If the preset interval is the third preset interval, a backup system corresponding to the system to be evaluated is determined, and the system to be evaluated is switched to the backup system.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the health assessment method as described in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the health assessment method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the health assessment method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Performance index health degree monitoring method and device, equipment and storage medium
CN112801434A
System health degree intelligent monitoring and evaluation method based on anomaly detection
CN116383645A
Abnormal log detection method and device
CN118798167A
Computer-Implemented Systems And Methods For Loan Evaluation Using A Credit Assessment Framework
US20090299911A1
Method for health evaluation based on intelligent operation and maintenance scenarios, and device thereof
US20250036541A1
Cited By
Battery health degree evaluation method and system based on dual-channel attention mechanism
CN120779279A
A Battery Health Assessment Method and System Based on a Dual-Channel Attention Mechanism
CN120779279B
Log collection method and electronic equipment
CN120780575A
System health degree assessment method and system, terminal and medium
CN120822134A
Chip fault prediction method, device and equipment and computer readable medium
CN120822153A