IT resource health degree quantitative evaluation system and method based on business perspective
Through the IT resource health metric quantitative evaluation system based on the business perspective, quantitative mapping, dynamic baseline generation and root cause positioning analysis are used to solve the problem of low correlation between IT resource data and business demand analysis, and the efficient correlation between IT resource data and business demand is achieved, and the accuracy of problem source determination is improved.
Patent Information
- Application Number
- CN202510708086.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, IT resource data and business demand analysis have low correlation degree, and lack in-depth correlation analysis between business and IT resources, resulting in low correlation degree between IT resource data and business demand analysis.
Provides an IT resource health metric quantitative evaluation system based on a business perspective, including a business monitoring indicator quantification module, a dynamic baseline adaptive module and a root cause positioning module. Through quantitative mapping, dynamic baseline generation and root cause positioning analysis, the correlation between IT resource data and business needs is improved.
The correlation between IT resource data and business demand analysis has been improved, the fit of business monitoring indicators has been accurately evaluated, and timely measures have been taken to reduce the misreport rate at the source of problems and improve the accuracy of determining the source of problems.
Smart Images

Figure CN120276955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a quantitative evaluation system and method for the health of IT resources from a business perspective. Background Art
[0002] With the rapid development of information technology, the IT resources of enterprises have become the core assets supporting business operations and innovation. However, traditional IT resource management mainly focuses on the technical perspective and ignores the close combination with business goals. A quantitative evaluation system for the health of IT resources from a business perspective has emerged, aiming to evaluate and optimize the health status of IT resources starting from business requirements. By quantitatively evaluating the performance, availability, and reliability of IT resources, enterprises can more accurately identify potential risks and bottlenecks, ensure that the technical architecture can flexibly respond to the rapidly changing market environment, and promote the sustainable and healthy development of the business.
[0003] Existing quantitative evaluation technologies for the health of IT resources mostly focus on performance monitoring and fault warning at the technical level, such as indicators in aspects like server load, network bandwidth, and storage capacity. These methods mainly focus on the stability and efficiency of the IT system, but often ignore the matching degree between business requirements and technical resources. Although some systems analyze by integrating business indicators and technical data, their core is still technology-oriented and lacks in-depth consideration of business goals and strategies. The existing technologies cannot comprehensively reflect the actual support and contribution of IT resources to business operations, restricting the precise scheduling and optimization of IT resources by enterprises during the process of digital transformation. Therefore, a quantitative evaluation method for the health of IT resources from a business perspective urgently needs to be developed to ensure the deep integration of technology and business.
[0004] For example, the method and system for predicting and evaluating the health of business application system resources based on Informer disclosed in the patent application with the publication number of CN118034919A include: obtaining the host resource utilization data of the business application system and performing preprocessing to form data sets at hourly and minute levels, and the indicators include CPU utilization rate, memory indicators, disk occupancy size, network bandwidth data, and host status, etc.; performing batch normalization operation on the preprocessed data and introducing positional encoding to form a time series with positional encoding information; inputting the data into the Informer model, using the encoder to perform preliminary feature extraction through a probabilistic sparse attention mechanism, and using the decoder to synthesize the preliminary features and generate a prediction output; after performing batch normalization operation on the model output, converting it into the actual resource utilization rate value and evaluating and analyzing the system health based on this.
[0005] For example, a method for evaluating the health of a business system for IT centralized monitoring announced in the invention patent announcement with the announcement number of CN111274087B includes: obtaining the indicators of key business points of the business system for IT centralized monitoring, performing a health assessment according to preset data to be evaluated, the IT centralized monitoring server performing a weighted calculation on the scoring results of each data category, the health maintenance of the business system for IT centralized monitoring, re-evaluating and setting the health value of the business system for IT centralized monitoring, and daily health check of the hardware and software of the business system for IT centralized monitoring.
[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, it is found that the above technology has at least the following technical problems: In the prior art, since many health measurement systems set fixed thresholds and standards for dynamically changing business requirements and IT resources (such as considering resource tension when the CPU usage rate exceeds 80%), and often lack a more in-depth correlation analysis between business and IT resources, there is a problem of low correlation between IT resource data and business requirement analysis. Summary of the Invention
[0007] The embodiments of the present application provide a system and method for quantitatively evaluating the health of IT resources from a business perspective, solve the problem of low correlation between IT resource data and business requirement analysis in the prior art, and achieve an improvement in the correlation between IT resource data and business requirement analysis.
[0008] The embodiments of the present application provide a system for quantitatively evaluating the health of IT resources from a business perspective, including: a business monitoring index quantification module, a dynamic baseline adaptation module, and a root cause location module: among them, the business monitoring index quantification module is used to perform quantitative mapping on the obtained IT resource data to obtain business monitoring indexes, obtain business requirement data to obtain a business monitoring index fitting degree index, and take corresponding quantitative mapping measures; the dynamic baseline adaptation module is used to generate a baseline fluctuation range for each business monitoring index through dynamic baseline generation, obtain an alarm message according to each business monitoring index and the corresponding baseline fluctuation range, and determine whether to take system emergency measures; the root cause location module is used to obtain a root cause location data set based on the alarm message and the corresponding IT resource data, and obtain alarm evaluation data based on the root cause location data set for topology-driven analysis to obtain the problem source.
[0009] Further, the service demand data includes the service IT ratio, the service KPI correlation degree, and the number of service emergency measures; the service demand ratio represents the ratio of service growth to IT resource consumption, the service KPI correlation degree represents the correlation degree between service monitoring indicators and service KPIs, and the number of service emergency measures represents the number of abnormal events corresponding to service monitoring indicators with emergency measures; the alarm evaluation data includes the average alarm level, the average alarm frequency, the average alarm duration, the average baseline difference value, and the number of alarm trigger types; the average baseline difference value represents the average value of the deviation of the service monitoring indicator corresponding to the alarm information from the baseline fluctuation range.
[0010] Further, the process of obtaining the service monitoring indicator fitting degree index from the service demand data is as follows: Obtain the service fitting weights and the reference service busy degree from the preset database, where the service fitting weights include the service busy weight, the service availability weight, and the service health weight; if the service busy degree is not less than the corresponding reference service busy degree, perform a busy degree deviation operation on the service busy degree and the corresponding reference service busy degree to obtain a busy degree deviation, and perform a busy degree balance ratio operation on the busy degree deviation and the service IT ratio to obtain a first service fitting analysis indicator. The busy degree deviation operation is used to quantify the deviation degree of the service busy degree from the reference service busy degree, and the busy degree balance ratio operation is used to quantify the fitting degree of the service monitoring indicator from the perspective of the service busy degree; perform an availability balance ratio operation on the service KPI correlation degree and the service availability to obtain a second service fitting analysis indicator. The availability balance ratio operation is used to quantify the fitting degree of the service monitoring indicator from the perspective of the service availability; perform a normalization operation on the number of service emergency measures to obtain a service emergency degree, and perform an availability balance ratio operation on the service emergency degree and the service health degree to obtain a third service fitting analysis indicator. The normalization operation is used to quantify the response degree of the number of service emergency measures to abnormal events, and the availability balance ratio operation is used to quantify the fitting degree of the service monitoring indicator from the perspective of the service health degree; combine the first service fitting analysis indicator, the second service fitting analysis indicator, the third service fitting analysis indicator, and the corresponding service fitting weights to perform a service fitting indicator ratio distribution operation to obtain the service monitoring indicator fitting degree index. The service fitting indicator ratio distribution operation is used to perform a ratio distribution on the influence degree of the fitting degree of each service monitoring indicator and the IT resource data; if the service busy degree is less than the corresponding reference service busy degree, perform a secondary service fitting indicator ratio distribution operation based on the second service fitting analysis indicator, the third service fitting analysis indicator, and the corresponding service fitting weights to obtain the service monitoring indicator fitting degree index.
[0011] Further, the corresponding quantization mapping measures are taken, and the specific process is as follows: If the fitness index of the service monitoring index is not less than the fitness threshold, the dynamic baseline adaptive module continues to be executed; if the fitness index of the service monitoring index is less than the fitness threshold, the IT resource data is re-obtained for quantization mapping, and the corresponding quantization mapping measures are taken. The specific process is as follows: The difference between the fitness index of the service monitoring index and the fitness threshold is processed to obtain a fitness deviation. The difference processing represents a way to quantify the gap between the fitness index of the service monitoring index and the fitness threshold; a frequency mapping set between the fitness deviation and the frequency adjustment multiple is constructed, and the real-time fitness deviation is input into the frequency mapping set to obtain the corresponding frequency adjustment multiple speed, and the acquisition frequency of the IT resource data and the frequency adjustment multiple speed are subjected to a frequency multiple scaling operation to obtain an adjustment frequency. When the adjustment frequency is lower than the update frequency limit value, the adjustment frequency is used to obtain the IT resource data, otherwise the update frequency limit value is used; a weight mapping set between the fitness deviation and the data ratio adjustment value is constructed, and the real-time fitness deviation is input into the weight mapping set to obtain the corresponding data ratio adjustment value, and the acquisition weight of the IT resource data and the data ratio adjustment value are subjected to a ratio adjustment process to obtain an adjustment ratio, and the corresponding IT resource data is obtained based on the adjustment ratio; when the number of times of re-quantization mapping is higher than the adjustment times obtained from the preset database, data filling measures are taken.
[0012] Further, the alarm information is obtained according to the baseline fluctuation range corresponding to each service monitoring index, and the specific steps are as follows: The baseline fluctuation range corresponding to each service monitoring index is obtained through dynamic baseline generation, and each service monitoring index is compared with the baseline fluctuation range: If the service monitoring index is within the corresponding baseline fluctuation range, no alarm information is generated; if the service monitoring index is not within the corresponding baseline fluctuation range, the corresponding alarm information is generated, and the alarm level of the alarm information is judged; if the alarm level is higher than the preset alarm level, emergency processing is carried out according to the alarm level corresponding to the alarm information, otherwise problem reliability analysis is carried out. The emergency processing is used to automatically provide corresponding system emergency measures for the corresponding alarm level.
[0013] Further, the specific process of the problem reliability analysis is as follows: A root cause location data set is obtained based on the alarm information corresponding to all service monitoring indexes and the corresponding IT resource data; if the data volume of the root cause location data set is higher than the preset data volume, the root cause location data set is automatically subjected to topology-driven analysis to obtain the problem root cause, otherwise the alarm evaluation data is obtained; the corresponding alarm warning evaluation index is obtained based on the alarm evaluation data, and the alarm warning evaluation index is compared with the warning threshold obtained from the preset database: If the alarm warning evaluation index is less than the warning threshold, the alarm information continues to be obtained; if the alarm warning evaluation index is not less than the warning threshold, the root cause reliability index is obtained to judge whether to perform topology-driven analysis.
[0014] Further, the specific process of obtaining the alarm warning evaluation index is as follows: Obtain the alarm evaluation weights from a preset database, where the alarm evaluation weights include alarm level weights, alarm frequency weights, alarm time weights, baseline difference weights, and alarm type weights; perform data preprocessing on the alarm evaluation data, and the data preprocessing is used to de-unify the alarm evaluation data; obtain the baseline fluctuation range, where the baseline fluctuation range represents the range corresponding to the maximum value and the minimum value of the baseline fluctuation range; perform a range deviation operation on each service monitoring indicator and the corresponding baseline fluctuation range to obtain an average baseline proximity difference, and the range deviation operation is used to obtain data measuring the degree of deviation between the service monitoring indicator and the baseline fluctuation range; perform a trend fitting quantization operation on the alarm evaluation data to obtain a corresponding alarm quantization value, and the trend fitting quantization operation is used to unify the trend of each alarm evaluation data and obtain the corresponding quantization data; perform a proportion distribution and balancing operation based on the alarm quantization value and the corresponding alarm evaluation weights to obtain the alarm warning evaluation index, where the alarm warning evaluation index represents data used to quantify the risk degree of the alarm information in the root cause location dataset, and the proportion distribution and balancing operation is used to quantify the risk degree of the alarm information in the root cause location dataset reflected by the alarm warning evaluation data.
[0015] Further, the specific process of the root cause reliability index is as follows: Obtain the root cause analysis weights from a preset database, where the root cause analysis weights include data weights and index weights; perform a data deviation operation on the data volume of the root cause location dataset and the preset data volume to obtain a data deviation value, and the data deviation operation is used to quantify the reverse deviation degree of the data volume of the root cause location dataset from the preset data volume; perform an index deviation operation on the alarm warning evaluation index and the warning threshold to obtain an index deviation value, and the index deviation operation is used to quantify the positive deviation degree of the alarm warning evaluation index from the warning threshold; perform a root cause analysis proportion distribution operation based on the data deviation value, the index deviation value, and the corresponding root cause analysis weights to obtain the root cause reliability index, where the root cause analysis proportion distribution operation is a processing method for comprehensively quantifying the reliability degree of the problem source obtained by topological-driven analysis based on the current root cause location dataset, and the root cause reliability index is used to quantify the reliability degree of the problem source obtained by topological-driven analysis based on the current root cause location dataset.
[0016] Further, the specific process of determining whether to perform topology-driven analysis is as follows: construct a mapping set between the root cause reliability index and the reliability probability, input the real-time root cause reliability index into the mapping set to obtain the corresponding reliability probability; compare the reliability probability with the preset probability obtained from the preset database: if the reliability probability is not less than the preset probability, perform topology-driven analysis on the current root cause location data set to obtain the problem source and feedback the problem source to the preset staff; if the reliability probability is less than the preset probability, continue to obtain alarm information.
[0017] In the embodiments of the present application, a method for quantitatively evaluating the health of IT resources from a business perspective is provided, and the specific steps are as follows: perform quantitative mapping on the obtained IT resource data to obtain business monitoring indicators, obtain the business requirement data to obtain the business monitoring indicator fitting degree index to take corresponding quantitative mapping measures; generate a dynamic baseline to obtain the baseline fluctuation range of each business monitoring indicator, obtain alarm information based on each business monitoring indicator and the corresponding baseline fluctuation range, and determine whether to take system emergency measures; obtain the root cause location data set based on the alarm information and the corresponding IT resource data, and obtain alarm evaluation data based on the root cause location data set to perform topology-driven analysis to obtain the problem source.
[0018] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Perform quantitative mapping on the obtained IT resource data to obtain business monitoring indicators, and then combine the obtained business requirement data to obtain the business monitoring indicator fitting degree index to take corresponding quantitative mapping measures. Then, generate a dynamic baseline to obtain the baseline fluctuation range of each business monitoring indicator, and then obtain alarm information based on each business monitoring indicator, and determine whether to take system emergency measures. Finally, obtain the root cause location data set based on the alarm information and the corresponding IT resource data to obtain alarm evaluation data, and perform topology-driven analysis to obtain the problem source, thereby increasing the association degree between the determination of IT resource problems and business requirements, and further realizing the improvement of the association degree between IT resource data and business requirement analysis, effectively solving the problem of low association degree between IT resource data and business requirement analysis in the prior art; 2. By obtaining the business fit weight and reference business busyness from a preset database, if the business busyness is not less than the reference business busyness, then the first business fit analysis indicator is obtained based on the business busyness, reference business busyness, and business IT ratio. Then, the second business fit analysis indicator and the third business fit analysis indicator are obtained in sequence based on the data required for each business and business monitoring indicators. Finally, the business monitoring indicator fit index is obtained by combining the above-obtained business fit analysis indicators and business fit weights. Otherwise, the business monitoring indicator fit index is obtained through the second business fit analysis indicator and the third business fit analysis indicator, thus more accurately evaluating the fit of business monitoring indicators, and then achieving timely adoption of corresponding measures to obtain more fitting business monitoring indicators; 3. By obtaining the alarm evaluation weight from a preset database and performing data preprocessing on the alarm evaluation data, then obtaining the baseline fluctuation range and getting the average baseline proximity difference between each business monitoring indicator and the corresponding baseline fluctuation range. Then, trend fitting quantization operations are performed on the alarm evaluation data to obtain the corresponding alarm quantization value. Finally, the alarm warning evaluation index is obtained based on the alarm quantization value and the corresponding alarm evaluation weight, thus more accurately evaluating the alarm risk degree of the current root cause location data set, and then achieving a more comprehensive problem analysis of alarm information; 4. If the business monitoring indicator fit index is not less than the fit threshold, then the dynamic baseline adaptive module is continued to be executed. Otherwise, IT resource data is re-obtained for quantization mapping, and corresponding quantization mapping measures are taken. First, the fit deviation is obtained based on the business monitoring indicator fit index and the fit threshold, and a frequency mapping set between the fit deviation and the frequency adjustment multiple is constructed to obtain the frequency adjustment multiple. Thus, the adjustment frequency is obtained by combining the acquisition frequency of IT resource data. When the adjustment frequency is lower than the update frequency limit value, the adjustment frequency is used to obtain IT resource data. Otherwise, the update frequency limit value is used. Then, a weight mapping set between the fit deviation and the data ratio adjustment value is constructed to obtain the data ratio adjustment value, and the adjustment ratio is obtained by combining the acquisition weight of IT resource data to obtain IT resource data. At the same time, when the number of re-quantization mappings is higher than the adjustment times obtained from the preset database, then data filling measures are taken, thus achieving a more accurate optimization of IT resource data acquisition, and then obtaining more fitting business monitoring indicators; 5. By constructing a mapping set between the root cause reliability index and the reliability probability, and inputting the real-time root cause reliability index into the mapping set to obtain the corresponding reliability probability. Then, the reliability probability is compared with the preset probability obtained from the preset database. If the reliability probability is not less than the preset probability, then topological-driven analysis is performed on the current root cause location data set to obtain the problem source and the problem source is fed back to the preset staff. Otherwise, alarm information is continuously obtained, thus further ensuring the accuracy of problem source determination, and then reducing the misreport rate of the problem source. Description of the Drawings
[0019] Figure 1 It is a schematic structural diagram of a quantitative evaluation system for the health of IT resources from the perspective of business provided by an embodiment of the present application. Detailed Implementation Manner
[0020] By providing a quantitative evaluation system and method for the health of IT resources from the perspective of business, the embodiment of the present application solves the problem of low correlation between IT resource data and business requirement analysis in the prior art. By performing quantitative mapping on the obtained IT resource data to obtain business monitoring indicators, and by obtaining the business compliance weight and reference business busyness from a preset database, if the business busyness is not less than the reference business busyness, then according to the business busyness, the reference business busyness and the business IT ratio, a first business compliance analysis indicator is obtained. Then, based on the data required for each business and the business monitoring indicators, a second business compliance analysis indicator and a third business compliance analysis indicator are obtained in sequence. Then, by combining the above-obtained business compliance analysis indicators and the business compliance weight, a business monitoring indicator compliance index is obtained. Otherwise, the business monitoring indicator compliance index is obtained through the second business compliance analysis indicator and the third business compliance analysis indicator, and corresponding quantitative mapping measures are taken. Then, the baseline fluctuation range of each business monitoring indicator is generated through dynamic baseline, and based on this, alarm information is obtained for each business monitoring indicator, and it is judged whether to take system emergency measures. Finally, a root cause location data set is obtained based on the alarm information and the corresponding IT resource data to obtain alarm evaluation data. By obtaining the alarm evaluation weight from a preset database and performing data preprocessing on the alarm evaluation data, then obtaining the baseline fluctuation range and obtaining the average baseline proximity difference between each business monitoring indicator and the corresponding baseline fluctuation range, and then performing trend fitting quantization operation on the alarm evaluation data to obtain the corresponding alarm quantization value. At the same time, according to the alarm quantization value and the corresponding alarm evaluation weight, an alarm early warning evaluation index is obtained to perform topology-driven analysis to obtain the problem source, realizing the improvement of the correlation between IT resource data and business requirement analysis.
[0021] The technical solution in the embodiment of the present application is to solve the problem of low correlation between IT resource data and business requirement analysis, and the overall idea is as follows: By performing quantitative mapping on IT resource data to obtain business monitoring indicators, and based on this, combining with business requirement data to obtain a business monitoring indicator compliance index to take quantitative mapping measures. Then, the baseline fluctuation range of each business monitoring indicator is generated through dynamic baseline and combined with each business monitoring indicator to obtain alarm information. Then, it is judged whether to take system emergency measures. Finally, a root cause location data set is obtained based on the alarm information and IT resource data to obtain alarm evaluation data, and topology-driven analysis is performed to obtain the problem source, achieving the effect of improving the correlation between IT resource data and business requirement analysis.
[0022] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0023] As Figure 1 shown, it is a schematic structural diagram of an IT resource health quantification evaluation system based on the business perspective provided by an embodiment of the present application. The IT resource health quantification evaluation system based on the business perspective provided by an embodiment of the present application includes: a business monitoring index quantification module, a dynamic baseline adaptation module, and a root cause location module. Among them, the business monitoring index quantification module is used to perform quantization mapping on the obtained IT resource data to obtain business monitoring indexes, obtain the business monitoring index compliance index from the obtained business requirement data to take corresponding quantization mapping measures; the dynamic baseline adaptation module is used to generate a baseline fluctuation range for each business monitoring index through dynamic baseline generation, obtain an alarm message based on each business monitoring index and the corresponding baseline fluctuation range, and determine whether to take system emergency measures; the root cause location module is used to obtain a root cause location data set based on the alarm message and the corresponding IT resource data, and obtain alarm evaluation data based on the root cause location data set for topology-driven analysis to obtain the problem source.
[0024] In this embodiment, the business perspective means evaluating and managing IT resources from the business requirements and business goals. For example, in an online education system, the core business goal is to provide a stable and efficient online learning experience. Then the business goal is to ensure that students can smoothly watch video courses and can quickly participate in online exams. The business indicators are video loading time, course viewing fluency, student concurrent access volume, etc. The online education system will monitor the health status of the video stream service (such as the CPU and memory usage of the server, network bandwidth, etc.) and establish an association between these technical indicators and the video loading time and viewing fluency. The online education system will judge whether the server resources are sufficient to support the concurrent access volume of the platform and the transmission quality of the video stream. When the system finds that the CPU usage of the server is too high or the bandwidth is approaching saturation, the online education system will evaluate the impact of these IT problems on the video playback experience. If it affects the viewing experience of a large number of students, the online education system will trigger an alarm to remind the preset IT personnel to take emergency measures. If an alarm occurs, the online education system will not only give a technical alarm such as "insufficient server resources", but will also propose the impact at the business level, such as students may experience stuttering when watching videos, which affects the learning progress and may reduce customer satisfaction. Therefore, the preset IT personnel will receive more specific business-related warnings and can handle them targeted.
[0025] Establish the relationship between IT resource data and business requirement metrics through machine learning, especially deep learning and reinforcement learning. For example, time series modeling techniques such as LSTM (Long Short-Term Memory) have realized the analysis and prediction of the dynamic relationship between IT resource usage and business performance, and then obtained quantitative metrics such as business busyness (TPS, Transactions Per Second), business availability (SLA compliance rate, service level agreement), and business health (abnormal event density); similarly, by using LSTM (Long Short-Term Memory) to model time series data and training the LSTM model with historical business data, based on the trained LSTM model, the LSTM model is used to predict real-time data to generate dynamic baseline values. For example, LSTM will generate an expected health value (i.e., dynamic baseline) or a prediction range (i.e., dynamic baseline range) for each business monitoring metric; Topology-driven analysis is specifically implemented through the correlation analysis algorithm between business topology and alarms. When the business monitoring metrics and dynamic baseline generation technology provide data and alarm information that meet the preset data volume, topology-driven analysis is used to analyze the root cause and obtain the problem source; Through the implementation of this system, it helps to more accurately evaluate the health of IT resources from the business perspective, is also conducive to more accurately determining the root cause of problems, and at the same time realizes the improvement of the correlation between IT resource data and business requirement analysis.
[0026] Furthermore, business requirement data includes business IT ratio, business KPI correlation degree, and the number of business emergency measures; the business requirement ratio represents the ratio of business growth to IT resource consumption, the business KPI correlation degree represents the correlation between business monitoring metrics and business KPIs, and the number of business emergency measures represents the number of abnormal events corresponding to business monitoring metrics with emergency measures; alarm evaluation data includes average alarm level, average alarm frequency, average alarm duration, average baseline difference value, and the number of alarm trigger types; the average baseline difference value represents the average value of the deviation difference between the business monitoring metric corresponding to the alarm information and the baseline fluctuation range.
[0027] In this embodiment, the business IT ratio is obtained by performing a ratio operation on the statistically obtained business growth volume and IT resource consumption (such as CPU, memory, and bandwidth, etc.). The number of business KPIs associated with the business monitoring metrics is counted and a ratio operation is performed with the total number of business KPIs to obtain the business KPI correlation degree. The number of abnormal events corresponding to the business monitoring metrics with emergency measures is counted to obtain the business emergency measure quantity; the average alarm level is obtained by performing an average operation on the alarm levels corresponding to all alarm messages in the root cause location dataset, the average alarm frequency is obtained by performing an average operation on the alarm frequencies of all alarm messages in the root cause location dataset, the average alarm duration is obtained by performing an average operation on the alarm durations of all alarm messages in the root cause location dataset, and the average baseline difference value is obtained by performing an average operation on the difference between the business monitoring metrics corresponding to all alarm messages in the root cause location dataset and the baseline dynamic range. The number of alarm trigger types is counted for all alarm messages corresponding to the root cause location dataset. The alarm level, alarm frequency, alarm duration, and alarm type are all recorded in the corresponding system logs; through the acquisition of the above data, a data basis is provided for the quantitative evaluation of IT resource health, ensuring the accuracy of subsequent analysis.
[0028] Further, the business monitoring metric compliance index is obtained by acquiring business requirement data, and the specific process is as follows: The business compliance weights and reference business busyness are obtained from a preset database. The business compliance weights include business busyness weight, business availability weight, and business health weight; classification analysis is performed based on the business busyness and the corresponding reference business busyness. If the business busyness is not less than the corresponding reference business busyness, a busyness deviation operation is performed on the business busyness and the corresponding reference business busyness to obtain the busyness deviation (i.e., ), and a busyness balance ratio operation is performed on the busyness deviation and the business IT ratio to obtain the first business compliance analysis index (i.e., ). The busyness deviation operation is used to quantify the deviation degree of the business busyness from the reference business busyness, and the busyness balance ratio operation is used to quantify the compliance degree of the business monitoring metrics from the perspective of business busyness; an availability balance ratio operation is performed on the business KPI correlation degree and business availability to obtain the second business compliance analysis index (i.e., ). The availability balance ratio operation is used to quantify the compliance degree of the business monitoring metrics from the perspective of business availability; a normalization operation is performed on the number of business emergency measures to obtain the business emergency degree, i.e., , and an availability balance ratio operation is performed on the business emergency degree and business health degree to obtain the third business compliance analysis index (i.e., ), the normalization operation is used to quantify the response degree of the number of business emergency measures to abnormal events, and the availability balance ratio operation is used to quantify the degree of fit of business monitoring indicators from the perspective of business health; combining the first business fit analysis indicator, the second business fit analysis indicator, the third business fit analysis indicator, and the corresponding business fit weights to perform a business fit indicator ratio allocation operation to obtain a business monitoring indicator fit index. The business fit indicator ratio allocation operation is used to allocate proportions to the influence degrees of the fit degrees of each business monitoring indicator and IT resource data. The business monitoring indicator fit index represents the quantitative data of the fit degree of the business monitoring indicator and IT resource data; if the business busyness is less than the corresponding reference business busyness, then based on the second business fit analysis indicator, the third business fit analysis indicator, and the corresponding business fit weights, perform a secondary business fit indicator ratio allocation operation to obtain a business monitoring indicator fit index. The secondary business fit indicator ratio allocation operation is used to allocate proportions to the influence degrees of the fit degrees of the second business fit analysis indicator and the third business fit analysis indicator and the corresponding IT resource data. The business monitoring indicators include business busyness, business availability, and business health.
[0029] The specific limiting expression of the business monitoring indicator fit index is as follows: ; In the formula, represents the business IT ratio, represents the business KPI correlation degree, represents the business emergency degree, represents the business busyness, represents the business availability, represents the business health, represents the reference business busyness, represents the business busyness weight, represents the business availability weight, represents the business health weight, represents the business monitoring indicator fit index.
[0030] In this embodiment, the algorithm combines business monitoring indicators, business requirement data, reference business busyness, and the corresponding business fit weights for comprehensive analysis to obtain a business monitoring indicator fit index. In the formula, by comparing the business busyness with the reference busyness, it is divided into two analysis cases, and the value range of the business IT ratio is (0, 1). Therefore, the difference between the value 1 and the business IT ratio (i.e., ) is always not zero.
[0031] When the business busyness is not greater than the reference business busyness, as the first business fitting analysis index, the second business fitting analysis index, and the third business fitting analysis index increase, the corresponding business monitoring index fitting degree index also increases, indicating that the fitting degree between the business monitoring index and the business requirements is higher; among them, when the business busyness is closer to the reference business busyness, the corresponding busyness deviation is smaller, and when the difference between the business IT ratio and the value 1 is smaller, it indicates that the IT resource data and the business growth are more balanced. Then, the larger the ratio of the busyness deviation to the difference between the business IT ratio and the value 1, the higher the IT resource health degree; when the business emergency degree is smaller, it indicates the availability of the known business for abnormal events, and when the business KPI correlation degree is higher, then the corresponding second business fitting analysis index is larger, indicating a higher IT resource availability; when the business emergency degree is larger and the business health degree is lower, the corresponding third business fitting analysis index is larger, indicating a higher healthy performance degree of the system in processing the business.
[0032] When the business busyness is greater than the reference business busyness, the first business fitting analysis index is assigned a value of 0 to obtain the corresponding business monitoring index fitting degree index; and the reference business busyness is specifically set by the preset IT personnel according to the actual IT resource situation; through the analysis of the business monitoring index fitting degree index, it is helpful to more accurately evaluate the compliance degree between the business monitoring index obtained by quantifying and mapping the IT business resources and the business requirements, so as to take measures in a timely manner to remap the business monitoring index that does not meet the business requirements, and then obtain a business monitoring index that is more in line with the business requirements, providing more accurate analysis data for the subsequent analysis based on the business monitoring index.
[0033] Specifically, the business fitting weight is obtained from the preset database, and the business fitting weight represents the influence degree of the business requirement data on the business monitoring index fitting degree index. There is a unique mapping relationship between each business requirement data and the business fitting weight, and the value range is between 0 and 1; for example, a mapping set of the business requirement data and the preset business fitting weight is constructed, and the real-time business IT ratio, business KPI correlation degree, and business emergency measure quantity are input into the mapping set to obtain the business busyness weight, business availability weight, and business health weight respectively, which represent the influence degrees of the business IT ratio, business KPI correlation degree, and business emergency measure quantity on the business monitoring index fitting degree index, and the sum of the three is 1.
[0034] Further, corresponding quantization mapping measures are taken, and the specific process is as follows: Compare the compliance index of the business monitoring metrics with the compliance threshold obtained from the preset database: If the compliance index of the business monitoring metrics is not less than the compliance threshold, continue to execute the dynamic baseline adaptive module; If the compliance index of the business monitoring metrics is less than the compliance threshold, re-obtain the IT resource data for quantization mapping and take corresponding quantization mapping measures. The specific process is as follows: Perform difference processing on the compliance index of the business monitoring metrics and the compliance threshold to obtain a compliance deviation. Difference processing represents a way to quantify the gap between the compliance index of the business monitoring metrics and the compliance threshold; Construct a frequency mapping set between the compliance deviation and the frequency adjustment multiple, input the real-time compliance deviation into the frequency mapping set to obtain the corresponding frequency adjustment multiple speed, and perform a frequency multiple scaling operation on the acquisition frequency of the IT resource data and the frequency adjustment multiple speed to obtain an adjusted frequency. When the adjusted frequency is lower than the update frequency limit value, use the adjusted frequency to obtain the IT resource data, otherwise use the update frequency limit value; Construct a weight mapping set between the compliance deviation and the data ratio adjustment value, input the real-time compliance deviation into the weight mapping set to obtain the corresponding data ratio adjustment value, and perform a ratio adjustment process on the acquisition weight of the IT resource data and the data ratio adjustment value to obtain an adjusted ratio, and obtain the corresponding IT resource data based on the adjusted ratio; When the number of times of re-quantization mapping is higher than the adjustment times obtained from the preset database, then take data filling measures. The data filling measures represent sending a historical data filling request to the preset staff, and the historical data represents the IT resource data in the same time period in the historical data.
[0035] In this embodiment, the difference processing represents performing a difference operation on the compliance index of the business monitoring metrics and the compliance threshold. The frequency multiple scaling operation represents performing a multiplication operation on the acquisition frequency of the IT resource data and the frequency adjustment multiple speed. The ratio adjustment process represents performing a sum operation on the acquisition weight of the IT resource data and the data ratio adjustment value. Among them, the acquisition weight represents the proportion of the data volume obtained from the front end and the back end of the IT resource data, and by default, the sum operation is performed on the front-end data ratio and the data ratio adjustment value, and the difference operation is performed on the back-end data ratio and the data ratio adjustment value; Also, the update frequency limit value and the adjustment times are both set by the preset IT personnel according to the specific situation of the IT resource data; Through the quantization mapping measures, business monitoring metrics with a higher degree of compliance with business requirements are obtained, laying a more accurate data foundation for subsequent analysis.
[0036] Specifically, the compliance threshold is obtained from the preset database. In a specific embodiment, substitute the business requirement data and business monitoring metrics corresponding to the accurate problem root cause obtained from the historical data into the specific limit expression of the business monitoring metrics compliance index to obtain the corresponding data set, and perform a mean operation on the data set to obtain the compliance threshold.
[0037] Furthermore, alarm information is obtained based on the fluctuation ranges of each service monitoring metric and their corresponding baselines. The specific steps are as follows: The fluctuation ranges of each service monitoring metric corresponding to the baselines are obtained through dynamic baseline generation, and each service monitoring metric is compared with the fluctuation range of the baseline. If the service monitoring metric is within the corresponding baseline fluctuation range, no alarm information is generated. If the service monitoring metric is not within the corresponding baseline fluctuation range, the corresponding alarm information is generated, and the alarm level of the alarm information is judged. If the alarm level is higher than the preset alarm level, emergency handling is performed according to the alarm level corresponding to the alarm information; otherwise, problem reliability analysis is carried out. Emergency handling is used to automatically provide corresponding system emergency measures for the corresponding alarm level.
[0038] In this embodiment, the alarm levels include the information level (i.e., level one), the warning level (i.e., level two), the emergency level (i.e., level three), and the critical level (i.e., level four). Usually, the second level is set as the preset alarm level. Among them, the information level is usually some non-urgent notifications used to monitor the health status or changes of the monitoring system. The warning level indicates that the value of a certain service monitoring metric is close to the dynamic baseline range, which may cause problems but does not currently affect the service. The emergency level indicates that the value of a certain service monitoring metric has exceeded the dynamic baseline range, and the system may have problems, which may lead to service failures. The critical level indicates that the system has serious failures, which have led to service interruptions or major business problems.
[0039] For alarm information exceeding the preset alarm level, corresponding emergency handling is taken. The emergency handling that can be taken for the emergency level includes, but is not limited to, notifying the preset IT personnel, continuously monitoring the progress of the problem, and starting the preset emergency response process. The emergency handling that can be taken for the critical level includes, but is not limited to, immediately notifying the entire team, including management. Usually, a full-team emergency response needs to be initiated, the system needs to be comprehensively checked, critical services need to be restored, and external services need to be suspended. Through the above analysis, it is further ensured that problems can be monitored more timely and corresponding emergency handling measures can be taken when problems occur, reducing the time and cost of continuous system failures.
[0040] Further, the specific process for problem reliability analysis is as follows: Obtain a root cause location data set based on the alarm information corresponding to all business monitoring metrics and the corresponding IT resource data; Compare the data volume of the root cause location data set with the preset data volume obtained from a preset database: If the data volume of the root cause location data set is higher than the preset data volume, automatically perform topology-driven analysis on the root cause location data set to obtain the problem root cause, otherwise obtain alarm evaluation data; Obtain the corresponding alarm warning evaluation index based on the alarm evaluation data, and compare the alarm warning evaluation index with the warning threshold obtained from the preset database: If the alarm warning evaluation index is less than the warning threshold, continue to obtain alarm information; If the alarm warning evaluation index is not less than the warning threshold, obtain the root cause reliability index to determine whether to perform topology-driven analysis.
[0041] In this embodiment, the preset data volume is set by preset IT personnel according to the specific IT resource data situation; The data in the root cause location data set includes data related to alarm information, the corresponding business monitoring metrics, and the corresponding dynamic baseline range, and the IT resource data corresponding to the business monitoring metrics; When a problem fault occurs, the process of this embodiment not only focuses on the data volume of the root cause location data set, but also makes a dual judgment based on the alarm warning evaluation index and the root cause reliability index to ensure that the finally obtained problem source has higher reliability, and by timely discovering and accurately analyzing the problem root cause, it can effectively reduce the long-tail impact of system failures, improve the stability and reliability of the system, and at the same time can also convert business monitoring metrics and alarm information into effective data information.
[0042] Specifically, the warning threshold is obtained from the preset database. In a specific embodiment, the alarm evaluation data corresponding to the accurate problem root cause obtained from historical data is substituted into the specific limit expression of the alarm warning evaluation index to obtain a corresponding data set, and the mean value operation is performed on the data set to obtain the warning threshold.
[0043] Further, the specific process for obtaining the alarm warning evaluation index is as follows: Obtain the alarm evaluation weights from the preset database, and the alarm evaluation weights include alarm level weight, alarm frequency weight, alarm time weight, baseline difference weight, and alarm type weight; Perform data preprocessing on the alarm evaluation data, and the data preprocessing is used to de-unitize the alarm evaluation data; Obtain the baseline fluctuation range, and the baseline fluctuation range represents the range corresponding to the maximum value and the minimum value of the baseline fluctuation range; Perform range deviation operation on each business monitoring metric and the corresponding baseline fluctuation range to obtain the average baseline proximity difference (i.e., )), and the range deviation operation is used to obtain data measuring the deviation degree between the business monitoring metric and the baseline fluctuation range; Perform trend fitting quantization operation on the alarm evaluation data to obtain the corresponding alarm quantization value (i.e., ), the trend fitting quantization operation is used to uniformly trend each alarm evaluation data and obtain the corresponding quantization data; according to the alarm quantization value and the corresponding alarm evaluation weight, a proportion allocation and balancing operation is performed to obtain an alarm early warning evaluation index, where the alarm early warning evaluation index represents the data used to quantify the risk degree of the alarm information in the root cause location dataset, and the proportion allocation and balancing operation is used to quantify the risk degree of the alarm information in the root cause location dataset reflected by the alarm early warning evaluation data.
[0044] The specific limit expression of the alarm early warning evaluation index is: ; ; ; In the formula, i represents the number of the alarm evaluation data type, , represents the alarm evaluation data of the i-th type, n represents the number of the alarm information, , represents the total number of the alarm information, represents the alarm evaluation weight, represents the alarm evaluation data of the 4th type corresponding to the nth alarm information, represents the maximum value of the dynamic baseline range corresponding to the nth alarm information, represents the minimum value of the dynamic baseline range corresponding to the nth alarm information, represents the alarm early warning evaluation index.
[0045] In this embodiment, the algorithm combines the alarm evaluation data, the dynamic baseline range, and the alarm evaluation weight for comprehensive analysis to obtain the alarm early warning evaluation index. In the formula, represents the alarm evaluation data of the 1st type, i.e., the average alarm level, represents the alarm evaluation data of the 2nd type, i.e., the average alarm frequency, represents the alarm evaluation data of the 3rd type, i.e., the average alarm duration, represents the alarm evaluation data of the 4th type, i.e., the average baseline difference value, It represents the alarm evaluation data of the 5th type, that is, the number of alarm trigger types. As the alarm evaluation data increases, the corresponding alarm warning evaluation index becomes larger, indicating that the risk degree of the alarm information in the current root cause location dataset is higher. Among them, in the operation of the average baseline difference value, when the business monitoring index corresponding to the alarm information is higher than the maximum value of the dynamic baseline range or lower than the minimum value of the dynamic baseline range, the corresponding average baseline difference value is larger, indicating that the business monitoring index deviates more from the dynamic baseline range; through the analysis of the alarm warning evaluation index, not only the alarm prevention for the root cause location dataset that does not meet the preset data volume is carried out, but also the reliability of the problem root cause obtained from the root cause location dataset that does not meet the preset data volume is analyzed, reducing the false alarm rate of the problem and improving the accuracy of determining the problem root cause.
[0046] Specifically, the alarm evaluation weight is obtained from the preset database, and the alarm evaluation weight represents the influence degree of the alarm evaluation data on the alarm warning evaluation index. There is a unique mapping relationship between the alarm evaluation data and the alarm evaluation weight, and the value range is between 0 and 1; for example, a mapping set of the alarm evaluation data and the preset alarm evaluation weight is constructed, and the real-time average alarm level, average alarm frequency, average alarm duration, average baseline difference value, and the number of alarm trigger types are input into the mapping set to obtain the alarm level weight, alarm frequency weight, alarm time weight, baseline difference weight, and alarm type weight respectively, which represent the influence degrees of the average alarm level, average alarm frequency, average alarm duration, average baseline difference value, and the number of alarm trigger types on the alarm warning evaluation index, and the sum of the five is 1.
[0047] Furthermore, the specific process of the root cause reliability index is as follows: obtain the root cause analysis weight from the preset database, and the root cause analysis weight includes the data weight and the index weight; perform a data deviation operation on the data volume of the root cause location dataset and the preset data volume to obtain a data deviation value (i.e., ), and the data deviation operation is used to quantify the reverse deviation degree of the data volume of the root cause location dataset from the preset data volume; perform an index deviation operation on the alarm warning evaluation index and the warning threshold to obtain an index deviation value (i.e., ), and the index deviation operation is used to quantify the positive deviation degree of the alarm warning evaluation index from the warning threshold; perform a root cause analysis proportion allocation operation based on the data deviation value, the index deviation value, and the corresponding root cause analysis weight to obtain the root cause reliability index. The root cause analysis proportion allocation operation is used to comprehensively quantify the processing method of the reliable degree of the problem source obtained by the topology-driven analysis based on the current root cause location dataset, and the root cause reliability index is used to quantify the reliable degree of the problem source obtained by the topology-driven analysis based on the current root cause location dataset.
[0048] The specific limit expression of the root cause reliability index is as follows; ; In the formula, represents the data volume of the root cause location dataset, represents the preset data volume, represents the alarm warning evaluation index, represents the warning threshold, represents the data weight, represents the index weight, represents the root cause reliability index.
[0049] In this embodiment, the algorithm combines the data volume, the alarm warning evaluation index, the corresponding preset data volume, and the warning threshold for comprehensive analysis to obtain the root cause reliability index. In the formula, when the data volume of the root cause location dataset is lower than the preset data volume, the root cause reliability index is smaller, indicating that the reliability of the problem source obtained based on the current root cause location dataset is lower; when the alarm warning evaluation index is greater than the warning threshold, it indicates that the risk degree of the alarm information in the current root cause location dataset is higher, and the corresponding root cause reliability index is higher, indicating that the reliability of the obtained problem source is higher; at the same time, the root cause reliability index is calculated based on the condition that the data volume of the root cause location dataset is less than the preset data volume. Therefore, the difference between the data volume of the root cause location dataset and the preset data volume (i.e., ) is always not zero; through the analysis of the root cause reliability index, the analysis result of the problem source is made more rigorous, reducing the misreport rate of the problem source, and also further ensuring the real-time risk assessment of the root cause location dataset, improving the accurate judgment of the health status of IT resources and the efficiency of determining the problem source.
[0050] Specifically, the root cause analysis weight is obtained from the preset database, and the root cause analysis weight represents the influence degree of the business requirement data on the root cause reliability index. There is a unique mapping relationship between the data volume, the alarm warning evaluation index and the corresponding root cause analysis weight, and the value range is between 0 and 1; for example, a mapping set of the data volume, the alarm warning evaluation index and the preset root cause analysis weight is constructed, and the real-time data volume and alarm warning evaluation index are input into the mapping set to obtain the data weight and the index weight respectively, which respectively represent the influence degree of the data volume and the alarm warning evaluation index on the root cause reliability index, and the sum of the two is 1.
[0051] Further, the specific process for determining whether to perform topology-driven analysis is as follows: Construct a mapping set between the root cause reliability index and the reliability probability, and input the real-time root cause reliability index into the mapping set to obtain the corresponding reliability probability; Compare the reliability probability with the preset probability obtained from the preset database: If the reliability probability is not less than the preset probability, perform topology-driven analysis on the current root cause location data set to obtain the problem source and feedback the problem source to the preset staff; If the reliability probability is less than the preset probability, continue to obtain alarm information.
[0052] In this embodiment, the preset probability is set by the preset IT personnel according to the specific IT resource data situation; The real-time input of the root cause reliability index and the corresponding reliability probability obtained through the mapping set enable the system to quickly and accurately evaluate the reliability obtained for each problem source; At the same time, based on the comparison between the reliability probability and the preset probability, it helps to decide whether to perform topology-driven analysis on the basis of judging whether the root cause has sufficient reliability. This mechanism effectively avoids unnecessary topology-driven analysis, thereby saving computing resources and processing time.
[0053] The embodiment of the present application provides a method for quantitatively evaluating the health of IT resources from a business perspective, and the specific steps are as follows: Perform quantitative mapping on the obtained IT resource data to obtain business monitoring indicators, obtain business requirement data to obtain the compliance index of the business monitoring indicators, and take corresponding quantitative mapping measures; Generate dynamic baselines to obtain the baseline fluctuation ranges of each business monitoring indicator, obtain alarm information based on each business monitoring indicator and the corresponding baseline fluctuation range, and determine whether to take system emergency measures; Obtain the root cause location data set based on the alarm information and the corresponding IT resource data, and obtain alarm evaluation data based on the root cause location data set to perform topology-driven analysis to obtain the problem source.
[0054] In this embodiment, through the evaluation of the health of IT resources from a business perspective, combined with methods such as dynamic baseline generation, accurate alarm analysis, and root cause location analysis, the accuracy of system monitoring, the efficiency of problem response, and the matching degree between IT resources and business requirements are improved, and the correlation between IT resource data and business requirement analysis is increased.
[0055] In summary, in the embodiments of the present application, the IT resource data obtained is quantitatively mapped to obtain business monitoring indicators, and based on this, the business monitoring indicator compliance index is obtained by combining the obtained business requirement data to take corresponding quantitative mapping measures. Then, the baseline fluctuation range of each business monitoring indicator is generated through dynamic baseline generation. Based on this, alarm information is obtained from each business monitoring indicator, and it is determined whether to take system emergency measures. Finally, a root cause location data set is obtained based on the alarm information and the corresponding IT resource data to obtain alarm evaluation data, and topology-driven analysis is performed to obtain the problem source, thereby increasing the relevance between the determination of IT resource problems and business requirements, and further realizing the improvement of the relevance between IT resource data and business requirement analysis, effectively solving the problem of low relevance between IT resource data and business requirement analysis in the prior art.
[0056] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0057] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows Figure 1 or blocks.
[0058] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more flows Figure 1 or blocks.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or multiple processes and / or one block or multiple blocks. Figure 1 one process or multiple processes and / or blocks Figure 1 steps for realizing the functions specified in one block or multiple blocks.
[0060] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0061] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An IT resource health quantification and evaluation system from a business perspective, characterized in that, Including: A business monitoring metric quantification module, a dynamic baseline adaptation module, and a root cause localization module: Among them, the business monitoring metric quantification module is used to obtain business monitoring metrics through quantitative mapping of the obtained IT resource data, obtain a business monitoring metric compliance index from business requirement data to take corresponding quantitative mapping measures; The dynamic baseline adaptation module is used to obtain the baseline fluctuation range of each business monitoring metric through dynamic baseline generation, obtain an alarm message based on each business monitoring metric and the corresponding baseline fluctuation range, and determine whether to take system emergency measures; The root cause localization module is used to obtain a root cause localization data set based on the alarm message and the corresponding IT resource data, and obtain alarm evaluation data based on the root cause localization data set to perform topology-driven analysis to obtain the problem source.
2. The IT resource health quantification and evaluation system based on the business perspective according to claim 1, characterized in that: The business requirement data includes the business IT ratio, the business KPI correlation degree, and the number of business emergency measures; The business requirement ratio represents the ratio of business growth to IT resource consumption, the business KPI correlation degree represents the correlation between business monitoring metrics and business KPIs, and the number of business emergency measures represents the number of abnormal events corresponding to business monitoring metrics with emergency measures; The alarm evaluation data includes the average alarm level, the average alarm frequency, the average alarm duration, the average baseline difference value, and the number of alarm trigger types; The average baseline difference value represents the average value of the deviation difference between the business monitoring metric corresponding to the alarm message and the baseline fluctuation range.
3. The IT resource health quantification evaluation system based on the business perspective according to claim 2, wherein: The process of obtaining the business monitoring metric compliance index from the business requirement data is as follows: Obtain the business compliance weight and the reference business busyness from a preset database, and the business compliance weight includes the business busyness weight, the business availability weight, and the business health weight; If the business busyness is not less than the corresponding reference business busyness, perform a busyness deviation operation on the business busyness and the corresponding reference business busyness to obtain a busyness deviation, and perform a busyness balance ratio operation on the busyness deviation and the business IT ratio to obtain a first business compliance analysis index. The busyness deviation operation is used to quantify the deviation degree of the business busyness from the reference business busyness, and the busyness balance ratio operation is used to quantify the compliance degree of the business monitoring metric from the perspective of business busyness; Perform an availability balance ratio operation on the business KPI correlation degree and the business availability to obtain a second business compliance analysis index. The availability balance ratio operation is used to quantify the compliance degree of the business monitoring metric from the perspective of business availability; Perform a normalization operation on the number of business emergency measures to obtain a business emergency degree, and perform an availability balance ratio operation on the business emergency degree and the business health degree to obtain a third business compliance analysis index. The normalization operation is used to quantify the response degree of the number of business emergency measures to abnormal events, and the availability balance ratio operation is used to quantify the compliance degree of the business monitoring metric from the perspective of business health. Perform the proportional distribution operation of business compliance indicators by combining the first business compliance analysis indicator, the second business compliance analysis indicator, the third business compliance analysis indicator, and the corresponding business compliance weights to obtain the business monitoring indicator compliance index. The proportional distribution operation of business compliance indicators is used to proportionally distribute the influence degree of the compliance degree of each business monitoring indicator and IT resource data. If the business busyness is less than the corresponding reference business busyness, perform the secondary business compliance indicator proportional distribution operation based on the second business compliance analysis indicator, the third business compliance analysis indicator, and the corresponding business compliance weights to obtain the business monitoring indicator compliance index. The secondary business compliance indicator proportional distribution operation is used to proportionally distribute the influence degree of the compliance degree of the second business compliance analysis indicator and the third business compliance analysis indicator and the corresponding IT resource data.
4. The IT resource health quantification and evaluation system based on the business perspective according to claim 3, characterized in that: The corresponding quantization mapping measures are taken as follows: If the business monitoring indicator compliance index is not less than the compliance threshold, continue to execute the dynamic baseline adaptive module. If the business monitoring indicator compliance index is less than the compliance threshold, re-obtain the IT resource data for quantization mapping and take the corresponding quantization mapping measures as follows: Perform difference processing on the business monitoring indicator compliance index and the compliance threshold to obtain the compliance deviation. The difference processing represents a way to quantify the gap between the business monitoring indicator compliance index and the compliance threshold. Construct a frequency mapping set between the compliance deviation and the frequency adjustment multiple. Input the real-time compliance deviation into the frequency mapping set to obtain the corresponding frequency adjustment multiple speed, and perform a frequency multiple scaling operation on the acquisition frequency of the IT resource data and the frequency adjustment multiple speed to obtain the adjustment frequency. When the adjustment frequency is lower than the update frequency limit value, use the adjustment frequency to obtain the IT resource data, otherwise use the update frequency limit value. Construct a weight mapping set between the compliance deviation and the data ratio adjustment value. Input the real-time compliance deviation into the weight mapping set to obtain the corresponding data ratio adjustment value, and perform a ratio adjustment process on the acquisition weight of the IT resource data and the data ratio adjustment value to obtain the adjustment ratio, and obtain the corresponding IT resource data based on the adjustment ratio. When the number of re-quantization mappings is higher than the adjustment times obtained from the preset database, then take data filling measures.
5. The IT resource health quantification and evaluation system based on the business perspective according to claim 1, wherein: The specific steps for obtaining the alarm information according to each business monitoring indicator and the corresponding baseline fluctuation range are as follows: Generate the baseline fluctuation range corresponding to each business monitoring indicator through dynamic baseline generation, and compare each business monitoring indicator with the baseline fluctuation range: If the business monitoring indicator is within the corresponding baseline fluctuation range, no alarm information is generated. If the business monitoring indicator is not within the corresponding baseline fluctuation range, generate the corresponding alarm information and judge the alarm level of the alarm information. If the alarm level is higher than the preset alarm level, perform emergency processing according to the alarm level corresponding to the alarm information, otherwise perform problem reliability analysis. The emergency processing is used to automatically provide the corresponding system emergency measures for the corresponding alarm level.
6. The IT resource health quantification and evaluation system based on the business perspective according to claim 5, wherein: The specific process of performing problem reliability analysis is as follows: Obtain a root cause location data set based on the alarm information corresponding to all service monitoring metrics and the corresponding IT resource data; If the data volume of the root cause location data set is higher than the preset data volume, automatically perform topology-driven analysis on the root cause location data set to obtain the problem root cause, otherwise obtain alarm evaluation data; Obtain the corresponding alarm warning evaluation index based on the alarm evaluation data, and compare the alarm warning evaluation index with the warning threshold obtained from the preset database: If the alarm warning evaluation index is less than the warning threshold, continue to obtain alarm information; If the alarm warning evaluation index is not less than the warning threshold, obtain the root cause reliability index to determine whether to perform topology-driven analysis.
7. The IT resource health quantification and evaluation system based on the business perspective as claimed in claim 6, wherein: The specific process of obtaining the alarm warning evaluation index is as follows: Obtain the alarm evaluation weights from the preset database, and the alarm evaluation weights include alarm level weight, alarm frequency weight, alarm time weight, baseline difference weight, and alarm type weight; Perform data preprocessing on the alarm evaluation data, and the data preprocessing is used to de-unitize the alarm evaluation data; Obtain the baseline fluctuation range, and the baseline fluctuation range represents the range corresponding to the maximum value and the minimum value of the baseline fluctuation range; Perform range deviation operation on each service monitoring metric and the corresponding baseline fluctuation range to obtain the average baseline proximity difference, and the range deviation operation is used to obtain data for measuring the deviation degree of the service monitoring metric from the baseline fluctuation range; Perform trend fitting quantization operation on the alarm evaluation data to obtain the corresponding alarm quantization value, and the trend fitting quantization operation is used to unify the trend of each alarm evaluation data and obtain the corresponding quantization data; Perform proportion allocation and balancing operation according to the alarm quantization value and the corresponding alarm evaluation weights to obtain the alarm warning evaluation index. The alarm warning evaluation index represents the data for quantifying the risk degree of the alarm information in the root cause location data set, and the proportion allocation and balancing operation is used to quantify the risk degree of the alarm information in the root cause location data set reflected by the alarm warning evaluation data.
8. The IT resource health quantification and evaluation system based on the business perspective according to claim 6, wherein: The specific process of the root cause reliability index is as follows: Obtain the root cause analysis weights from the preset database, and the root cause analysis weights include data weight and index weight; Perform data deviation operation on the data volume of the root cause location data set and the preset data volume to obtain the data deviation value, and the data deviation operation is used to quantify the reverse deviation degree of the data volume of the root cause location data set from the preset data volume; Perform index deviation operation on the alarm warning evaluation index and the warning threshold to obtain the index deviation value, and the index deviation operation is used to quantify the positive deviation degree of the alarm warning evaluation index from the warning threshold; Perform root cause analysis proportion allocation operation based on the data deviation value, the index deviation value, and the corresponding root cause analysis weights to obtain the root cause reliability index. The root cause analysis proportion allocation operation is used to comprehensively quantify the processing method for the reliability degree of the problem source obtained by performing topology-driven analysis based on the current root cause location data set, and the root cause reliability index is used to quantify the reliability degree of the problem source obtained by performing topology-driven analysis based on the current root cause location data set.
9. The IT resource health quantification and evaluation system based on the business perspective according to claim 8, wherein: The specific process of determining whether to perform topology-driven analysis is as follows: Construct a mapping set between the root cause reliability index and the reliability probability, and input the real-time root cause reliability index into the mapping set to obtain the corresponding reliability probability; Compare the reliability probability with the preset probability obtained from the preset database: If the reliability probability is not less than the preset probability, perform topology-driven analysis on the current root cause location data set to obtain the problem source and feedback the problem source to the preset staff; If the reliability probability is less than the preset probability, continue to obtain the alarm information.
10. A method for quantitatively evaluating the health of IT resources from a business perspective, characterized in that: The specific steps are as follows: Obtain business monitoring indicators through quantitative mapping of the obtained IT resource data, and obtain the business monitoring indicator compliance index from the business requirement data to take corresponding quantitative mapping measures; Generate the baseline fluctuation range of each business monitoring indicator through dynamic baseline generation, obtain the alarm information based on each business monitoring indicator and the corresponding baseline fluctuation range, and judge whether to take system emergency measures; Obtain the root cause location data set based on the alarm information and the corresponding IT resource data, and obtain the alarm evaluation data based on the root cause location data set to perform topology-driven analysis to obtain the problem source.
Citation Information
Patent Citations
A Health Assessment Method for Business Systems with Centralized IT Monitoring
CN111274087B
Business application system resource health degree prediction and evaluation method and system based on Informer
CN118034919A
Method and device for evaluating high availability of application system
CN117851192A
Intelligent alarm analysis method and system based on AIOps
CN118245923A
Operation monitoring method and device of energy internet marketing service system
CN119089349A