Software system monitoring index threshold determination method and electronic equipment
By using machine learning models to analyze monitoring indicator time series data and metadata in the software monitoring system, the threshold is automatically determined, and the subjective impact and high cost problems caused by manual setting of thresholds are solved, and the accuracy and efficiency of the monitoring system are improved.
Patent Information
- Application Number
- CN202510328716.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, software monitoring systems rely on manual setting monitoring indicator thresholds, which have problems such as subjective influence, high labor costs and poor prediction results.
By using the time window to obtain the time series data and related metadata of the monitoring indicators, the machine learning model is used to analyze these data, and the monitoring indicator threshold is automatically determined.
It improves the accuracy of the threshold recommendation of monitoring indicators, reduces the dependence on manual experience, and improves the efficiency of the monitoring system.
Smart Images

Figure CN119961104A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of software monitoring technology, and in particular to a method for determining a threshold value of a software system monitoring indicator, a computer-readable storage medium, a computer program product, and an electronic device. Background Art
[0002] The software monitoring system refers to a system used to monitor and evaluate the running status of computer software, aiming to detect indicators of software performance, stability and security in real time and provide effective early warning. Among them, the monitoring system determines whether the software system is running normally through the indicator threshold. Once the indicator exceeds the set threshold range, the alarm mechanism will be triggered to notify relevant personnel to handle it.
[0003] In the use of monitoring systems, the traditional method of setting thresholds for monitoring indicators relies on the subjective settings of operation and maintenance personnel. This method of manually setting thresholds for monitoring indicators has some disadvantages. First, manually setting thresholds requires relying on personal subjective judgment. Different users may have different experiences, preferences, and understandings, resulting in inconsistency in threshold settings. This subjectivity may lead to overly loose or overly strict thresholds, which in turn affects the accuracy and reliability of the monitoring system in detecting abnormal situations. In addition, for complex systems or large-scale monitoring indicators, manually setting thresholds is more difficult because it requires an in-depth understanding of the relationship and impact between various indicators, which may be a tedious and time-consuming task for users. In addition, manually setting thresholds for monitoring indicators is prone to human errors and omissions. Users may set unreasonable thresholds due to negligence, lack of experience, or wrong estimates, resulting in the monitoring system being unable to accurately detect abnormal situations and delaying the discovery and resolution of problems. Especially in large-scale enterprise-level monitoring systems, the trial-and-error cost of threshold setting is very high. Manually setting thresholds may require processing a large number of indicators, which is prone to configuration errors or omissions of important indicators, further exacerbating the risk and stability issues of the monitoring system.
[0004] In summary, manually setting monitoring indicator thresholds by operation and maintenance personnel has many shortcomings, such as high subjective influence factors, prone to human errors and omissions, etc. These problems limit the accuracy and efficiency of most current monitoring systems. Summary of the invention
[0005] The main purpose of the present application is to provide a method for determining a threshold value of a software system monitoring indicator, a computer-readable storage medium, a computer program product, and an electronic device, so as to at least solve the problems of subjective influence, high labor cost, and poor prediction effect caused by manually setting thresholds in software monitoring in the prior art.
[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for determining a threshold value of a monitoring indicator of a software system is provided, comprising: using a time window to obtain monitoring indicator time series data, and obtaining metadata related to the monitoring indicator, the monitoring indicator time series data representing the value of the monitoring indicator in the software system that changes over time within a certain time range; using a machine learning model to analyze the monitoring indicator time series data and the metadata to obtain a monitoring indicator threshold, wherein the machine learning model is trained by machine learning using multiple groups of training data, each group of the training data includes sample monitoring indicator time series data, sample metadata, and sample monitoring indicator threshold; according to the monitoring indicator threshold, determining whether an abnormality occurs in the software system.
[0007] Optionally, a machine learning model is used to analyze the monitoring indicator time series data and the metadata, including: performing data preprocessing on the monitoring indicator time series data and the metadata, the data preprocessing including data cleaning and data normalization; performing feature engineering on the monitoring indicator time series data and the metadata after data preprocessing; and using the machine learning model to analyze the monitoring indicator time series data and the metadata after feature engineering.
[0008] Optionally, feature engineering processing is performed on the preprocessed monitoring indicator time series data and the metadata, including: using rolling statistical calculations to extract temporal statistical features of the monitoring indicator time series data; and using label encoding to convert text data in the metadata into numerical features.
[0009] Optionally, before using the machine learning model to analyze the monitoring indicator time series data and the metadata, the method also includes: constructing the machine learning model, the machine learning model includes an adaptive learning module, an optimization module and a fusion decision module, the adaptive learning module includes an XGBoost sub-model, a LightGBM sub-model and a CatBoost sub-model, the optimization module is used to search and optimize the hyperparameters of each sub-model in the adaptive learning module, and the fusion decision module is used to fuse the sub-models in the adaptive learning module.
[0010] Optionally, constructing the machine learning model includes: constructing the adaptive learning module, the loss function corresponding to the adaptive learning module includes the root mean square error; constructing the optimization module, the optimization module is used to optimize the adaptive learning module; constructing the fusion decision module, the fusion decision module is used to fuse the sub-models in the optimized adaptive learning module to obtain the machine learning model.
[0011] Optionally, constructing the optimization module includes: constructing the optimization module, wherein the optimization module is used to use a particle swarm algorithm to perform heuristic search optimization on the hyperparameters of each sub-model in the adaptive learning module to obtain the optimal hyperparameter combination corresponding to the adaptive learning module.
[0012] Optionally, constructing the fusion decision module includes: constructing the fusion decision module, wherein the fusion decision module is used to fuse the sub-models in the optimized adaptive learning module by using a Stacking algorithm.
[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the methods for determining the threshold value of a software system monitoring indicator.
[0014] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, any one of the methods for determining a threshold value of a software system monitoring indicator is implemented.
[0015] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for determining a threshold value of a software system monitoring indicator for executing any one of the above-described methods.
[0016] Applying the technical solution of this application, firstly, the monitoring indicator time series data is obtained by using the time window, and the metadata related to the monitoring indicator is obtained, and then the monitoring indicator time series data and metadata are analyzed by using the machine learning model to obtain the monitoring indicator threshold, and finally, according to the monitoring indicator threshold, it is determined whether the software system is abnormal. Compared with the subjective influence, high labor cost and poor prediction effect caused by relying on manual setting of thresholds in software monitoring in the prior art, this application comprehensively introduces the monitoring indicator time series data and the metadata of the monitoring indicator at the data level, expands the data dimension, and improves the modeling ability of the monitoring indicator threshold recommendation process from the feature level compared with the traditional manual expert experience method and the machine learning method that only considers the historical data of the indicator, thereby improving the accuracy of the monitoring indicator threshold recommendation, and by using the machine learning model to automatically analyze the monitoring indicator time series data and related metadata, the threshold of the monitoring indicator can be automatically determined, which reduces the dependence on manual experience and improves the efficiency of the monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings constituting part of the present application are used to provide a further understanding of the present application. The exemplary embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 A hardware structure block diagram of a mobile terminal for executing a method for determining a threshold value of a software system monitoring indicator provided in an embodiment of the present application is shown;
[0019] Figure 2 A flowchart of a method for determining a threshold value of a software system monitoring indicator provided in accordance with an embodiment of the present application is shown;
[0020] Figure 3 A flowchart of a specific method for determining a threshold value of a software system monitoring indicator provided in accordance with an embodiment of the present application is shown.
[0021] The above drawings include the following reference numerals:
[0022] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION
[0023] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] As introduced in the background technology, the software monitoring in the prior art relies on manually setting thresholds, which brings about subjective influences, high labor costs and poor prediction effects. To solve the above problems, the embodiments of the present application provide a method for determining a software system monitoring indicator threshold, a computer-readable storage medium, a computer program product and an electronic device.
[0027] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0028] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal of a method for determining a threshold value of a software system monitoring indicator according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the threshold value of the software system monitoring index in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific examples of the network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0030] In this embodiment, a method for determining the threshold value of a software system monitoring indicator running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0031] Figure 2 1 is a flowchart of a method for determining a threshold value of a software system monitoring indicator according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:
[0032] Step S201, using a time window to obtain monitoring indicator time series data, and obtain metadata related to the monitoring indicator, the above monitoring indicator time series data represents the value of the above monitoring indicator in the software system that changes over time within a certain time range;
[0033] Step S202: using a machine learning model to analyze the monitoring indicator time series data and the metadata to obtain a monitoring indicator threshold, wherein the machine learning model is trained by machine learning using multiple sets of training data, and each set of the training data includes sample monitoring indicator time series data, sample metadata, and sample monitoring indicator threshold;
[0034] Step S203: Determine whether an abnormality occurs in the software system according to the monitoring indicator threshold.
[0035] Through the above embodiments, firstly, the monitoring indicator time series data is obtained by using the time window, and the metadata related to the monitoring indicator is obtained, and then the monitoring indicator time series data and metadata are analyzed by using the machine learning model to obtain the monitoring indicator threshold, and finally, according to the monitoring indicator threshold, it is determined whether the software system is abnormal. Compared with the subjective influence, high labor cost and poor prediction effect caused by relying on manual setting of thresholds in the software monitoring of the prior art, this application comprehensively introduces the monitoring indicator time series data and the metadata of the monitoring indicator at the data level, expands the data dimension, and improves the modeling ability of the monitoring indicator threshold recommendation process from the feature level compared with the traditional manual expert experience method and the machine learning method that only considers the historical data of the indicator, thereby improving the accuracy of the monitoring indicator threshold recommendation, and by using the machine learning model to automatically analyze the monitoring indicator time series data and related metadata, the threshold of the monitoring indicator can be automatically determined, which reduces the dependence on manual experience and improves the efficiency of the monitoring system.
[0036] It should be noted that the software monitoring system is a system used to monitor and manage the running status of the software. It evaluates the performance and stability of the software by collecting and analyzing various indicator data, such as CPU utilization, memory usage, network traffic, etc. The monitoring system can provide real-time data display, anomaly detection and alarm functions to help administrators discover and solve problems in a timely manner.
[0037] Specifically, monitoring indicator time series data refers to a series of values recorded over time as monitoring indicators of software systems or hardware devices change over time within a certain time range. These data are usually collected automatically by the monitoring system and reflect the operating status and performance of the system at different time points, such as CPU utilization, memory usage, network traffic, response time, error rate and other key performance indicators (KPIs). The characteristic of time series data is that it is arranged in chronological order and can be used to analyze the operating trend, periodicity, seasonal changes and anomaly detection of the system.
[0038] Specifically, metadata mainly includes: the category of the indicator monitoring object (for example, transaction monitoring indicators, operation monitoring indicators, physical equipment monitoring indicators), the indicator's alarm level, the indicator's threshold type (normal threshold, year-on-year threshold, quarter-on-quarter threshold), indicator collection frequency, etc.
[0039] In an optional solution, a machine learning model is used to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, including: performing data preprocessing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned data preprocessing includes data cleaning and data normalization; performing feature engineering processing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata after data preprocessing; and using the above-mentioned machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata after feature engineering processing. In this embodiment, by performing data preprocessing on the monitoring indicator time series data and metadata, including data cleaning and data normalization, outliers, missing values and noise can be removed, and the accuracy and consistency of the data can be improved, thereby providing higher quality data input for subsequent machine learning model analysis; using the data processed by feature engineering for machine learning model analysis can more accurately capture the potential laws and patterns in the data, thereby further improving the accuracy of the model's prediction of the monitoring indicator threshold.
[0040] According to some exemplary embodiments of the present application, feature engineering is performed on the preprocessed monitoring indicator time series data and the metadata, including: extracting the time series statistical features of the monitoring indicator time series data using rolling statistical calculations; and converting the text data in the metadata into numerical features using label encoding. In this embodiment, by extracting the time series statistical features of the time series data, the time series dependency and trend information in the data can be captured. These features can provide richer information to the machine learning model, thereby further improving the accuracy of the model's prediction of the monitoring indicator thresholds. The text data in the metadata is converted into numerical features, so that the machine learning model can process non-numerical data, enhancing the model's generalization ability for different types of data.
[0041] In other embodiments, before using the machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned method also includes: constructing the above-mentioned machine learning model, the above-mentioned machine learning model includes an adaptive learning module, an optimization module and a fusion decision module, the above-mentioned adaptive learning module includes an XGBoost sub-model, a LightGBM sub-model and a CatBoost sub-model, the above-mentioned optimization module is used to search and optimize the hyperparameters of each sub-model in the above-mentioned adaptive learning module, and the above-mentioned fusion decision module is used to fuse each of the above-mentioned sub-models in the above-mentioned adaptive learning module. In this embodiment, the machine learning model integrates multiple models through three constituent modules: an adaptive learning module, an optimization module and a fusion decision module, limits the tendency of a single model in representation learning, reduces the risk of overfitting, and combines the respective advantages of different models to improve the prediction performance, generalization and robustness of the overall model.
[0042] Specifically, the monitoring indicator time series data and metadata are input into the XGBoost sub-model, the monitoring indicator time series data and metadata are input into the LightGBM sub-model, and the monitoring indicator time series data and metadata are input into the CatBoost sub-model to obtain the preliminary regression prediction results of each sub-model.
[0043] Specifically, the adaptive learning module uses XGBoost, LightGBM and CatBoost models. These three models have shown excellent performance on different types of features and data sets, and can automatically adjust to adapt to the complexity and diversity of data. Together, they form a powerful learning framework that can learn and extract complex patterns and rules from monitoring indicator time series data and metadata; the optimization module searches and optimizes the hyperparameters of each sub-model in the adaptive learning module. Hyperparameters are key parameters that affect model behavior and generalization capabilities. This step is crucial to improving the predictive performance of the model; the fusion decision module fuses the optimized XGBoost, LightGBM and CatBoost models. Model fusion can reduce the bias and variance of a single model and improve the consistency and reliability of the overall prediction. Through the comprehensive prediction of multiple sub-models, the risk of overfitting can be reduced, the generalization ability of the model can be improved, and the performance of the recommended monitoring indicator threshold on unknown data can also be ensured to be consistent.
[0044] According to some other exemplary embodiments of the present application, constructing the above-mentioned machine learning model includes: constructing the above-mentioned adaptive learning module, the loss function corresponding to the above-mentioned adaptive learning module includes the root mean square error; constructing the above-mentioned optimization module, the above-mentioned optimization module is used to optimize the above-mentioned adaptive learning module; constructing the above-mentioned fusion decision module, the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the above-mentioned adaptive learning module after optimization to obtain the above-mentioned machine learning model. In this embodiment, by using the root mean square error as the loss function of the adaptive learning module, the regression performance of the model is further optimized directly targeting the task characteristics of the indicator threshold prediction.
[0045] In some other optional schemes of the present application, constructing the above-mentioned optimization module includes: constructing the above-mentioned optimization module, and the above-mentioned optimization module is used to use the particle swarm algorithm to perform heuristic search optimization on the above-mentioned hyperparameters of each of the above-mentioned sub-models in the above-mentioned adaptive learning module to obtain the optimal hyperparameter combination corresponding to the above-mentioned adaptive learning module. In this embodiment, the construction of the optimization module realizes efficient search of hyperparameters through the particle swarm algorithm, avoids blind trial and error, avoids the subjectivity and time-consuming of manually adjusting hyperparameters, and further improves the efficiency of model training and prediction accuracy.
[0046] In some other optional schemes of the present application, constructing the above-mentioned fusion decision module includes: constructing the above-mentioned fusion decision module, and the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the optimized above-mentioned adaptive learning module using the Stacking algorithm. In this embodiment, the construction of the fusion decision module integrates the prediction results of multiple sub-models through the Stacking algorithm, further improving the stability and prediction accuracy of the model.
[0047] Specifically, the adaptive learning module is used to characterize and learn the characteristics of monitoring indicators and to make preliminary regression predictions on the thresholds of monitoring indicators. The optimization module introduces a particle swarm algorithm to perform heuristic search and optimization on the hyperparameters of the adaptive learning module sub-models to find the optimal hyperparameter combination of the sub-models. The fusion decision module introduces a Stacking algorithm to fuse the optimized adaptive learning module sub-models, and by combining the prediction results of multiple sub-models, improves the accuracy, stability and generalization ability of the overall model, and makes final recommendations on the thresholds of monitoring indicators.
[0048] Specifically, Figure 3 As shown, the specific steps of the method for determining the threshold value of the software system monitoring indicator of the present application include: Step 1: Acquire the collected historical time series data of the monitoring indicator (i.e., the sample monitoring indicator time series data), the metadata related to the monitoring indicator (i.e., the sample metadata), and the configured historical monitoring indicator threshold value (i.e., the sample monitoring indicator threshold value) from the monitoring system database; Step 2: Perform feature preprocessing on the collected data, mainly including data cleaning and normalization, and then perform feature engineering on the preprocessed data, wherein the feature engineering is mainly divided into two parts: for the historical time series data of the monitoring indicator, the rolling statistical calculation is mainly used to extract the data; Time series statistical features (such as standard deviation, mean deviation, variance, quartile, weekday and holiday features, etc.), where the rolling window length T = 3 days, and for indicator metadata, label encoding is used to convert text data into numerical features that can be input into the model; Step 3: Build a data set, and use the feature engineering data to form a training set and a test set in a ratio of 8:2; Step 4: Build an adaptive learning module. The sub-model of the adaptive learning module consists of the XGBoost model, LightGBM model, and CatBoost model, which are widely used in many fields and have leading performance. The training loss function selects the root mean square error (RMSE); Step 5: Optimize the adaptive learning module, introduce the particle swarm algorithm to search for the optimal combination of sub-model hyperparameters in the adaptive learning module, where a hyperparameter combination is used as a particle in the particle swarm algorithm, and the loss function of the adaptive learning module is used as the objective function of the optimization algorithm. The mathematical representation of the optimization process is: Where x represents the training data, f xgboost 、flightgbm 、f catboost Represents the three sub-models of the adaptive learning module; Step 6: Model training and fusion, introduce the Stacking algorithm to use cross-validation to divide the training data set into multiple subsets, train the XGBoost sub-model, LightGBM sub-model and CatBoost sub-model on each subset, and use their prediction results to build new features, and then integrate the prediction results of each sub-model through a multi-layer perceptron (MLP) model. This process is mathematically expressed as where F stacking Represents the stacking algorithm, Represents the three sub-models of the adaptive learning module after optimization by the particle swarm algorithm; Step 7: After steps 1 to 6, the training of the fusion optimization model (i.e., the machine learning model) is completed. When the user configures the monitoring indicator threshold, the current time is taken as the starting point, and the corresponding monitoring indicator time series data with a window length of T = 3 days is retrieved forward, and the time series statistical features are calculated in real time, and input into the model together with the metadata features of this indicator. The model outputs the recommended threshold of the monitoring indicator (i.e., the monitoring indicator threshold) according to the calculation. The process is mathematically expressed as follows: Where y is the recommended threshold of the final output (i.e., the monitoring indicator threshold).
[0049] In summary, this application first starts from the data level, and introduces more comprehensive monitoring indicator-related data by combining the collection of monitoring indicator time series data and monitoring indicator metadata, and performing effective feature engineering, expanding the data dimension to improve the modeling ability of the indicator threshold recommendation process, and improving the accuracy of model prediction from the feature level; this application also proposes a new machine learning model, which combines multiple models by integrating a decision-making module and an adaptive learning module, constrains the tendency of a single model in representation learning, reduces the risk of overfitting, and improves the generalization ability and robustness of the model. It also fully combines the advantages of different models, reduces the limitations of single model prediction, and improves the overall prediction performance. At the same time, a particle swarm algorithm is introduced to the model The optimal hyperparameter combination of the model is heuristically searched. Compared with the method of adjusting hyperparameters through manual experiments, it improves efficiency and can also calculate the complex nonlinear relationship between model hyperparameters that is difficult to fully consider manually, further improving the performance of the model. At the same time, this machine learning algorithm is introduced into the field of indicator threshold recommendation of monitoring system for effective application. The application can quickly and automatically optimize and generate reasonable monitoring indicator thresholds, help operation and maintenance personnel to configure monitoring indicator thresholds, improve the accuracy, sensitivity and reliability of the monitoring system, thereby improving the performance of the monitoring system and user experience. It has broad application prospects in large-scale and complex monitoring systems, and can also provide effective solutions for predictive analysis tasks with similar data scenarios.
[0050] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0051] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the method for determining the threshold value of the software system monitoring indicator.
[0052] Specifically, the method for determining the threshold of the software system monitoring indicator includes:
[0053] Step S201, using a time window, obtaining monitoring indicator time series data and metadata related to the monitoring indicator, wherein the monitoring indicator time series data represents the value of the monitoring indicator in the software system that changes over time within a certain time range;
[0054] Step S202: using a machine learning model to analyze the monitoring indicator time series data and the metadata to obtain a monitoring indicator threshold, wherein the machine learning model is trained by machine learning using multiple sets of training data, and each set of the training data includes sample monitoring indicator time series data, sample metadata, and sample monitoring indicator threshold;
[0055] Step S203: Determine whether an abnormality occurs in the software system according to the monitoring indicator threshold.
[0056] Optionally, a machine learning model is used to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, including: performing data preprocessing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned data preprocessing includes data cleaning and data normalization; performing feature engineering processing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata after data preprocessing; and using the above-mentioned machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata after feature engineering processing.
[0057] Optionally, feature engineering processing is performed on the preprocessed monitoring indicator time series data and the metadata, including: using rolling statistical calculations to extract time series statistical features of the monitoring indicator time series data; and using label encoding to convert text data in the metadata into numerical features.
[0058] Optionally, before using the machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned method also includes: constructing the above-mentioned machine learning model, the above-mentioned machine learning model includes an adaptive learning module, an optimization module and a fusion decision module, the above-mentioned adaptive learning module includes an XGBoost sub-model, a LightGBM sub-model and a CatBoost sub-model, the above-mentioned optimization module is used to search and optimize the hyperparameters of each sub-model in the above-mentioned adaptive learning module, and the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the above-mentioned adaptive learning module.
[0059] Optionally, constructing the above-mentioned machine learning model includes: constructing the above-mentioned adaptive learning module, the loss function corresponding to the above-mentioned adaptive learning module includes the root mean square error; constructing the above-mentioned optimization module, the above-mentioned optimization module is used to optimize the above-mentioned adaptive learning module; constructing the above-mentioned fusion decision module, the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the optimized above-mentioned adaptive learning module to obtain the above-mentioned machine learning model.
[0060] Optionally, constructing the above-mentioned optimization module includes: constructing the above-mentioned optimization module, and the above-mentioned optimization module is used to use a particle swarm algorithm to perform heuristic search optimization on the above-mentioned hyperparameters of each of the above-mentioned sub-models in the above-mentioned adaptive learning module to obtain the optimal hyperparameter combination corresponding to the above-mentioned adaptive learning module.
[0061] Optionally, constructing the above-mentioned fusion decision module includes: constructing the above-mentioned fusion decision module, and the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the optimized above-mentioned adaptive learning module by using a Stacking algorithm.
[0062] The present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement at least the following method steps: step S201, using a time window to obtain monitoring indicator time series data and metadata related to the monitoring indicator, wherein the monitoring indicator time series data represents the value of the monitoring indicator in the software system that changes over time within a certain time range; step S202, using a machine learning model to analyze the monitoring indicator time series data and the metadata to obtain a monitoring indicator threshold, wherein the machine learning model is trained by machine learning using multiple groups of training data, and each group of the training data includes sample monitoring indicator time series data, sample metadata and sample monitoring indicator threshold; step S203, determining whether an abnormality occurs in the software system based on the monitoring indicator threshold.
[0063] Optionally, a machine learning model is used to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, including: performing data preprocessing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned data preprocessing includes data cleaning and data normalization; performing feature engineering processing on the above-mentioned monitoring indicator time series data and the above-mentioned metadata after data preprocessing; and using the above-mentioned machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata after feature engineering processing.
[0064] Optionally, feature engineering processing is performed on the preprocessed monitoring indicator time series data and the metadata, including: using rolling statistical calculations to extract time series statistical features of the monitoring indicator time series data; and using label encoding to convert text data in the metadata into numerical features.
[0065] Optionally, before using the machine learning model to analyze the above-mentioned monitoring indicator time series data and the above-mentioned metadata, the above-mentioned method also includes: constructing the above-mentioned machine learning model, the above-mentioned machine learning model includes an adaptive learning module, an optimization module and a fusion decision module, the above-mentioned adaptive learning module includes an XGBoost sub-model, a LightGBM sub-model and a CatBoost sub-model, the above-mentioned optimization module is used to search and optimize the hyperparameters of each sub-model in the above-mentioned adaptive learning module, and the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the above-mentioned adaptive learning module.
[0066] Optionally, constructing the above-mentioned machine learning model includes: constructing the above-mentioned adaptive learning module, the loss function corresponding to the above-mentioned adaptive learning module includes the root mean square error; constructing the above-mentioned optimization module, the above-mentioned optimization module is used to optimize the above-mentioned adaptive learning module; constructing the above-mentioned fusion decision module, the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the optimized above-mentioned adaptive learning module to obtain the above-mentioned machine learning model.
[0067] Optionally, constructing the above-mentioned optimization module includes: constructing the above-mentioned optimization module, and the above-mentioned optimization module is used to use a particle swarm algorithm to perform heuristic search optimization on the above-mentioned hyperparameters of each of the above-mentioned sub-models in the above-mentioned adaptive learning module to obtain the optimal hyperparameter combination corresponding to the above-mentioned adaptive learning module.
[0068] Optionally, constructing the above-mentioned fusion decision module includes: constructing the above-mentioned fusion decision module, and the above-mentioned fusion decision module is used to fuse the above-mentioned sub-models in the optimized above-mentioned adaptive learning module by using a Stacking algorithm.
[0069] An embodiment of the present application also provides an electronic device, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for determining a threshold value of a software system monitoring indicator for executing any one of the above-mentioned software system monitoring indicators.
[0070] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0071] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0076] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0077] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0078] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0079] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0080] In the method for determining the threshold of the monitoring indicator of the software system of the present application, firstly, the monitoring indicator time series data is obtained by using the time window, and the metadata related to the monitoring indicator is obtained, and then the monitoring indicator time series data and metadata are analyzed by using the machine learning model to obtain the monitoring indicator threshold, and finally, according to the monitoring indicator threshold, it is determined whether the software system is abnormal. Compared with the subjective influence, high labor cost and poor prediction effect caused by relying on manual setting of thresholds in the software monitoring of the prior art, the present application comprehensively introduces the monitoring indicator time series data and the metadata of the monitoring indicator at the data level, expands the data dimension, and improves the modeling ability of the monitoring indicator threshold recommendation process from the feature level compared with the traditional manual expert experience method and the machine learning method that only considers the historical data of the indicator, thereby improving the accuracy of the monitoring indicator threshold recommendation, and by using the machine learning model to automatically analyze the monitoring indicator time series data and related metadata, the threshold of the monitoring indicator can be automatically determined, reducing the dependence on manual experience and improving the efficiency of the monitoring system.
[0081] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining a threshold value of a software system monitoring indicator, characterized in that: include: Acquire monitoring indicator time series data using a time window and obtain metadata related to the monitoring indicator, wherein the monitoring indicator time series data represents the value of the monitoring indicator in the software system that changes over time within a certain time range; The monitoring indicator time series data and the metadata are analyzed using a machine learning model to obtain a monitoring indicator threshold, wherein the machine learning model is trained by machine learning using multiple sets of training data, and each set of training data includes sample monitoring indicator time series data, sample metadata, and sample monitoring indicator threshold; Determine whether an abnormality occurs in the software system based on the monitoring indicator threshold.
2. The method for determining the threshold value of software system monitoring indicators according to claim 1, characterized in that: The monitoring indicator time series data and the metadata are analyzed using a machine learning model, including: Performing data preprocessing on the monitoring indicator time series data and the metadata, wherein the data preprocessing includes data cleaning and data normalization; Performing feature engineering processing on the monitoring indicator time series data and the metadata after data preprocessing; The machine learning model is used to analyze the monitoring indicator time series data and the metadata after feature engineering processing.
3. The method for determining the threshold value of software system monitoring indicators according to claim 2, characterized in that: Performing feature engineering processing on the preprocessed monitoring indicator time series data and the metadata, including: Extracting the time series statistical features of the monitoring indicator time series data by rolling statistical calculation; Label encoding is used to convert text data in the metadata into numerical features.
4. The method for determining the threshold value of software system monitoring indicators according to claim 1, characterized in that: Before using the machine learning model to analyze the monitoring indicator time series data and the metadata, the method further includes: Construct the machine learning model, which includes an adaptive learning module, an optimization module and a fusion decision module. The adaptive learning module includes an XGBoost sub-model, a LightGBM sub-model and a CatBoost sub-model. The optimization module is used to search and optimize the hyperparameters of each sub-model in the adaptive learning module. The fusion decision module is used to fuse the sub-models in the adaptive learning module.
5. The method for determining the threshold value of software system monitoring indicators according to claim 4, characterized in that: Constructing the machine learning model includes: Constructing the adaptive learning module, wherein the loss function corresponding to the adaptive learning module includes a root mean square error; Constructing the optimization module, wherein the optimization module is used to optimize the adaptive learning module; Construct the fusion decision module, which is used to fuse the sub-models in the optimized adaptive learning module to obtain the machine learning model.
6. The method for determining the threshold value of software system monitoring indicators according to claim 5, characterized in that: Constructing the optimization module includes: The optimization module is constructed, and the optimization module is used to use the particle swarm algorithm to perform heuristic search optimization on the hyperparameters of each sub-model in the adaptive learning module to obtain the optimal hyperparameter combination corresponding to the adaptive learning module.
7. The method for determining the threshold value of software system monitoring indicators according to claim 5, characterized in that: Constructing the fusion decision module includes: Construct the fusion decision module, which is used to fuse the sub-models in the optimized adaptive learning module using the Stacking algorithm.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for determining the threshold value of a software system monitoring indicator according to any one of claims 1 to 7.
9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method for determining the threshold value of a software system monitoring indicator as described in any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for determining a threshold value of a software system monitoring indicator as described in any one of claims 1 to 7.
Citation Information
Cited By
Monitoring index self-adaptive full-process processing method and system
CN120492280A