Automatic performance monitoring and tuning method and device for intelligent operation and maintenance platform
Through the automated performance monitoring and tuning methods of the intelligent operation and maintenance platform, using deep learning and Markov decision-making processes, the problems of incomplete monitoring and tuning dependence on manual experience in traditional operation and maintenance methods in large-scale and high-complex scenarios are solved, achieving efficient and stable operation and maintenance results.
Patent Information
- Application Number
- CN202510152747.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional operation and maintenance methods cannot effectively deal with large-scale and high-complex operation and maintenance scenarios, performance monitoring is not comprehensive and accurate enough, tuning depends on manual experience, it is difficult to ensure the results, and it is difficult to adapt to the dynamic and complex environments brought by cloud computing, big data and other technologies.
Using an intelligent operation and maintenance platform, by obtaining the performance parameters, configuration parameters and environmental condition data of the operation and maintenance objects, using the performance tuning decision model trained by the deep deterministic strategy gradient algorithm, combined with the Markov decision-making process and the long-term and short-term memory neural network, automated performance monitoring and tuning are achieved.
It realizes multi-dimensional and comprehensive performance monitoring and tuning, improves operation and maintenance efficiency, reduces the risk of manual intervention, and ensures the stable and efficient operation of the information system.
Smart Images

Figure CN120276930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of performance monitoring and tuning, and particularly relates to an automated performance monitoring and tuning method and device for an intelligent operation and maintenance platform. Background Art
[0002] With the rapid development of information technology, the scale and complexity of various information systems and operation and maintenance objects (such as servers, databases, network devices, application programs, etc.) have increased day by day, which poses higher requirements for operation and maintenance management. Traditional operation and maintenance methods mainly rely on manual monitoring and tuning, which are not only inefficient but also difficult to handle large-scale and high-complexity operation and maintenance scenarios. Especially when facing sudden performance problems or the need to continuously optimize system performance, traditional operation and maintenance methods are often powerless.
[0003] In current information system operation and maintenance practices, performance monitoring and tuning are two crucial links. Performance monitoring aims to obtain the running state data of operation and maintenance objects in real time and timely discover performance bottlenecks or anomalies, while performance tuning aims to improve system performance by adjusting the configuration parameters of operation and maintenance objects or taking other optimization measures to ensure the stable and efficient operation of the system.
[0004] However, traditional performance monitoring and tuning methods have many deficiencies. First, in terms of performance monitoring, traditional monitoring tools often can only provide single performance index data and lack the comprehensive analysis ability of multi-dimensional data such as environmental conditions and operation and maintenance object configuration parameters, resulting in incomplete and inaccurate monitoring results. Second, in terms of performance tuning, traditional tuning methods mainly rely on the experience and intuition of operation and maintenance personnel and lack scientific and systematic tuning strategies, resulting in difficult-to-guarantee tuning effects and even possible new problems.
[0005] In addition, with the rise of technologies such as cloud computing, big data, and artificial intelligence, the environmental conditions and configuration parameters of operation and maintenance objects have become more dynamic and complex, and traditional operation and maintenance methods are no longer able to adapt to this change. Therefore, there is an urgent need for an intelligent and automated performance monitoring and tuning method to achieve comprehensive, real-time, and accurate monitoring and tuning of operation and maintenance objects, improve operation and maintenance efficiency, reduce operation and maintenance costs, and ensure the stable and efficient operation of information systems. Summary of the Invention
[0006] In order to overcome the above defects, the present invention proposes an automated performance monitoring and tuning method and device for an intelligent operation and maintenance platform.
[0007] In a first aspect, an automated performance monitoring and tuning method for an intelligent operation and maintenance platform is provided, and the automated performance monitoring and tuning method for the intelligent operation and maintenance platform includes:
[0008] Obtain the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0009] Use the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as the input of a pre-trained performance tuning decision model to obtain a tuning operation strategy output by the pre-trained performance tuning decision model;
[0010] Use the tuning operation strategy to control the operation and maintenance object.
[0011] Preferably, the performance parameters include at least one of the following: CPU usage rate, memory occupancy rate, disk I / O, network bandwidth, response time, throughput, error rate, and system log; the environmental conditions include at least one of the following: temperature, humidity, power supply stability; the configuration parameters include at least one of the following: hardware configuration, software version, system architecture, network configuration, and security settings.
[0012] Preferably, the pre-trained performance tuning decision model is trained using the deep deterministic policy gradient algorithm.
[0013] Preferably, the state space corresponding to the pre-trained performance tuning decision model includes: the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object, and the action space corresponding to the pre-trained performance tuning decision model includes a set of tuning operations.
[0014] Further, the set of tuning operations includes: increasing or decreasing computing resources, adjusting system or application configuration parameters, restarting or resetting services, migrating loads, adjusting network configuration, optimizing database performance, adjusting power management strategies, implementing load balancing strategies, and adjusting security settings.
[0015] Preferably, the reward function corresponding to the pre-trained performance tuning decision model is as follows:
[0016] R = w1 * R_perf + w2 * R_util + w3 * R_stab + w4 * R_cost + w5 * R_dev
[0017] In the above formula, R is the reward function value, R_perf represents the reward term for performance improvement, R_util represents the reward term for resource utilization optimization, R_stab represents the reward term for enhanced stability, R_cost represents the reward term for reduced operation and maintenance costs, R_dev represents the reward term or penalty term for reduced performance deviation, and w1, w2, w3, w4, w5 are the weight coefficients of each sub-reward term respectively.
[0018] Further, the process of obtaining the performance deviation includes:
[0019] Obtain the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0020] Taking the configuration parameters and environmental condition data as the input of a pre-trained performance prediction model to obtain the predicted performance parameter value output by the pre-trained performance prediction model;
[0021] Timestamp-align the performance parameters of the operation and maintenance object with the predicted performance parameter value, and use the deviation value exceeding the threshold as the performance deviation.
[0022] Further, the training process of the pre-trained performance prediction model includes:
[0023] Obtaining the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0024] Constructing training data based on the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0025] Training a long short-term memory neural network using the training data to obtain the pre-trained performance prediction model.
[0026] Further, during the process of training the long short-term memory neural network using the training data, the backpropagation algorithm and the gradient descent method are used to optimize the model parameters.
[0027] Further, during the process of optimizing the model parameters using the gradient descent method, the model parameters are updated according to the following formula:
[0028]
[0029] In the above formula, θ represents the model parameters, ω represents the learning rate, represents the gradient of the loss function J(θ) with respect to the parameter θ, and := represents the assignment operation.
[0030] In a second aspect, there is provided an intelligent operation and maintenance platform automated performance monitoring and tuning device, and the intelligent operation and maintenance platform automated performance monitoring and tuning device includes:
[0031] An acquisition module, configured to acquire the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0032] An analysis module, configured to take the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as the input of a pre-trained performance tuning decision model to obtain a tuning operation strategy output by the pre-trained performance tuning decision model;
[0033] A regulation module, configured to regulate the operation and maintenance object using the tuning operation strategy.
[0034] In a third aspect, there is provided a computer device, including: one or more processors;
[0035] The processor is used to store one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the automated performance monitoring and tuning method of the intelligent operation and maintenance platform is implemented.
[0037] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the automated performance monitoring and tuning method of the intelligent operation and maintenance platform is implemented.
[0038] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:
[0039] The present invention relates to the technical field of performance monitoring and tuning. Specifically, an automated performance monitoring and tuning method and device for an intelligent operation and maintenance platform are provided, including: obtaining performance parameters, configuration parameters, and environmental condition data of an operation and maintenance object; using the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as inputs to a pre-trained performance tuning decision model to obtain a tuning operation strategy output by the pre-trained performance tuning decision model; and using the tuning operation strategy to regulate the operation and maintenance object. By collecting time series data of performance parameters of the operation and maintenance object during historical operation, as well as environmental conditions and configuration parameters during operation, the present invention provides a multi-dimensional and comprehensive data basis, making performance monitoring more comprehensive and accurate. The data preprocessing and feature extraction steps ensure the cleanliness and consistency of the data, further improving the accuracy of monitoring and prediction.
[0040] The constructed performance prediction model can predict the change trend of the performance parameters of the operation and maintenance object in the future period based on historical data and current environmental conditions, providing a forward-looking performance warning for operation and maintenance personnel, facilitating timely measures to prevent potential problems. Real-time obtaining of the current operation state data and environmental condition data of the operation and maintenance object and comparing them with the predicted performance parameters, through timestamp alignment and curve comparison, realizes the accurate identification of performance deviations, providing strong support for quickly responding to performance problems.
[0041] Modeling the performance tuning process as a Markov decision process, combined with the reward function and policy selection designed according to the operation and maintenance objectives, realizes scientific and systematic tuning decisions. The initial simple rule-based policy ensures the feasibility of the tuning process. As the model matures, it gradually transitions to a deep learning-based policy, further improving the intelligence level and effect of tuning. The automated performance monitoring and tuning method significantly improves the operation and maintenance efficiency and reduces the risk of manual intervention and misoperation. By optimizing performance, improving resource utilization, enhancing stability, and reducing operation and maintenance costs, the present invention provides strong guarantee for the efficient and stable operation of information systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the main step process of the intelligent operation and maintenance platform automation performance monitoring and tuning method in an embodiment of the present invention;
[0043] Figure 2 It is a design schematic diagram of the reward function in an embodiment of the present invention;
[0044] Figure 3 It is a training flow chart of the performance prediction model in an embodiment of the present invention. Detailed implementation manners
[0045] The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1
[0048] The intelligent operation and maintenance platform automation performance monitoring and tuning method in the embodiments of the present invention mainly includes the following steps:
[0049] Obtain the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0050] Use the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as the input of a pre-trained performance tuning decision model to obtain a tuning operation strategy output by the pre-trained performance tuning decision model;
[0051] Use the tuning operation strategy to regulate and control the operation and maintenance object.
[0052] In one embodiment, referring to the attached Figure 1 , Figure 1 It is a schematic diagram of the main step process of the intelligent operation and maintenance platform automation performance monitoring and tuning method in an embodiment of the present invention. As Figure 1 shown, the method includes:
[0053] Step 1: Through the monitoring agents deployed on the operation and maintenance objects or by using existing monitoring tools, comprehensively collect the time series data of performance parameters of various operation and maintenance objects (including but not limited to servers, databases, network devices, application programs, etc.) during historical operation. These data include but are not limited to key performance indicators such as CPU usage rate, memory occupancy rate, disk I / O, network throughput, etc. At the same time, it is also necessary to collect the environmental condition data during operation, such as environmental temperature, humidity, power supply status, etc., and the configuration parameters of the operation and maintenance objects, such as server model, operating system version, application program configuration, etc.
[0054] Step 2: Preprocess the collected data. This includes data cleaning to remove invalid, duplicate or incorrect data; data format unification to ensure seamless docking of data from different sources; and data standardization to convert the original data into a format acceptable to the model. After preprocessing, perform feature extraction to extract valuable features for performance prediction and optimization from the original data to form a feature dataset. The method for standardizing the collected data is: use the Z-score method to subtract the mean of each numerical data and divide it by its standard deviation, so that the processed data conforms to the standard normal distribution. The formula for standardization is: Z = (X - μ) / σ, where X is the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data.
[0055] Step 3: Based on the feature dataset, combined with environmental conditions and operation and maintenance object configuration parameters, build a performance prediction model. This model can adopt machine learning algorithms such as time series analysis, regression analysis, etc. to predict the change trend of performance parameters of operation and maintenance objects in the future period. After the model training is completed, it is necessary to conduct verification and testing to ensure its prediction accuracy.
[0056] Step 4: Involve real-time acquisition of the current operation status data and environmental condition data of the operation and maintenance objects. This can be collected in real time through monitoring agents and transmitted to the operation and maintenance platform through the network.
[0057] Step 5: Input the real-time acquired data into the performance prediction model and receive the predicted performance parameters output by the model. These predicted parameters will be used to compare with the actual performance parameters to evaluate the performance status of the operation and maintenance objects.
[0058] Step 6: Conduct real-time performance monitoring on the operation and maintenance objects and obtain the actual performance parameter data. At the same time, perform anomaly detection on the original monitoring data, and use statistical methods or machine learning algorithms to identify and eliminate or correct outliers to ensure the accuracy and reliability of the data.
[0059] Step 7: Plot the real-time performance parameter data and the predicted performance parameter data as curves varying with time respectively, and perform timestamp alignment. Then, place the two curves in the same coordinate system for comparison to visually observe the differences between the actual performance and the predicted performance.
[0060] Step 8: Involve setting a performance deviation threshold and determining whether there is a deviation beyond the threshold for each corresponding point on the two curves according to this threshold. If there is a deviation, mark it as a deviation point for subsequent tuning reference.
[0061] Step 9: Build a performance tuning decision model. Model the performance tuning process of the operation and maintenance object as a Markov decision process, and clearly define the state space, action space, reward function, and policy. The state space includes the combination of real-time performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object; the action space is the set of executable tuning operations; the reward function is designed according to the operation and maintenance objectives, including reward mechanisms for performance improvement, resource utilization optimization, stability enhancement, and operation and maintenance cost reduction, and combines the performance deviation data monitored in Step 8 to give additional rewards to operations that reduce deviations and additional penalties to operations that increase deviations; the policy initially adopts a simple rule-based policy, and gradually transitions to a deep learning-based policy as the model matures and data accumulates to improve the intelligence and accuracy of decision-making.
[0062] Step 10: Perform real-time tuning based on the trained performance tuning decision model. Collect the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object in real time as the model input. According to the current state and policy, the model selects the optimal tuning action and automatically executes the decision result or transmits it to the operation and maintenance personnel through the operation and maintenance platform for confirmation and then execution. In this way, through continuous performance monitoring and tuning, ensure that the operation and maintenance object always maintains the best performance state.
[0063] The present invention will be further described below:
[0064] The present invention provides an intelligent operation and maintenance platform automated performance monitoring and tuning method, aiming to ensure the efficient and stable operation of the operation and maintenance object through comprehensive performance monitoring and intelligent tuning decisions. The implementation manner of this method will be described in detail below in combination with specific examples and problems to be solved.
[0065] First, in terms of performance monitoring, the present invention defines a wide range of performance parameter sets, including but not limited to CPU usage rate, memory occupancy rate, disk I / O, network bandwidth, response time, throughput, error rate, and system logs. These parameters can comprehensively reflect the operating status and performance level of the operation and maintenance object. At the same time, the impact of environmental conditions on the performance of the operation and maintenance object is also considered, such as the temperature, humidity, and power supply stability of the operating environment, as well as the configuration parameters of the operation and maintenance object, including hardware configuration, software version, system architecture, network configuration, and security settings. These data and parameters are collected in real time through monitoring agents deployed on the operation and maintenance object or by using existing monitoring tools, and are transmitted to the operation and maintenance platform for unified processing and analysis.
[0066] Taking the server cluster of a large e-commerce platform as an example, key performance indicators such as the CPU usage rate and memory occupancy rate of each server, as well as environmental condition data such as the temperature and humidity of the computer room, are collected in real time through monitoring agents. At the same time, configuration parameters such as the hardware configuration and operating system version of the server are recorded to provide data support for subsequent performance prediction and optimization.
[0067] In terms of constructing a performance prediction model, the present invention constructs a performance prediction model based on the collected historical performance parameter time series data, environmental conditions, and configuration parameters by using machine learning algorithms (such as time series analysis, regression analysis, etc.). This model can predict the change trend of the performance parameters of the operation and maintenance object in the future period of time, providing forward-looking performance warnings and decision-making bases for operation and maintenance personnel.
[0068] Continuing with the example of the e-commerce platform server cluster, a performance prediction model is trained using historical performance data and environmental condition data. The training process is as Figure 3 shown. The model can predict the change trend of performance indicators such as the CPU usage rate and memory occupancy rate of each server in the next week, helping operation and maintenance personnel to identify potential performance bottlenecks and failure risks in advance.
[0069] In terms of performance optimization, the present invention defines a series of executable optimization operation sets, including increasing or decreasing computing resources, adjusting system or application configuration parameters, restarting or resetting services, migrating loads, adjusting network configuration, optimizing database performance, adjusting power management strategies, implementing load balancing strategies, and adjusting security settings, etc. These operations can cover all aspects of the performance optimization of the operation and maintenance object, meeting different operation and maintenance requirements.
[0070] Regarding the performance issues of the e-commerce platform server cluster, the operation and maintenance personnel can select corresponding tuning operations for performance optimization based on the output results of the performance prediction model and real-time monitoring data. For example, when it is predicted that the CPU usage rate of a certain server will continue to increase in the next few days and may reach the bottleneck, the operation and maintenance personnel can increase the number of CPU cores of the server in advance or adjust configuration parameters such as the thread pool size to avoid the occurrence of performance bottlenecks.
[0071] To guide the formulation and execution of tuning decisions, the present invention also designs a performance tuning decision model based on the Markov decision process. By defining elements such as the state space, action space, reward function, and policy, the performance tuning process of the operation and maintenance object is modeled as a sequential decision-making problem. Among them, the reward function is the core of the tuning decision. As Figure 2 shown, it consists of multiple sub-reward items, including performance improvement, resource utilization optimization, stability enhancement, operation and maintenance cost reduction, and performance deviation reduction. These sub-reward items are weighted and summed to obtain the total reward value, which is used to evaluate the effects of different tuning operations.
[0072] Specifically, the formula for the reward function R is: R = w1 * R_perf + w2 * R_util + w3 * R_stab + w4 * R_cost + w5 * R_dev. Among them, R_perf represents the reward item for performance improvement, which is calculated according to the improvement degree of the performance indicators KPIs (such as response time, throughput, etc.) of the operation and maintenance object; R_util represents the reward item for resource utilization optimization, which is calculated according to the improvement degree of resource usage efficiency (such as CPU, memory utilization, etc.); R_stab represents the reward item for stability enhancement, which is calculated according to the duration of stable operation of the system or service and the indicator of reduced failure rate; R_cost represents the reward item for operation and maintenance cost reduction, which is calculated according to the cost savings of reduced resource consumption and reduced operation and maintenance operation frequency; R_dev represents the reward item or penalty item for performance deviation reduction. According to the performance deviation data monitored in step 8, a positive reward is given to the operation that reduces the deviation, and a penalty is given to the operation that increases the deviation. w1, w2, w3, w4, and w5 are the weight coefficients of each sub-reward item, respectively, which are used to adjust the proportion of different reward items in the total reward value.
[0073] In practical applications, operation and maintenance personnel can adjust the weight coefficients of each sub-reward item according to specific operation and maintenance objectives and requirements to obtain a reward function that better conforms to the actual situation. Then, based on the trained performance tuning decision model, the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object are collected in real time as the model input. The model selects the optimal tuning action according to the current state and policy, and automatically executes the decision result or transmits it to the operation and maintenance personnel through the operation and maintenance platform for confirmation and then execution. In this way, through continuous performance monitoring and tuning, it is ensured that the operation and maintenance object always maintains the best performance state, improving operation and maintenance efficiency, reducing operation and maintenance costs, and ensuring the stable and efficient operation of the information system.
[0074] The present invention uses the Deep Deterministic Policy Gradient (DDPG) algorithm to train the performance tuning decision model, which is illustrated by a specific example as follows:
[0075] The training of the performance tuning decision model specifically includes the following steps:
[0076] S1: Initialize the actor network and the critic network
[0077] Construct two deep neural networks, namely the actor network and the critic network. The actor network is responsible for generating resource allocation actions according to the current system state, and its output layer is a continuous action space, representing the allocation ratio of different server resources. The critic network is used to evaluate the value of a given state-action pair, that is, to predict the cumulative reward that can be obtained after taking this action, and its output is a scalar value. Both networks adopt deep neural network structures, such as multi-layer perceptrons (MLP), and initialize the network weights.
[0078] S2: Initialize the experience replay buffer D
[0079] Set the size of the experience replay buffer D to N, for example, N = 10000. This buffer is used to store the state transition experiences generated during the performance tuning process of the intelligent operation and maintenance platform. Each experience includes the current state s, the action a taken, the reward r obtained, the next state s', and the termination flag done. The state s can include performance metrics such as the CPU usage rate, memory occupancy rate, and network bandwidth of the server; the action a is the amount of resource allocation adjustment; the reward r is calculated according to the degree of system performance improvement; and done indicates whether the current round is over.
[0080] S3: For each training round
[0081] S301: Initialize the state s at the start of the round
[0082] Obtain the performance metrics of the current server from the intelligent operation and maintenance platform as the state s at the start of the round.
[0083] S302: Before the round ends, loop and execute the following steps: i) The actor network generates an action a: Input the current state s into the actor network to obtain the resource allocation action a.
[0084] ii) Execute the action and observe the result: Execute the action a on the intelligent operation and maintenance platform, that is, adjust the server resource allocation, and then observe the change of the system's performance metrics to obtain the reward r, the next state s', and whether the termination condition is reached (such as reaching the maximum number of iterations or the system performance is stable).
[0085] iii) Store the experience: Store the experience (s, a, r, s', done) in the experience replay buffer D.
[0086] iv) Experience replay and sampling: Randomly sample a batch of experiences from the buffer D for updating the network parameters.
[0087] v) Update the critic network: Use the sampled experiences, through the backpropagation algorithm and the gradient descent method, to update the parameters of the critic network to minimize the value function prediction error. The value function can use the mean squared error loss function.
[0088] vi) Update the actor network: Use the sampled experiences and the evaluation results of the critic network, through the backpropagation algorithm and the gradient ascent method, to update the parameters of the actor network to maximize the expected value of the cumulative reward. Specifically, the goal of the actor network is to maximize the Q value given by the critic network.
[0089] vii) Update the state: Update the current state s to the next state s', and continue the next round of iteration.
[0090] S303: Reset the state when the round ends
[0091] When the termination condition is reached (such as done = True), reset the state s to the starting state of a new round, that is, re-obtain the performance metrics of the current server from the intelligent operation and maintenance platform.
[0092] S4: Performance evaluation and parameter adjustment
[0093] During the training process, every certain number of rounds (such as every 100 rounds), perform a performance evaluation on the performance tuning decision model. The evaluation method can be simulation testing or application in the actual operation and maintenance environment, and calculate the proportion of system performance improvement. According to the evaluation results, adjust the weight coefficient of the reward function (such as increasing the weight of the performance improvement item), the learning rate of the model (such as gradually decreasing the learning rate as the training progresses), and the magnitude of the exploration noise (such as setting a larger noise at the beginning to promote exploration and decreasing the noise later to stabilize the training).
[0094] S5: Repeat training
[0095] Repeat steps S3 and S4 until a preset number of training rounds (e.g., 10,000 rounds) is reached, or the model performance reaches a preset goal (e.g., the system performance improvement reaches a specific threshold).
[0096] In order to build a performance prediction model based on Long Short-Term Memory (LSTM) and achieve automated performance monitoring and tuning in the intelligent operation and maintenance platform, the following will elaborate on the specific implementation method and training steps of the model, and illustrate it with a practical application scenario.
[0097] ① Model construction:
[0098] Input layer design: The input layer is responsible for receiving the feature data set, which covers the performance metric data of time series (such as CPU usage, memory occupancy, disk I / O, etc.), environmental parameter data (such as temperature, humidity, etc.), system configuration data (such as CPU model, memory capacity, etc.), and time-related features (such as hour, day of the week, month, etc.). After preprocessing, these data form a sequence of fixed length as the input of the LSTM model.
[0099] Hidden layer design: The hidden layer contains at least one LSTM hidden layer, and each layer consists of several LSTM units. Inside the LSTM unit, there are a forget gate, an input gate, and an output gate. These gating mechanisms can effectively control the flow of information, enabling the LSTM to capture long-term dependencies. After the hidden layer, a ReLU (Rectified Linear Unit) activation function is connected to perform a non-linear transformation on the output of the LSTM unit to enhance the model's expressive ability.
[0100] Output layer design: The output layer receives the output of the hidden layer and maps the output of the LSTM network to the predicted performance metric value through a fully connected layer or a regression layer. For the performance prediction task, the mean squared error (MSE) or the mean absolute error (MAE) is usually used as the loss function to measure the difference between the model prediction value and the true value.
[0101] ② Model training:
[0102] Data division: Divide the feature data set into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for model validation and parameter tuning, and the test set is used for model evaluation. Usually, the training set accounts for 60%-80% of the entire data set, and the validation set and the test set each account for 10%-20%.
[0103] Parameter setting: Set the number of network layers, the number of hidden layer units in each layer, the activation function, and the optimizer. The number of network layers and the number of hidden layer units are adjusted according to the complexity of the task and the data scale. The activation function selects ReLU, and the optimizer can select Adam, SGD, etc.
[0104] Iterative training: The LSTM model is iteratively trained using the training set. In each iteration, a batch of feature datasets is fed into the LSTM model, and the loss between the predicted value and the true value of the model is calculated. The model parameters are optimized through the backpropagation algorithm and the gradient descent method. The specific update formula is as follows:
[0105]
[0106] where θ represents the model parameters, including all weights and biases that the model needs to learn; ω represents the learning rate, which is used to control the step size of parameter update; represents the gradient of the loss function J(θ) with respect to the parameter θ, is a vector pointing in the direction where the loss function grows fastest; := represents the assignment operation, that is, updating the value of the parameter θ; The iteration process continues until a preset number of iterations is reached or the loss converges.
[0107] Model validation and tuning: During the training process, the validation set is used to validate the model and evaluate its performance. According to the validation results, hyperparameters such as the network structure, learning rate, and batch size are adjusted to obtain better model performance.
[0108] Model evaluation: The trained model is evaluated using the test set, and metrics such as the prediction accuracy, recall rate, and F1 score of the model are calculated to evaluate the generalization ability of the model.
[0109] Taking server performance prediction as an example, a feature dataset is constructed by collecting historical performance metric data of the server (such as CPU usage rate, memory occupancy rate, etc.), environmental parameter data (such as computer room temperature, humidity, etc.), and system configuration data (such as CPU model, memory capacity, etc.). The above LSTM performance prediction model is used to predict the future performance of the server. Through the prediction results, performance bottlenecks can be discovered in advance, and resource scheduling or configuration optimization can be carried out to ensure the stable operation of the server.
[0110] To achieve fine-grained monitoring and timely tuning of system performance, this embodiment details the specific implementation method of step 8, which includes setting a performance deviation threshold and performing deviation detection and marking.
[0111] Setting the performance deviation threshold
[0112] To accurately determine whether there are abnormal fluctuations in system performance, it is first necessary to set reasonable deviation thresholds for each performance metric. This step is based on the following considerations:
[0113] Historical performance monitoring data: Analyze the performance monitoring data of the system over a period of time in the past to understand the normal fluctuation range of each performance metric.
[0114] System Requirements: Based on the performance index requirements of the system design, determine the standard values or expected ranges that each performance index should achieve.
[0115] Operation and Maintenance Experience: Combine the experience of the operation and maintenance team, consider the special situations or performance bottlenecks that the system may encounter during actual operation, and appropriately adjust the deviation threshold.
[0116] When setting the deviation threshold, it is necessary to ensure that performance anomalies can be detected in a timely manner and that alarms are not frequently triggered due to normal performance fluctuations. For example, for the performance index of CPU utilization rate, a deviation threshold can be set, such as ±5%. That is, when the real-time CPU utilization rate differs from the predicted value or the normal value by more than 5%, it is regarded as an anomaly.
[0117] Deviation Detection and Marking
[0118] After setting the performance deviation threshold, the next step is to perform deviation detection and marking. The specific steps are as follows:
[0119] Data Alignment: First, ensure that the real-time performance monitoring data and the predicted data are aligned in time, that is, there is corresponding real-time data and predicted data for each timestamp.
[0120] Deviation Calculation: Traverse the aligned data, compare the real-time performance monitoring data and the predicted data for each timestamp, and calculate the deviation value between the two. The deviation value can be obtained through simple difference calculation or determined through more complex statistical methods (such as standard deviation, percentile, etc.).
[0121] Deviation Judgment: Compare the calculated deviation value with the set performance deviation threshold. If the deviation value at a certain timestamp exceeds the set threshold, it is determined that there is a performance anomaly at that timestamp.
[0122] Marking and Recording: For the timestamps with performance anomalies, mark them as deviation points and record the specific deviation values, corresponding performance indicators, and occurrence times. This information will be used for subsequent fault troubleshooting, performance tuning, and the generation of operation and maintenance reports.
[0123] Taking the monitoring of the CPU utilization rate of a certain server as an example, assume that the set deviation threshold is ±5%. During a certain monitoring period, the real-time performance monitoring data shows that the CPU utilization rate at a certain timestamp is 75%, while the predicted data is 68%. The calculated deviation value is 7%, which exceeds the set threshold. Therefore, this timestamp is marked as a deviation point, and the deviation value of 7%, the performance indicator of CPU utilization rate, and the occurrence time of the corresponding specific time of this timestamp are recorded. The operation and maintenance personnel can, based on this information, promptly troubleshoot the server, find out the reason for the abnormal increase in CPU utilization rate, and take corresponding tuning measures to ensure the stability of the server performance.
[0124] Example 2
[0125] Based on the same inventive concept, the present invention also provides an intelligent operation and maintenance platform automatic performance monitoring and tuning device, and the intelligent operation and maintenance platform automatic performance monitoring and tuning device includes:
[0126] An acquisition module, configured to acquire performance parameters, configuration parameters, and environmental condition data of an operation and maintenance object;
[0127] An analysis module, configured to use the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as inputs of a pre-trained performance tuning decision model, and obtain a tuning operation strategy output by the pre-trained performance tuning decision model;
[0128] A regulation module, configured to regulate the operation and maintenance object by using the tuning operation strategy.
[0129] Preferably, the performance parameters include at least one of the following: CPU usage rate, memory occupancy rate, disk I / O, network bandwidth, response time, throughput, error rate, and system log; the environmental conditions include at least one of the following: temperature, humidity, power supply stability; the configuration parameters include at least one of the following: hardware configuration, software version, system architecture, network configuration, and security settings.
[0130] Preferably, the pre-trained performance tuning decision model is trained by using a deep deterministic policy gradient algorithm.
[0131] Preferably, the state space corresponding to the pre-trained performance tuning decision model includes: performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object, and the action space corresponding to the pre-trained performance tuning decision model includes a set of tuning operations.
[0132] Further, the set of tuning operations includes: increasing or decreasing computing resources, adjusting system or application configuration parameters, restarting or resetting services, migrating loads, adjusting network configuration, optimizing database performance, adjusting power management strategies, implementing load balancing strategies, and adjusting security settings.
[0133] Preferably, the reward function corresponding to the pre-trained performance tuning decision model is as follows:
[0134] R = w1 * R_perf + w2 * R_util + w3 * R_stab + w4 * R_cost + w5 * R_dev
[0135] In the above formula, R is the reward function value, R_perf represents the reward item for performance improvement, R_util represents the reward item for resource utilization optimization, R_stab represents the reward item for stability enhancement, R_cost represents the reward item for reducing operation and maintenance costs, R_dev represents the reward item or penalty item for reducing performance deviation, and w1, w2, w3, w4, and w5 are the weight coefficients of each sub-reward item respectively.
[0136] Furthermore, the process of obtaining the performance deviation includes:
[0137] Obtain the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0138] Use the configuration parameters and environmental condition data as the input of a pre-trained performance prediction model to obtain the predicted performance parameter value output by the pre-trained performance prediction model;
[0139] Align the time stamps of the performance parameters of the operation and maintenance object with the predicted performance parameter values, and use the deviation values exceeding the threshold as the performance deviation.
[0140] Furthermore, the training process of the pre-trained performance prediction model includes:
[0141] Obtain the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0142] Construct training data based on the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object;
[0143] Use the training data to train a long short-term memory neural network to obtain the pre-trained performance prediction model.
[0144] Furthermore, during the process of using the training data to train the long short-term memory neural network, the backpropagation algorithm and the gradient descent method are used to optimize the model parameters.
[0145] Furthermore, during the process of using the gradient descent method to optimize the model parameters, the model parameters are updated according to the following formula:
[0146]
[0147] In the above formula, θ represents the model parameters, ω represents the learning rate, represents the gradient of the loss function J(θ) with respect to the parameter θ, and := represents the assignment operation.
[0148] Embodiment 3
[0149] Based on the same inventive concept, the present invention also provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of an intelligent operation and maintenance platform automatic performance monitoring and optimization method in the above embodiments.
[0150] Embodiment 4
[0151] Based on the same inventive concept, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of an intelligent operation and maintenance platform automatic performance monitoring and optimization method in the above embodiments.
[0152] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still modifications or equivalent substitutions can be made to the specific embodiments of the present invention, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An automated performance monitoring and tuning method for an intelligent operation and maintenance platform, characterized in that, The method includes: Obtaining the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object; Using the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as the input of a pre-trained performance tuning decision model to obtain a tuning operation strategy output by the pre-trained performance tuning decision model; Using the tuning operation strategy to control the operation and maintenance object.
2. The method according to claim 1, wherein The performance parameters include at least one of the following: CPU usage rate, memory occupancy rate, disk I / O, network bandwidth, response time, throughput, error rate, and system log; the environmental conditions include at least one of the following: temperature, humidity, power supply stability; the configuration parameters include at least one of the following: hardware configuration, software version, system architecture, network configuration, and security settings.
3. The method according to claim 1, wherein The pre-trained performance tuning decision model is trained using the deep deterministic policy gradient algorithm.
4. The method according to claim 1, wherein The state space corresponding to the pre-trained performance tuning decision model includes: the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object, and the action space corresponding to the pre-trained performance tuning decision model includes a set of tuning operations.
5. The method according to claim 4, characterized in that, The set of tuning operations includes: increasing or decreasing computing resources, adjusting system or application configuration parameters, restarting or resetting services, migrating loads, adjusting network configuration, optimizing database performance, adjusting power management strategies, implementing load balancing strategies, and adjusting security settings.
6. The method according to claim 1, characterized in that The reward function corresponding to the pre-trained performance tuning decision model is as follows: R = w1 * R_perf + w2 * R_util + w3 * R_stab + w4 * R_cost + w5 * R_dev In the above formula, R is the reward function value, R_perf represents the reward item for performance improvement, R_util represents the reward item for resource utilization optimization, R_stab represents the reward item for enhanced stability, R_cost represents the reward item for reduced operation and maintenance costs, R_dev represents the reward item or penalty item for reduced performance deviation, and w1, w2, w3, w4, w5 are the weight coefficients of each sub-reward item respectively.
7. The method according to claim 6, wherein The process of obtaining the performance deviation includes: Obtaining the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object; Using the configuration parameters and environmental condition data as the input of a pre-trained performance prediction model to obtain a predicted value of the performance parameters output by the pre-trained performance prediction model; Aligning the time stamps of the performance parameters of the operation and maintenance object with the predicted values of the performance parameters, and taking the deviation values exceeding the threshold as the performance deviation.
8. The method according to claim 7, wherein The training process of the pre-trained performance prediction model includes: Obtaining the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object; Constructing training data based on the historical performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object; Using the training data to train a long short-term memory neural network to obtain the pre-trained performance prediction model.
9. The method according to claim 8, characterized in that, During the process of training the long short-term memory neural network using the training data, the backpropagation algorithm and the gradient descent method are used to optimize the model parameters.
10. The method according to claim 8, characterized in that During the process of optimizing the model parameters using the gradient descent method, the model parameters are updated according to the following formula: In the above formula, θ represents the model parameter, ω represents the learning rate, represents the gradient of the loss function J(θ) with respect to the parameter θ, and := represents the assignment operation.
11. An apparatus for the automated performance monitoring and tuning method of the intelligent operation and maintenance platform according to any one of claims 1-10, characterized in that, The device includes: An acquisition module, configured to acquire performance parameters, configuration parameters, and environmental condition data of an operation and maintenance object; An analysis module, configured to use the performance parameters, configuration parameters, and environmental condition data of the operation and maintenance object as inputs to a pre-trained performance tuning decision model, and obtain a tuning operation strategy output by the pre-trained performance tuning decision model; A regulation module, configured to regulate the operation and maintenance object by using the tuning operation strategy.
12. A computer device, characterized in that, It includes: One or more processors; The processor is configured to execute one or more programs; When the one or more programs are executed by the one or more processors, the intelligent operation and maintenance platform automated performance monitoring and tuning method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, the intelligent operation and maintenance platform automated performance monitoring and tuning method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Terminal software comprehensive analysis method and system based on artificial intelligence
CN120448247A