Intelligent software system performance monitoring and automatic tuning method and system

By collecting and preprocessing a variety of performance data of the intelligent software system, building dynamic mode fragments and system behavior maps, combining Markov decision-making process to achieve real-time tuning, the challenges of intelligent software system performance monitoring and tuning are solved, and accurate monitoring and automatic optimization of system performance are achieved.

CN120162220APending Publication Date: 2025-06-17WUXI QITENG INTELLECTUAL PROPERTY AGENCY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510331472.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

It is difficult for the existing technology to fully and in-depth understanding of the operating status of intelligent software systems, especially under different load conditions, the dynamic changes in system performance indicators cannot be tracked in real time, resulting in the inability to deal with performance problems in a timely manner, affecting user experience and business efficiency.

Method used

By collecting a variety of runtime performance data, annotating and preprocessing, building a multi-dimensional state sequence, using an adaptive segmented aggregation algorithm to generate dynamic pattern fragments, perform clustering analysis based on pattern entropy, build a system behavior map and pattern transfer probability matrix, calculate the matching degree in real time, and calculate the optimal tuning strategy through the Markov decision-making process, and automatically adjust the system parameters to optimize performance.

Benefits of technology

It realizes comprehensive and accurate performance monitoring and automatic tuning of intelligent software systems, can promptly detect potential performance risks, avoid system failures, and improve the stability and reliability of system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162220A_ABST
    Figure CN120162220A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of system performance monitoring, and discloses an intelligent software system performance monitoring and automatic tuning method and system. The method comprises the steps of collecting performance data during operation, marking and preprocessing the performance data, and dividing the performance data into a training set and a real-time monitoring set; constructing a multi-dimensional state sequence by using the training set data, generating a dynamic mode fragment, and constructing a system behavior map and a mode transition probability matrix based on mode entropy clustering analysis; inputting the real-time monitoring set fragments into the map to calculate the matching degree, and outputting a real-time resource allocation scheme through a Markov decision process; system parameters are adjusted according to the scheme, and iterative optimization is carried out if the difference exceeds a threshold value. The system comprises a data acquisition and preprocessing module, a dynamic mode construction module, a mode optimization module, a real-time tuning decision module and an iterative optimization module. The system performance can be comprehensively monitored, intelligent automatic tuning is realized, the stability and performance of the system are improved, and the method is particularly suitable for a distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of system performance monitoring, and particularly to a method and system for intelligent software system performance monitoring and automatic tuning. Background Art

[0002] In today's digital age, intelligent software systems are widely used in various fields, from finance, healthcare to industrial manufacturing, Internet services, etc. Their performance directly affects user experience, business efficiency, and the competitiveness of enterprises. As the scale of software systems continues to expand and functions become increasingly complex, system performance monitoring and optimization face unprecedented challenges.

[0003] Traditional software system performance monitoring means are relatively single, often only able to obtain limited performance indicators, such as simple CPU usage, memory occupancy, etc., and it is difficult to comprehensively and deeply understand the running state of the system. Moreover, most of these monitoring methods are static and cannot track the dynamic changes of the system under different load conditions in real time. For example, during an e-commerce promotion event, the system load suddenly increases significantly, and traditional monitoring tools may not be able to promptly capture abnormal fluctuations in key performance indicators, resulting in the inability to quickly take measures to respond, thereby affecting the user shopping experience and even causing business losses.

[0004] In terms of performance tuning, existing methods usually rely on manual experience. Operation and maintenance personnel need to analyze and adjust the system based on long-term accumulated knowledge and practical experience. This method is inefficient and is easily affected by subjective factors. Different operation and maintenance personnel have different understandings and handling methods of system problems, and it is difficult to ensure the consistency of tuning effects. In addition, manual tuning is overwhelmed when faced with complex system architectures and massive data. Take a large-scale distributed system as an example. It contains multiple service components and a large number of nodes. Each component is interrelated and interacts with each other. Manually troubleshooting performance problems and performing tuning not only takes a lot of time and effort but may also cause new problems due to incomplete consideration.

[0005] With the development of artificial intelligence and big data technologies, some performance monitoring and tuning methods based on machine learning have gradually emerged. However, these methods still have many problems in practical applications. Some methods have overly high requirements for the quality and quantity of training data. If the data is incomplete or contains noise, it will seriously affect the accuracy and reliability of the model. Moreover, these methods often lack effective utilization of system context information and are difficult to accurately identify the associations and evolutions between system behavior patterns. For example, in complex business scenarios, system performance problems may be caused by multiple factors jointly. Existing methods cannot fully explore these potential factors, thus unable to achieve precise performance optimization. Summary of the Invention

[0006] The object of the present invention is to provide a method and system for monitoring and automatically tuning the performance of an intelligent software system, so as to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solution: A method for monitoring and automatically tuning the performance of an intelligent software system, the method comprising:

[0008] S1. Collect the runtime performance data of the intelligent software system, label the data, and clarify the system state categories corresponding to the performance data at different timestamps; perform data preprocessing, including noise filtering and normalization processing, and divide the preprocessed data into a training set and a real-time monitoring set;

[0009] S2. Construct the data in the training set into a multi-dimensional state sequence, and use an adaptive piecewise aggregation algorithm to generate multiple dynamic pattern segments containing context information;

[0010] S3. Perform clustering analysis on the dynamic pattern segments based on pattern entropy, eliminate redundant pattern segments, construct the remaining pattern segments into a system behavior map, and generate a pattern transition probability matrix;

[0011] S4. Input the dynamic pattern segments corresponding to the real-time monitoring set into the system behavior map, calculate their matching degrees with each pattern segment, and generate a matching degree distribution vector; calculate the optimal tuning strategy through a Markov decision process, and output a real-time resource allocation plan;

[0012] S5. Adjust the system parameters according to the real-time resource allocation plan, calculate the difference in performance indicators before and after the adjustment, and if the difference exceeds the dynamic fault tolerance threshold, trigger an iterative optimization process until the system performance is stable.

[0013] Preferably, in step S1, the following steps are specifically included:

[0014] S101. Collect the runtime performance data of the intelligent software system, including CPU utilization rate, memory occupancy rate, request response time, thread block rate, and communication delay between distributed nodes;

[0015] S102. Label the collected data with system state labels, and the labels include "normal load", "resource competition", "potential deadlock", and "network congestion";

[0016] S103. Use wavelet transform to filter the noise of the data, and map the data to the [0,1] interval through maximum-minimum normalization; divide the processed data into a training set and a real-time monitoring set in chronological order.

[0017] Preferably, in step S2, the following steps are specifically included:

[0018] S201. Extract the performance data within a continuous time window from the training set and construct it into a multi-dimensional state sequence as follows:

[0019] S(t,W) = [s(t - W + 1), s(t - W + 2), …, s(t)]

[0020] Wherein, represents the timestamp, represents the window length, represents the d-dimensional performance metric vector, where d is the number of dimensions of the performance metric;

[0021] S202. Perform adaptive piecewise aggregation on the multi-dimensional state sequence, dynamically adjust the piecewise length according to the variance of the data within the window, and generate multiple dynamic pattern segments;

[0022] S203. Extract statistical features for each dynamic pattern segment, including mean, variance, skewness, and kurtosis, to form an enhanced feature matrix.

[0023] Preferably, in step S3, it specifically includes the following steps:

[0024] S301. Define the pattern entropy calculation formula as follows:

[0025]

[0026] Where H(P) represents the pattern entropy of the dynamic pattern segment P, and p i represents the frequency of the i-th type of statistical feature in the dynamic pattern segment P, and k represents the total number of statistical feature categories;

[0027] S302. Use the density clustering algorithm to group the dynamic pattern segments, and eliminate redundant segments with entropy values lower than the preset threshold;

[0028] S303. Construct the retained pattern segments into a system behavior map in the form of a directed weighted graph, where the nodes represent the pattern segments, and the edge weights are calculated from the pattern transition frequencies;

[0029] S304. Generate a pattern transition probability matrix according to the transition relationship between the nodes in the system behavior map where n is the number of pattern segments.

[0030] Preferably, in step S4, it specifically includes the following steps:

[0031] S401. Extract statistical features from the dynamic pattern segments of the real-time monitoring set and perform similarity matching with the pattern segments in the system behavior map to generate a matching degree distribution vector

[0032] S402. Based on the pattern transition probability matrix M, construct a Markov decision process model, define the state space as the set of pattern fragments, and the action as the resource allocation strategy;

[0033] S403. Solve the optimal policy function π through the value iteration algorithm * , and output the real-time resource allocation scheme, including the thread pool size, cache capacity, and network bandwidth limit.

[0034] Preferably, in step S403, the state value function update formula of the value iteration algorithm is:

[0035]

[0036] where, V k+1 (s) represents the value function of state s at the (k + 1)-th iteration, R(s, a) represents the immediate reward for executing action a in state s, γ is the discount factor, M(s, a, s ′ ) represents the probability of transferring from state s by executing action a to state s ′ , and V k (s ′ ) represents the value function of state s at the k-th iteration ′ .

[0037] Preferably, in step S5, it specifically includes the following steps:

[0038] S501. Adjust the system parameters according to the real-time resource allocation scheme, collect the adjusted performance data, and calculate the difference vector between it and the expected performance

[0039] S502. Define the dynamic fault tolerance threshold vector If any dimension in E exceeds the corresponding threshold, trigger the gradient descent optimization process;

[0040] S503. Update the weight parameters in the pattern transition probability matrix M through the backpropagation algorithm, and iteratively execute steps S4 to S5 until the difference vector E converges within the threshold.

[0041] Preferably, in step S503, the weight update formula of the backpropagation algorithm is:

[0042]

[0043] where, ΔM ij is the weight change of the element M ij in the pattern transition probability matrix, η is the learning rate, is the partial derivative of the difference vector with respect to the pattern transition weight.

[0044] Preferably, the method further includes the following steps:

[0045] S6. Collaboratively analyze the performance data of multiple nodes in the distributed system, and use the federated learning framework to aggregate local model parameters to generate a global system behavior map;

[0046] S7. Optimize the resource scheduling strategy according to the global map, and allocate task loads through the consistent hashing algorithm to avoid single-point performance bottlenecks.

[0047] Preferably, the present invention further includes an intelligent software system performance monitoring and automatic tuning system, and the system includes:

[0048] Data acquisition and preprocessing module: used to collect the runtime performance data of the intelligent software system, label the data to clarify the system state categories corresponding to the performance data at different timestamps, perform data preprocessing operations such as noise filtering and normalization, and divide the preprocessed data into a training set and a real-time monitoring set;

[0049] Dynamic mode construction module: construct the data in the training set into a multi-dimensional state sequence, and use the adaptive piecewise aggregation algorithm to generate multiple dynamic mode segments containing context information;

[0050] Mode optimization module: perform clustering analysis on the dynamic mode segments based on mode entropy, eliminate redundant mode segments, construct the remaining mode segments into a system behavior map, and generate a mode transition probability matrix;

[0051] Real-time tuning decision module: input the dynamic mode segments corresponding to the real-time monitoring set into the system behavior map, calculate their matching degrees with each mode segment, generate a matching degree distribution vector, and then calculate the optimal tuning strategy through the Markov decision process to output a real-time resource allocation plan;

[0052] Iterative optimization module: adjust the system parameters according to the real-time resource allocation plan, calculate the difference in performance indicators before and after the adjustment, and if the difference exceeds the dynamic fault tolerance threshold, trigger the iterative optimization process until the system performance is stable.

[0053] Compared with the prior art, the beneficial effects of the present invention are:

[0054] By collecting various runtime performance data such as CPU utilization rate, memory occupancy rate, request response time, thread block rate, and communication delay between distributed nodes, and performing annotation and preprocessing, it can comprehensively and accurately reflect the running state of the system. For example, in an online game server system, combining these indicators can accurately determine whether the player's game experience is smooth and whether there are problems such as game lag caused by resource competition. Compared with traditional monitoring methods, the method of this patent can obtain richer information, discover potential performance hazards in a timely manner, and avoid the occurrence of system failures.

[0055] Construct the training set data into a multi-dimensional state sequence, and use the adaptive piecewise aggregation algorithm to generate dynamic pattern segments containing context information. This method can effectively capture the dynamic changes during the system operation and extract representative behavior patterns. Based on pattern entropy, perform clustering analysis and redundant pattern elimination to further optimize the pattern library, making the system behavior map more concise and accurate. Taking an e-commerce system as an example, during different shopping peak hours and promotional activities, the system's behavior patterns will be different. This method can accurately identify these differences and provide a basis for subsequent precise optimization.

[0056] Match the dynamic pattern segments of the real-time monitoring set with the system behavior map, calculate the optimal tuning strategy through the Markov decision process, and output the real-time resource allocation plan. This process realizes automated and intelligent performance tuning and can quickly respond according to the real-time state of the system. For example, when the system detects that the request response time is extended and judges that it may be caused by insufficient thread pool size, it will automatically adjust parameters such as the thread pool size, cache capacity, and network bandwidth limit to ensure that the system performance is always in the best state. Compared with manual tuning, it greatly improves the tuning efficiency and accuracy and reduces business losses caused by performance problems.

[0057] After adjusting the system parameters according to the real-time resource allocation plan, trigger the iterative optimization process by calculating the performance index difference and comparing it with the dynamic fault tolerance threshold. This closed-loop feedback mechanism can continuously optimize the system performance and make it more stable and reliable. At the same time, the backpropagation algorithm updates the weight parameters of the pattern transition probability matrix, further improving the accuracy and adaptability of the model. During long-term operation, the system can continuously self-optimize and adapt to various complex and changeable business scenarios and load conditions.

[0058] For a distributed system, adopt a federated learning framework to aggregate local model parameters to generate a global system behavior map, and allocate task loads through the consistent hashing algorithm. This effectively avoids a single-point performance bottleneck and improves the overall performance and reliability of the distributed system. For example, in a cross-regional distributed data processing system, each node can work together to jointly optimize the resource scheduling strategy to ensure that data processing tasks can be efficiently and evenly distributed to each node, improving the system's processing capacity and response speed. Description of the Drawings

[0059] Figure 1 It is the working principle diagram of the intelligent software system performance monitoring and automatic tuning method described in the present invention;

[0060] Figure 2 It is the working flow chart of pattern optimization;

[0061] Figure 3Workflow diagram for updating the value iteration algorithm;

[0062] Figure 4 It is a step diagram for triggering iterative optimization in step S5. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Please refer to Figures 1-4 , the present invention provides a technical solution: a method for monitoring and automatically tuning the performance of an intelligent software system, the method includes:

[0065] S1: Collect the runtime performance data of the intelligent software system, label the data, and clarify the system state categories corresponding to the performance data at different timestamps; perform data preprocessing, including noise filtering and normalization processing, and divide the preprocessed data into a training set and a real-time monitoring set.

[0066] S2: Construct the data in the training set into a multi-dimensional state sequence, and use the adaptive piecewise aggregation algorithm to generate multiple dynamic pattern segments containing context information.

[0067] S3: Perform clustering analysis on the dynamic pattern segments based on pattern entropy, eliminate redundant pattern segments, construct the remaining pattern segments into a system behavior map, and generate a pattern transition probability matrix.

[0068] S4: Input the dynamic pattern segments corresponding to the real-time monitoring set into the system behavior map, calculate their matching degrees with each pattern segment, generate a matching degree distribution vector; calculate the optimal tuning strategy through the Markov decision process, and output a real-time resource allocation plan.

[0069] S5: Adjust the system parameters according to the real-time resource allocation plan, calculate the difference in performance indicators before and after the adjustment, and if the difference exceeds the dynamic fault tolerance threshold, trigger the iterative optimization process until the system performance is stable.

[0070] The present invention will be further described below in conjunction with Embodiments 1 to 5:

[0071] Embodiment 1:

[0072] This embodiment elaborates in detail the specific process of data collection and preprocessing to ensure that the collected data is accurate and effective, providing a reliable basis for subsequent analysis and tuning.

[0073] During the operation of the intelligent software system, the built-in performance monitoring tool or third-party monitoring plug-in of the system is used to collect runtime performance data. These data cover multiple key metrics, including CPU utilization, memory occupancy, request response time, thread block rate, and communication latency between distributed nodes. For example, CPU utilization and memory occupancy data can be obtained through the system call interface provided by the operating system; a timer is set in the network communication module to record the time when the request is sent and the response is received, so as to calculate the request response time; with the help of the thread management mechanism, the ratio of the duration when the thread is in the blocked state to the total duration is statistically calculated to obtain the thread block rate; for the communication latency between distributed nodes, a timestamp is added to the node communication message, and the communication latency is obtained by calculating the time difference between sending and receiving.

[0074] After the data is collected, in order to clarify the system state categories corresponding to the performance data at different timestamps, system state labels are assigned to the data. The labels are divided into four categories: "normal load", "resource contention", "potential deadlock", and "network congestion". The labeling process is based on preset thresholds and rules. For example, when the CPU utilization exceeds 80% for a long time and the memory occupancy is also continuously high, it is labeled as "resource contention"; if the thread block rate rises sharply within a certain period of time and is accompanied by a significant increase in the request response time, it is labeled as "potential deadlock".

[0075] Then data preprocessing is carried out. Wavelet transform is used to filter the noise of the collected data. Wavelet transform can decompose the signal into components of different frequencies. By selecting appropriate wavelet bases and thresholds, the noise components can be effectively removed and the useful information of the data can be retained. For example, for the request response time data, there may be noise generated by factors such as network jitter. Wavelet transform can filter out these noises and make the data smoother and more accurate.

[0076] After the noise filtering is completed, the data is mapped to the [0,1] interval through min-max normalization. The min-max normalization formula is: where X is the original data, X min and X max are the minimum and maximum values in the dataset respectively, and X norm is the normalized data. The purpose of this processing is to make the data of different dimensions have the same dimension, which is convenient for subsequent analysis and comparison. For example, the value ranges of CPU utilization and memory occupancy are different. After normalization, they are analyzed on the same scale.

[0077] Finally, the processed data is divided into a training set and a real-time monitoring set in chronological order. Usually, the data over a past period is used as the training set to build the model and learn the system behavior pattern; the currently real-time collected data is used as the real-time monitoring set to monitor the system status in real time and trigger optimization operations. For example, the performance data of the past week is selected as the training set, while the real-time monitoring set is the data collected every minute currently.

[0078] Embodiment 2:

[0079] This embodiment deeply introduces the steps of dynamic pattern construction, enabling the generated dynamic pattern segments to accurately reflect the running state and behavior pattern of the system, providing strong support for subsequent analysis and decision-making.

[0080] Extract the performance data within consecutive time windows from the training set to construct a multi-dimensional state sequence. Let the timestamp window length represent a d-dimensional performance metric vector, where d is the number of dimensions of the performance metrics. For example, d = 5 (corresponding to 5 metrics: CPU utilization rate, memory occupancy rate, request response time, thread block rate, and communication latency between distributed nodes). Then the formula for constructing the multi-dimensional state sequence is:

[0081] S(t,W) = [s(t - W + 1), s(t - W + 2), …, s(t)]

[0082] Taking an online trading system as an example, if the window length W = 10 minutes and at t = 100 minutes, S(100, 10) is the sequence of 5-dimensional performance metric vectors within these 10 minutes from the 91st minute to the 100th minute.

[0083] Perform adaptive piecewise aggregation on the multi-dimensional state sequence. This process dynamically adjusts the piecewise length according to the variance of the data within the window. Variance can reflect the degree of dispersion of the data. When the data fluctuates greatly, the variance is large, and at this time, the piecewise length is appropriately shortened to more finely capture the data changes; when the data is relatively stable, the variance is small, and the piecewise length can be appropriately increased. For example, during a certain period, the request response time fluctuates violently and the variance is large, then when performing piecewise aggregation, the piecewise length for the dimension of request response time will be correspondingly shortened. In this way, multiple dynamic pattern segments are generated.

[0084] Extract statistical features for each dynamic pattern segment, including mean, variance, skewness, and kurtosis, to form an enhanced feature matrix. The mean reflects the average level of the data, the variance reflects the degree of dispersion of the data, the skewness measures the asymmetry of the data distribution, and the kurtosis describes the peakedness of the data distribution. These features can more comprehensively characterize the characteristics of the dynamic pattern segment. For example, for a dynamic pattern segment representing CPU utilization, calculating its mean can understand the average load of the CPU during that time period; the variance can reflect the fluctuation size of the CPU utilization; the skewness and kurtosis can further analyze its distribution characteristics to determine whether there are abnormal situations. Combining these statistical features together to form an enhanced feature matrix provides richer information for subsequent clustering analysis and matching degree calculation.

[0085] Example 3:

[0086] This example details the process of pattern optimization. Through pattern entropy calculation and clustering analysis, redundant pattern segments are removed, a system behavior map and a pattern transition probability matrix are constructed, providing an accurate model basis for real-time optimization decisions.

[0087] Define the pattern entropy calculation formula as follows:

[0088]

[0089] where H(P) represents the pattern entropy of the dynamic pattern segment P, p i represents the frequency of the i-th type of statistical feature in the dynamic pattern segment P, and k represents the total number of statistical feature categories. Pattern entropy is used to measure the uncertainty and complexity of the dynamic pattern segment. The higher the entropy value, the more dispersed the feature distribution of the pattern segment and the greater the uncertainty; the lower the entropy value, the more concentrated the feature distribution and the higher the certainty of the pattern segment. For example, for two different dynamic pattern segments, in one segment, the distribution of various statistical features is relatively uniform, and its pattern entropy value is higher; in the other segment, one type of statistical feature dominates and other features are less, and its pattern entropy value is lower.

[0090] Use the density clustering algorithm to group the dynamic pattern segments. The density clustering algorithm is based on the density distribution of data points and divides data points with connected densities into the same class. During the clustering process, redundant segments with entropy values lower than the preset threshold are removed. The preset threshold is determined based on experience and a large number of experiments. For example, through multiple experiments, it is found that when the pattern entropy value is lower than 0.3, the corresponding dynamic pattern segment is often repetitive or similar and contributes less to the system behavior analysis, so it can be removed. In this way, data redundancy is reduced, and the efficiency and accuracy of subsequent analysis are improved.

[0091] Construct the reserved pattern fragments into a system behavior graph in the form of a directed weighted graph. In the system behavior graph, nodes represent pattern fragments, and the edge weights are calculated from the pattern transition frequencies. The pattern transition frequency refers to the number of times of transitioning from one pattern fragment to another. For example, during the operation of the system, if pattern fragment A appears and then pattern fragment B appears 10 times, and the total number of times of transitioning from A to other pattern fragments is 50 times, then the edge weight from A to B is 10÷50 = 0.2. By constructing the system behavior graph, the relationships between different behavior patterns of the system can be visually displayed.

[0092] Generate a pattern transition probability matrix according to the transition relationships between nodes in the system behavior graph. where n is the number of pattern fragments. The element M in the pattern transition probability matrix M ij represents the probability of transitioning from pattern fragment i to pattern fragment j. For example, for a system behavior graph containing 5 pattern fragments, the pattern transition probability matrix M is a 5×5 matrix, and M 34 represents the probability of transitioning from pattern fragment 3 to pattern fragment 4. This matrix provides the key transition probability information for the subsequent Markov decision process, which is used to calculate the optimal tuning strategy.

[0093] Example 4:

[0094] This example elaborates in detail the specific implementation of real-time tuning decisions. By calculating the matching degree and using the Markov decision process, a real-time resource allocation plan is generated to achieve real-time optimization of the system performance.

[0095] Extract statistical features from the dynamic pattern fragments of the real-time monitoring set, including mean, variance, skewness, and kurtosis. The calculation methods of these features are the same as those for the dynamic pattern fragments of the training set in Example 2. Then, these statistical features are matched with the pattern fragments in the system behavior graph to generate a matching degree distribution vector. Similarity matching can use methods such as Euclidean distance and cosine similarity. Taking Euclidean distance as an example, for the dynamic pattern fragment P of the real-time monitoring set real and the pattern fragment P in the system behavior graph map , its Euclidean distance calculation formula is: where m is the number of feature dimensions, and are respectively P real and P mapThe i-th eigenvalue. The smaller the calculated Euclidean distance, the more similar the two pattern fragments are and the higher the matching degree. By calculating the similarity between the dynamic pattern fragments of the real-time monitoring set and all pattern fragments in the system behavior graph, a matching degree distribution vector V is obtained. Each element in the vector represents the matching degree between the dynamic pattern fragment of the real-time monitoring set and the corresponding pattern fragment.

[0096] Based on the pattern transition probability matrix M, a Markov decision process model is constructed. The state space is defined as the set of pattern fragments, that is, all pattern fragments in the system behavior graph; the action is the resource allocation strategy, such as adjusting the thread pool size, cache capacity, and network bandwidth limit, etc. In this model, the state transition of the system is determined by the pattern transition probability matrix M, and taking different actions will result in different rewards and system state changes.

[0097] Solve the optimal policy function π through the value iteration algorithm * , and output the real-time resource allocation plan. The update formula for the state value function of the value iteration algorithm is:

[0098]

[0099] where V k+1 (s) represents the value function of state s at the (k + 1)-th iteration, R(s, a) represents the immediate reward for executing action a in state s, γ is the discount factor (usually taking values between 0 and 1, such as 0.9), M(s, a, s ′ ) represents the probability of transferring from state s by executing action a to state s ′ , and V k (s ′ ) represents the value function of state s at the k-th iteration ′ . In practical applications, the state value function is continuously iteratively updated until convergence to obtain the optimal policy function π * . According to the optimal policy function, determine the optimal resource allocation plan under the current system state. For example, when the system is in the state corresponding to a specific pattern fragment, π * indicates adjusting the thread pool size to 100, increasing the cache capacity to 512MB, and adjusting the network bandwidth limit to 10Mbps.

[0100] Example 5:

[0101] This example details the iterative optimization process and the processing method in a distributed system to ensure continuous and stable optimization of system performance and effectively address performance issues in a distributed environment.

[0102] Adjust the system parameters according to the real-time resource allocation plan, collect the performance data after adjustment, and calculate the difference vector between it and the expected performance Expected performance can be determined based on historical data and system design goals. For example, in an e-commerce system, based on performance data from previous peak sales periods and the response time goals of the system design, expected request response time, CPU utilization and other performance indicators are determined. When calculating the difference vector, for each performance indicator dimension, the expected value is subtracted from the adjusted actual value to obtain each element of the difference vector.

[0103] Define dynamic fault tolerance threshold vector If any dimension in E exceeds the corresponding threshold, the gradient descent optimization process is triggered. The dynamic fault tolerance threshold vector is dynamically adjusted according to the system's performance requirements and actual operating conditions. For example, when the system load is low, the fault tolerance threshold can be set relatively strictly; when the system load is high or in a special operating stage (such as during a system upgrade), the fault tolerance threshold can be appropriately relaxed. When an element in the difference vector E exceeds the corresponding threshold, it means that the system performance has a large deviation and needs to be optimized.

[0104] The weight parameters in the mode transition probability matrix M are updated by the back propagation algorithm. The weight update formula of the back propagation algorithm is:

[0105]

[0106] Among them, ΔM ij is the element M in the mode transition probability matrix ij The weight change, η is the learning rate (usually a small value, such as 0.01), is the partial derivative of the difference vector with respect to the mode transfer weight. By calculating the partial derivative, the influence of each weight in the mode transfer probability matrix on the performance difference is determined, and then the weight is adjusted according to the learning rate to make the mode transfer probability matrix more consistent with the actual behavior of the current system. Iteratively execute steps S4 to S5, that is, recalculate the matching degree between the dynamic mode fragment of the real-time monitoring set and the system behavior map, calculate the new resource allocation plan through the Markov decision process, adjust the system parameters and calculate the performance difference until the difference vector E converges to the threshold to ensure the stability of system performance.

[0107] For the collaborative analysis of the performance data of multiple nodes in a distributed system, a federated learning framework is adopted to aggregate local model parameters and generate a global system behavior map. In a distributed environment, each node has its own local performance data and local model. The federated learning framework allows each node to aggregate and generate a global model by exchanging model parameters without sharing the original data. For example, each node constructs a local system behavior map and a pattern transition probability matrix based on the locally collected performance data, and then uploads the model parameters to a central server. The central server uses a federated learning algorithm, such as the FedAvg algorithm, to aggregate the parameters of each node and generate a global system behavior map. This can comprehensively consider the behavior patterns of each node in the distributed system and more comprehensively understand the overall performance of the system.

[0108] Optimize the resource scheduling strategy according to the global map, and allocate task loads through the consistent hashing algorithm to avoid single-point performance bottlenecks. The consistent hashing algorithm evenly distributes tasks to each node in the distributed system. When a node fails or has too high a load, it can automatically re-allocate tasks to other nodes, thus avoiding single-point performance bottlenecks. For example, in a distributed file storage system, the consistent hashing algorithm is used to distribute file storage tasks to each storage node. According to the performance status and load conditions of each node in the global system behavior map, dynamically adjust the task allocation strategy to ensure the optimal overall performance of the system.

[0109] The present invention also includes an intelligent software system performance monitoring and automatic tuning system, and the system includes:

[0110] Data acquisition and preprocessing module: used to acquire the runtime performance data of the intelligent software system, annotate the data to clarify the system state categories corresponding to the performance data at different timestamps, perform data preprocessing operations such as noise filtering and normalization, and divide the preprocessed data into a training set and a real-time monitoring set;

[0111] Dynamic mode construction module: constructs the data in the training set into a multi-dimensional state sequence, and uses an adaptive piecewise aggregation algorithm to generate multiple dynamic mode segments containing context information;

[0112] Mode optimization module: performs clustering analysis on the dynamic mode segments based on mode entropy, eliminates redundant mode segments, constructs the remaining mode segments into a system behavior map, and generates a pattern transition probability matrix;

[0113] Real-time tuning decision module: inputs the dynamic mode segments corresponding to the real-time monitoring set into the system behavior map, calculates their matching degrees with each mode segment, generates a matching degree distribution vector, and then calculates the optimal tuning strategy through a Markov decision process and outputs a real-time resource allocation plan;

[0114] Iterative Optimization Module: Adjust the system parameters according to the real-time resource allocation scheme, calculate the difference in performance metrics before and after the adjustment. If the difference exceeds the dynamic fault tolerance threshold, trigger the iterative optimization process until the system performance stabilizes.

[0115] The implementation of this system refers to the above embodiments and will not be elaborated in the description of the specification.

[0116] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0117] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for monitoring and automatically optimizing performance of an intelligent software system, characterized in that: The method comprises the following steps: S1. Collect runtime performance data of intelligent software systems, annotate the data, and clarify the system status categories corresponding to performance data at different timestamps; perform data preprocessing, including noise filtering and normalization, and divide the preprocessed data into training sets and real-time monitoring sets; S2. The data in the training set is constructed into a multi-dimensional state sequence, and an adaptive segmentation aggregation algorithm is used to generate multiple dynamic pattern segments containing context information; S3. Perform cluster analysis on dynamic pattern fragments based on pattern entropy, remove redundant pattern fragments, construct the retained pattern fragments into a system behavior map, and generate a pattern transition probability matrix; S4. Input the dynamic pattern fragments corresponding to the real-time monitoring set into the system behavior map, calculate the matching degree between the dynamic pattern fragments and each pattern fragment, and generate a matching degree distribution vector; calculate the optimal tuning strategy through the Markov decision process, and output the real-time resource allocation plan; S5. Adjust system parameters according to the real-time resource allocation plan, calculate the difference in performance indicators before and after the adjustment, and if the difference exceeds the dynamic fault tolerance threshold, trigger the iterative optimization process until the system performance is stable.

2. The method according to claim 1, characterized in that In step S1, the following steps are specifically included: S101. Collect runtime performance data of intelligent software systems, including CPU utilization, memory occupancy, request response time, thread blocking rate, and communication delay between distributed nodes; S102. Label the collected data with system status tags, including "normal load", "resource competition", "potential deadlock" and "network congestion"; S103. Use wavelet transform to filter the noise of the data, and map the data to the [0,1] interval through maximum and minimum value normalization; divide the processed data into a training set and a real-time monitoring set in chronological order.

3. The method according to claim 1, characterized in that In step S2, the following steps are specifically included: S201. Extract performance data in a continuous time window from the training set and construct it into a multi-dimensional state sequence as follows: S(t,W)=[s(t-W+1),s(t-W+2),…,s(t)] in, Indicates the timestamp, represents the window length, represents a d-dimensional performance indicator vector, d is the number of dimensions of the performance indicator; S202. Adaptively aggregate the multidimensional state sequence into segments, dynamically adjust the segment length according to the variance of the data in the window, and generate multiple dynamic pattern segments; S203. Extract statistical features from each dynamic pattern segment, including mean, variance, skewness and kurtosis, to form an enhanced feature matrix.

4. The method according to claim 1, characterized in that: In step S3, the following steps are specifically included: S301. Define the formula for calculating pattern entropy as follows: Where H(P) represents the pattern entropy of the dynamic pattern segment P, p i represents the frequency of the i-th type of statistical features in the dynamic pattern segment P, and k represents the total number of statistical feature categories; S302. Using a density clustering algorithm to group dynamic pattern segments, and removing redundant segments whose entropy values ​​are lower than a preset threshold; S303. Construct the retained pattern fragments into a system behavior map in the form of a directed weighted graph, where the nodes represent the pattern fragments and the edge weights are calculated by the frequency of pattern transitions; S304. Generate a mode transition probability matrix based on the transition relationship between nodes in the system behavior graph Where n is the number of pattern fragments.

5. The method according to claim 1, characterized in that In step S4, the following steps are specifically included: S401. Extract statistical features from the dynamic pattern fragments of the real-time monitoring set, and perform similarity matching with the pattern fragments in the system behavior map to generate a matching degree distribution vector S402. Based on the mode transition probability matrix M, a Markov decision process model is constructed, and the state space is defined as a set of mode fragments, and the action is a resource allocation strategy; S403. Solve the optimal strategy function π through value iteration algorithm * , output real-time resource allocation plan, including thread pool size, cache capacity and network bandwidth limit.

6. The method according to claim 5, characterized in that In step S403, the state value function update formula of the value iteration algorithm is: Among them, V k+1 (s) represents the value function of state s at the k+1th iteration, R(s,a) represents the immediate benefit of executing action a in state s, γ is the discount factor, and M(s,a,s ′ ) means that after executing action a from state s, it transfers to state s ′ The probability of V k (s ′ ) represents the state s at the kth iteration ′ The value function of .

7. The method according to claim 1, characterized in that In step S5, the following steps are specifically included: S501. Adjust system parameters according to the real-time resource allocation plan, collect performance data after adjustment, and calculate the difference vector between it and the expected performance S502. Define dynamic fault tolerance threshold vector If any dimension in E exceeds the corresponding threshold, the gradient descent optimization process is triggered; S503. Update the weight parameters in the mode transition probability matrix M through the back propagation algorithm, and iterate steps S4 to S5 until the difference vector E converges to within the threshold.

8. The method according to claim 7, characterized in that In step S503, the weight update formula of the back propagation algorithm is: Among them, ΔM ij is the element M in the mode transition probability matrix ij The weight change, η is the learning rate, is the partial derivative of the difference vector with respect to the mode transfer weight.

9. The method according to claim 1, characterized in that: The following steps are also included: S6. Collaboratively analyze the performance data of multiple nodes in the distributed system, aggregate local model parameters using a federated learning framework, and generate a global system behavior map; S7. Optimize resource scheduling strategies based on the global graph, distribute task loads through consistent hashing algorithms, and avoid single-point performance bottlenecks.

10. An intelligent software system performance monitoring and automatic tuning system, characterized in that: include: Data collection and preprocessing module: used to collect runtime performance data of intelligent software systems, annotate data to clarify the system status category corresponding to performance data of different timestamps, perform data preprocessing operations such as noise filtering and normalization, and divide the preprocessed data into training sets and real-time monitoring sets; Dynamic pattern construction module: constructs the data in the training set into a multi-dimensional state sequence, and uses an adaptive segmentation aggregation algorithm to generate multiple dynamic pattern segments containing contextual information; Pattern optimization module: clusters dynamic pattern fragments based on pattern entropy, removes redundant pattern fragments, constructs the retained pattern fragments into a system behavior map, and generates a pattern transition probability matrix; Real-time tuning decision module: Input the dynamic pattern fragments corresponding to the real-time monitoring set into the system behavior map, calculate the matching degree with each pattern fragment, generate the matching degree distribution vector, and then calculate the optimal tuning strategy through the Markov decision process, and output the real-time resource allocation plan; Iterative optimization module: adjusts system parameters according to the real-time resource allocation plan, calculates the difference in performance indicators before and after the adjustment, and triggers the iterative optimization process if the difference exceeds the dynamic fault tolerance threshold until the system performance is stable.

Citation Information

Cited By

  • Metabolite fingerprint spectrum construction method of probiotic fermentation type Chinese herbal medicine preparation

    CN121164509A