Rule engine-based micro-service performance optimization diagnosis method
By constructing a microservice performance optimization diagnostic method based on a rule engine, the problem of neglecting inter-service dependencies in traditional monitoring methods is solved, enabling more accurate performance detection and fault location, providing optimization suggestions, and improving the stability and response speed of microservice systems.
Patent Information
- Application Number
- CN202411311053.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional microservice monitoring methods focus on the performance metrics of individual services while neglecting the interdependencies and influences between services, making it difficult to comprehensively monitor microservice performance and quickly locate problems.
A rule engine-based approach is adopted to construct a comprehensive rule base that includes rules for single performance indicators and rules for the correlation between multi-dimensional performance indicators. Big data analysis technology is used to reveal the potential correlation between different performance indicators, and rule parameters are automatically adjusted by combining rule learning models to achieve data preprocessing, rule base updates, and fault location.
It enables more comprehensive and accurate detection of performance problems, reduces false alarms and missed alarms, quickly locates faults, provides targeted optimization suggestions, improves the accuracy and timeliness of diagnosis, and reduces manual troubleshooting time and costs.
Smart Images

Figure CN120929994A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microservice architecture and performance optimization, and more specifically, to a microservice performance optimization diagnostic method based on a rule engine. Background Technology
[0002] With the acceleration of enterprise digital transformation, microservice architecture has become the mainstream approach for building modern cloud-native applications due to its flexibility, scalability, and independent deployment advantages. Microservice architecture improves development efficiency and system maintainability by decomposing applications into a series of small, loosely coupled services. However, this architecture also brings new challenges, especially in performance monitoring and fault diagnosis. The distributed nature of microservices leads to complex interactions between services and potential performance bottlenecks, placing higher demands on monitoring tools.
[0003] Traditional monitoring methods often focus on the performance metrics of a single service, neglecting the interdependencies and impacts between services. In a microservice environment, a performance issue in one service can quickly affect the stability and responsiveness of the entire system. Therefore, a solution is needed that can comprehensively monitor microservice performance, analyze inter-service relationships in real time, and quickly pinpoint problems. Summary of the Invention
[0004] The purpose of this invention is to provide a microservice performance optimization and diagnostic method based on a rule engine, in order to solve the problem that traditional monitoring methods mentioned in the background technology often focus on the performance indicators of a single service, while ignoring the interdependence and influence between services.
[0005] To achieve the above objectives, the present invention aims to provide a microservice performance optimization and diagnostic method based on a rule engine, comprising the following steps:
[0006] S1. Collect monitoring data from microservices and preprocess the data;
[0007] S2. Construct a comprehensive rule base that includes rules for single performance indicators and rules for the association between multi-dimensional performance indicators, and reveal the potential associations between different performance indicators through big data analysis technology;
[0008] S3. Match the collected data with the rules defined in the rule base. If the data meets the conditions of a certain rule, an alarm is generated.
[0009] As a further improvement to this technical solution, the specific steps of S1 are as follows:
[0010] S1.1 Use Prometheus to periodically crawl data from the data source. The data should include at least CPU usage, memory usage, network latency, and request response time.
[0011] S1.2 Use the standard score model to check and clean up outliers in the data, and use the missing value imputation model to fill in missing values;
[0012] S1.3. Use a normalization model to process the collected data in a unified format.
[0013] As a further improvement to this technical solution, the standard fraction mold body in S1.2 is:
[0014]
[0015] Where, x i Represents a single observation; μ represents the mean of the observations; σ represents the standard deviation of the observations; Z represents the standardized score.
[0016] Specifically, the missing fill model in S1.2 is as follows:
[0017]
[0018] Where, x i Represents the observed value; n represents the number of valid observations; mean and median represent the mean and median, respectively;
[0019] Specifically, the normalization model in S1.3 is as follows:
[0020]
[0021] Where x represents the original observation value; x′ represents the normalized value; min(x) and max(x) represent the minimum and maximum values of the observation value, respectively.
[0022] As a further improvement to this technical solution, the specific steps of S2 are as follows:
[0023] S2.1 Determine the normal operating range of each performance indicator and set thresholds;
[0024] S2.2. Use a correlation analysis model to mine the potential correlations between different performance indicators from historical monitoring data;
[0025] S2.3. Rules extracted through actual microservice environment testing are adjusted based on test results;
[0026] S2.4 Store the validated rules in a queryable rule base;
[0027] S2.5 Design a rule learning model so that the rule base can automatically adjust rule parameters according to the system's operating status, thereby improving the accuracy and timeliness of diagnosis.
[0028] As a further improvement to this technical solution, the correlation analysis model in S2.2 is specifically as follows:
[0029]
[0030] Where, r xy Represents the correlation coefficient; x i and y i These represent the observed values of variables X and Y, respectively. and These represent the average values of variables X and Y, respectively.
[0031] As a further improvement to this technical solution, the rule learning model in S2.5 is specifically as follows:
[0032]
[0033] Where Q(s, a) represents the expected reward of performing action a in state s; s represents the current performance index state of the microservice; a represents the optimization measure taken; r represents the effect evaluation after taking the optimization measure; α represents the learning rate; γ represents the discount rate for future rewards; s′ represents the performance index state after taking the optimization measure; a′ represents the further optimization measures that may be taken in the new state. This represents the expected effect of taking the optimal action under the new state.
[0034] As a further improvement to this technical solution, the specific steps of S3 are as follows:
[0035] S3.1 Match the collected data with the rules defined in the rule base;
[0036] S3.2 If a performance indicator in the monitoring data exceeds a preset threshold, or if the correlation between multiple performance indicators reaches a preset correlation coefficient, an alarm will be triggered.
[0037] S3.3 If the rule matching is successful, the fault location model is used to locate the specific location of the fault.
[0038] S3.4 Based on the fault location results, the system generates optimization suggestions for specific problems.
[0039] As a further improvement to this technical solution, the fault location model in S3.3 is specifically as follows:
[0040]
[0041] Where L represents the comprehensive set of all fault locations; m represents the number of rules in the rule base; The set of fault locations representing the threshold rule matching for the i-th single performance metric; The set of fault locations representing the association rule matching among the i-th multiple performance metrics;
[0042] in, The calculation formula is:
[0043]
[0044] Where j represents the index in the monitoring data D; d j σ represents the observed value of the performance metric; D represents the monitoring dataset; i The θ represents the direction of the threshold rule for the performance metric; i Thresholds representing performance metrics:
[0045] in, The calculation formula is:
[0046]
[0047] Where j represents the index of the observation of the first performance metric; k represents the index of the observation of the second performance metric; d j The observed value representing the first performance metric; d k The observed value representing the second performance metric; μ j μ represents the average value of the first performance metric. k ρ represents the average value of the second performance metric. ij This represents the correlation coefficient threshold between the first and second performance metrics; this formula is used to determine the matching of association rules between multiple performance metrics.
[0048] As a further improvement to this technical solution, the specific steps of S3.4 are as follows:
[0049] S3.4.1 Build a database containing known failure modes and corresponding solutions;
[0050] S3.4.2 Based on the comprehensive fault location set L obtained in step S3.3, determine which service instances and components have failed, and the correlation between these faults;
[0051] S3.4.3. Use a matching algorithm to match the identified faults with known faults in the database;
[0052] S3.4.4 If a match is successful, the corresponding solution will be retrieved from the database, and specific optimization suggestions and operation guidelines will be generated.
[0053] As a further improvement to this technical solution, the matching algorithm in S3.4.3 is specifically as follows:
[0054]
[0055] Where |F ∩L| represents the size of the intersection between fault mode F and the currently determined set of fault locations L, i.e. the number of fault locations that the two have in common; |F| represents the total number of fault locations in fault mode F; |L| represents the total number of fault locations in the currently determined set of fault locations L; min(|F|, |L|) represents preventing excessively high scores due to a large number of fault locations in F or L.
[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0057] 1. By constructing a comprehensive rule base that includes rules for single performance indicators and rules for the correlation between multi-dimensional performance indicators, this invention can more comprehensively and accurately detect potential performance problems and faults, reducing the possibility of false alarms and false negatives. Through a rule learning model, this invention can automatically adjust rule parameters according to the system's operating status, enabling the rule base to continuously evolve, thereby improving the accuracy and timeliness of diagnosis and responding promptly to system changes.
[0058] 2. The fault location model proposed in this invention can effectively locate the specific location of the fault, reduce the time and cost of manual troubleshooting, and speed up the fault repair process. By matching the identified fault with known fault patterns in the database using a matching algorithm, this invention can provide targeted optimization suggestions to help maintenance personnel allocate resources more rationally and avoid resource waste. Attached Figure Description
[0059] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] Example 1:
[0062] Please see Figure 1 As shown, this embodiment provides a microservice performance optimization and diagnostic method based on a rule engine, including the following steps:
[0063] 1. A microservice performance optimization and diagnostic method based on a rule engine, characterized by the following steps:
[0064] S1. Collect monitoring data from microservices and preprocess the data.
[0065] The specific steps of S1 are as follows:
[0066] S1.1 Use Prometheus to periodically crawl data from the data source. The data should include at least CPU usage, memory usage, network latency, and request response time.
[0067] S1.2 Use the standard score model to check and clean up outliers in the data, and use the missing value imputation model to fill in missing values;
[0068] S1.3. Use a normalization model to process the collected data in a unified format.
[0069] The standard fraction mold in S1.2 is:
[0070]
[0071] Where, x i σ represents a single observation; μ represents the mean of the observations; σ represents the standard deviation of the observations; Z represents the standardized score; this formula is used to transform the raw data into a standard normal distribution, so that the mean of the data is 0 and the standard deviation is 1. Data points that deviate from the mean by more than 3 standard deviations are considered outliers.
[0072] Specifically, the missing fill model in S1.2 is as follows:
[0073]
[0074] Where, x i The values represent the observed values; n represents the number of valid observations; mean and median represent the mean and median, respectively; imputation is used to handle missing data points in the dataset.
[0075] Specifically, the normalization model in S1.3 is as follows:
[0076]
[0077] Where x represents the original observation; x′ represents the normalized value; min(x) and max(x) represent the minimum and maximum values of the observation, respectively; normalization scales the data to a specific range, which helps to eliminate differences in data scale, allowing data of different scales to be compared and processed at the same level.
[0078] S2. Construct a comprehensive rule base that includes rules for single performance indicators and rules for the association between multi-dimensional performance indicators. Use big data analytics to reveal the potential relationships between different performance indicators and achieve more comprehensive and accurate problem diagnosis.
[0079] The specific steps of S2 are as follows:
[0080] S2.1 Determine the normal operating range of each performance metric (such as CPU utilization, memory usage, etc.) and set thresholds. For example, if CPU utilization consistently exceeds 80%, it may indicate a performance bottleneck.
[0081] S2.2. Use correlation analysis models to mine the potential correlations between different performance indicators from historical monitoring data. For example, high CPU utilization may be accompanied by high memory utilization or high network latency.
[0082] S2.3. The rules extracted through actual microservice environment testing are adjusted based on the test results to make them more accurate;
[0083] S2.4 Store the verified rules in a queryable rule base. As new monitoring data is collected and analyzed, the rule base should be updated regularly to adapt to changes and developments in the system.
[0084] S2.5 Design a rule learning model so that the rule base can automatically adjust rule parameters according to the system's operating status, thereby improving the accuracy and timeliness of diagnosis.
[0085] The specific correlation analysis model in S2.2 is as follows:
[0086]
[0087] Where, r xy This represents the correlation coefficient, which ranges from -1 to +1. Positive values indicate a positive correlation, negative values indicate a negative correlation, and the larger the absolute value, the stronger the correlation. i and y i Let x represent the observed values of variables X and Y, respectively. i and y i These are measurements of different performance metrics (such as CPU utilization, memory usage, network latency, etc.). and These represent the average values of variables X and Y, respectively. By calculating this formula, we can quantify the strength of the relationship between two performance metrics. This is very useful for identifying potential performance bottlenecks and the mutual influence between microservices. For example, if a strong positive correlation is found between CPU utilization X and request response time Y, it may be necessary to further investigate whether the slow response is due to excessive CPU utilization.
[0088] The rule learning model in S2.5 is as follows:
[0089]
[0090] Where Q(s, a) represents the expected benefit of performing action a in state s; s represents the current performance index state of the microservice; a represents the optimization measures taken, such as adjusting resource allocation and optimizing load balancing; r represents the effect evaluation after taking optimization measures, such as the degree of performance improvement; α represents the learning rate, which determines how much new information will be used to update the Q value, and is used to weigh the importance of new information against old information; γ represents the discount rate for future rewards, used to balance short-term and long-term benefits; s′ represents the performance index state after taking optimization measures; a′ represents the further optimization measures that may be taken in the new state. This represents the expected effect of taking the optimal action under a new state. By continuously updating the Q value, the system can learn which optimization measures are most effective in which states and dynamically adjust the rule parameters according to the actual situation. In this way, the rule base can continuously evolve, improving the accuracy and timeliness of diagnosis.
[0091] S3. Match the collected data with the rules defined in the rule base. If the data meets the conditions of a certain rule, an alarm is generated.
[0092] The specific steps for S3 are as follows:
[0093] S3.1 Match the collected data with the rules defined in the rule base;
[0094] S3.2 If a performance indicator in the monitoring data exceeds a preset threshold, or if the correlation between multiple performance indicators reaches a preset correlation coefficient, an alarm is triggered. The alarm usually contains detailed information about the triggering rules, such as which performance indicator exceeds the normal range, or which performance indicators have an abnormal correlation.
[0095] S3.3 If the rule matching is successful, the fault location model is used to locate the specific location of the fault.
[0096] S3.4 Based on the fault location results, the system generates optimization suggestions for specific problems.
[0097] The fault location model in S3.3 is as follows:
[0098]
[0099] Where L represents the comprehensive set of all fault locations; m represents the number of rules in the rule base; The set of fault locations representing the threshold rule matching for the i-th single performance metric; L represents the set of fault locations matching the association rules between the i-th multiple performance indicators; this formula is used to combine the threshold rules of all single performance indicators and the results of the association rule matching between multiple performance indicators to form a comprehensive set of fault locations L.
[0100] in, The calculation formula is:
[0101]
[0102] Where j represents the index in the monitoring data D; d j σ represents the observed value of the performance metric; D represents the monitoring dataset; i The θ represents the direction of the threshold rule for the performance metric; i The threshold represents a performance metric; this formula is used to determine when a single performance metric exceeds the threshold. Specifically, it identifies which observations d in the monitoring dataset... j Threshold rules that satisfy a single performance metric;
[0103] in, The calculation formula is:
[0104]
[0105] Where j represents the index of the observation of the first performance metric; k represents the index of the observation of the second performance metric; d j The observed value representing the first performance metric; d k The observed value representing the second performance metric; μ j μ represents the average value of the first performance metric. k ρ represents the average value of the second performance metric. ij This represents the correlation coefficient threshold between the first and second performance metrics; this formula is used to determine the matching of association rules among multiple performance metrics, specifically, it identifies which observations d in the monitoring dataset... j and d k It satisfies the association rules between multiple performance indicators.
[0106] The specific steps in S3.4 are as follows:
[0107] S3.4.1 Build a database containing known failure modes and corresponding solutions;
[0108] S3.4.2 Based on the comprehensive fault location set L obtained in step S3.3, determine which service instances and components have failed, and the correlation between these faults;
[0109] S3.4.3. Use a matching algorithm to match the identified faults with known faults in the database;
[0110] S3.4.4 If a match is successful, the corresponding solution will be retrieved from the database, and specific optimization suggestions and operation guidelines will be generated.
[0111] The matching algorithm in S3.4.3 is as follows:
[0112]
[0113] Where |F ∩L| represents the size of the intersection between fault mode F and the currently determined set of fault locations L, i.e., the number of fault locations shared by both; |F| represents the total number of fault locations in fault mode F; |L| represents the total number of fault locations in the currently determined set of fault locations L; min(|F|, |L|) represents preventing excessively high scores due to a large number of fault locations in F or L; in this way, the best matching fault mode in the database can be found, and corresponding solutions can be provided based on the matching fault mode.
[0114] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A microservice performance optimization and diagnostic method based on a rule engine, characterized by: Includes the following steps: S1. Collect monitoring data from microservices and preprocess the data; S2. Construct a comprehensive rule base that includes rules for single performance indicators and rules for the association between multi-dimensional performance indicators, and reveal the potential associations between different performance indicators through big data analysis technology; S3. Match the collected data with the rules defined in the rule base. If the data meets the conditions of a certain rule, an alarm is generated.
2. The microservice performance optimization and diagnostic method based on a rule engine according to claim 1, characterized in that: The specific steps of S1 are as follows: S1.1 Use Prometheus to periodically crawl data from the data source. The data should include at least CPU usage, memory usage, network latency, and request response time. S1.2 Use the standard score model to check and clean up outliers in the data, and use the missing value imputation model to fill in missing values; S1.
3. Use a normalization model to process the collected data in a unified format.
3. The microservice performance optimization and diagnostic method based on a rule engine according to claim 2, characterized in that: The standard fraction mold body in S1.2 is: Where, x i Represents a single observation; μ represents the mean of the observations; σ represents the standard deviation of the observations; Z represents the standardized score. Specifically, the missing fill model in S1.2 is as follows: median = median(x1, x2, ..., x...) n ) Where, x i Represents the observed value; n represents the number of valid observations; mean and median represent the mean and median, respectively; Specifically, the normalization model in S1.3 is as follows: Where x represents the original observation value; x′ represents the normalized value; min(x) and max(x) represent the minimum and maximum values of the observation value, respectively.
4. The microservice performance optimization and diagnostic method based on a rule engine according to claim 1, characterized in that: The specific steps of S2 are as follows: S2.1 Determine the normal operating range of each performance indicator and set thresholds; S2.
2. Use a correlation analysis model to mine the potential correlations between different performance indicators from historical monitoring data; S2.
3. Rules extracted through actual microservice environment testing are adjusted based on test results; S2.4 Store the validated rules in a queryable rule base; S2.5 Design a rule learning model so that the rule base can automatically adjust rule parameters according to the system's operating status, thereby improving the accuracy and timeliness of diagnosis.
5. The microservice performance optimization and diagnostic method based on a rule engine according to claim 4, characterized in that: The specific correlation analysis model in S2.2 is as follows: Where, r xy Represents the correlation coefficient; x i and y i These represent the observed values of variables X and Y, respectively. and These represent the average values of variables X and Y, respectively.
6. The microservice performance optimization and diagnostic method based on a rule engine according to claim 4, characterized in that: The rule learning model in S2.5 is specifically as follows: Where Q(s, a) represents the expected reward of performing action a in state s; s represents the current performance index state of the microservice; a represents the optimization measure taken; r represents the effect evaluation after taking the optimization measure; α represents the learning rate; γ represents the discount rate for future rewards; s′ represents the performance index state after taking the optimization measure; a′ represents the further optimization measures that may be taken in the new state. This represents the expected effect of taking the optimal action under the new state.
7. The microservice performance optimization and diagnostic method based on a rule engine according to claim 1, characterized in that: The specific steps of S3 are as follows: S3.1 Match the collected data with the rules defined in the rule base; S3.2 If a performance indicator in the monitoring data exceeds a preset threshold, or if the correlation between multiple performance indicators reaches a preset correlation coefficient, an alarm will be triggered. S3.3 If the rule matching is successful, the fault location model is used to locate the specific location of the fault. S3.4 Based on the fault location results, the system generates optimization suggestions for specific problems.
8. The microservice performance optimization and diagnostic method based on a rule engine according to claim 7, characterized in that: The fault location model in S3.3 is specifically as follows: Where L represents the comprehensive set of all fault locations; m represents the number of rules in the rule base; The set of fault locations representing the threshold rule matching for the i-th single performance metric; The set of fault locations representing the association rule matching among the i-th multiple performance metrics; in, The calculation formula is: Where j represents the index in the monitoring data D; d j σ represents the observed value of the performance metric; D represents the monitoring dataset; i The θ represents the direction of the threshold rule for the performance metric; i Thresholds representing performance metrics; in, The calculation formula is: Where jj represents the index of the observation of the first performance metric; k represents the index of the observation of the second performance metric; d j The observed value representing the first performance metric; d k The observed value representing the second performance metric; μ j μ represents the average value of the first performance metric. k p represents the average value of the second performance metric. ij This represents the correlation coefficient threshold between the first and second performance metrics; this formula is used to determine the matching of association rules between multiple performance metrics.
9. The microservice performance optimization and diagnostic method based on a rule engine according to claim 7, characterized in that: The specific steps of S3.4 are as follows: S3.4.1 Build a database containing known failure modes and corresponding solutions; S3.4.2 Based on the comprehensive fault location set L obtained in step S3.3, determine which service instances and components have failed, and the correlation between these faults; S3.4.
3. Use a matching algorithm to match the identified faults with known faults in the database; S3.4.4 If a match is successful, the corresponding solution will be retrieved from the database, and specific optimization suggestions and operation guidelines will be generated.
10. The microservice performance optimization and diagnostic method based on a rule engine according to claim 7, characterized in that: The matching algorithm in S3.4.3 is as follows: Where |F∩L| represents the size of the intersection between fault mode F and the currently determined set of fault locations L, i.e. the number of fault locations that the two have in common; |F| represents the total number of fault locations in fault mode F; |L| represents the total number of fault locations in the currently determined set of fault locations L; min(|F|, |L|) represents preventing excessively high scores due to a large number of fault locations in F or L.
Citation Information
Patent Citations
Block generation method and device, storage medium and electronic equipment
CN114036237A
Monitoring method and device of micro-service system, storage medium and program product
CN114844771A
Application and cloud platform cross-layer fault analysis method and system based on micro-service deployment
CN116719664A
Fault monitoring method and device, electronic equipment, medium and product
CN116962144A
Comprehensive performance monitoring method based on cloud platform
CN117033158A
Cited By
SoC performance problem automatic locating and attribution method, device and equipment
CN122470431A