Middleware performance optimization method and device, electronic equipment and storage medium
By combining distributed probe components and predictive models with knowledge base rules to optimize middleware configuration, the problem of middleware performance optimization relying on manual operation is solved, and real-time monitoring and automated optimization of middleware are realized, thereby improving system stability and resource utilization.
Patent Information
- Application Number
- CN202510880797.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-28
AI Technical Summary
Middleware performance optimization relies on manual operation, which makes it difficult to detect performance faults in a timely manner, and configuration optimization lacks scientific basis, making it difficult to achieve optimal performance.
The system collects performance metrics through a distributed probe component, predicts performance inflection points using a predictive model, and generates optimization strategies by combining knowledge base rules and middleware dependency graphs, automatically adjusting middleware configurations to achieve optimal performance.
It enables real-time monitoring and optimization of middleware performance, improves system stability and resource utilization, and reduces false alarm rate and reliance on manual intervention.
Smart Images

Figure CN120849098A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular to a middleware performance optimization method, apparatus, electronic device, and storage medium. Background Technology
[0002] In backend systems, the performance and stability of middleware such as Redis, MySQL, and MQ directly impact the overall system operation. Traditional middleware optimization and monitoring rely on manual operation, which can easily fail to detect performance faults in a timely manner, leading to Redis memory overflows. Furthermore, configuration optimization based on human experience lacks scientific basis; approximately 70% of middleware is not configured optimally. For example, Redis frequently experiences Out of Memory (OOM) errors due to the lack of a memory eviction policy. These experience-based approaches make it difficult for middleware to achieve optimal performance. Summary of the Invention
[0003] This application provides a middleware performance optimization method, apparatus, electronic device, and storage medium to address the problem that middleware struggles to achieve optimal performance.
[0004] In a first aspect, this application provides a middleware performance optimization method, the method comprising:
[0005] The performance metrics of each middleware runtime are collected through a distributed probe component.
[0006] The performance indicators are predicted using a predictive model, and the performance inflection point when the prediction result reaches a dynamic threshold is determined, wherein the dynamic threshold is calculated by learning the latest business data online.
[0007] Based on the preset knowledge base rules and middleware dependency graph, the optimization strategy for each middleware at the corresponding performance inflection point is determined. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of changes to the current middleware on upstream and downstream middleware.
[0008] The middleware is optimized based on the optimization strategy to achieve optimal performance.
[0009] Optionally, the performance index is predicted using a prediction model, and the performance inflection point at which the prediction result reaches a dynamic threshold is determined, including:
[0010] The performance metrics of the middleware are input into the prediction model to obtain the performance prediction values within a future set time period. The Long Short-Term Memory network in the prediction model is used to capture the time-series dependencies of the performance metrics, and the Random Forest in the prediction model is used to handle nonlinear feature interactions.
[0011] The predicted performance value is compared with the corresponding dynamic threshold;
[0012] If the predicted performance value reaches the corresponding dynamic threshold, then the middleware is determined to have reached a performance inflection point.
[0013] Optionally, the optimization strategy for each middleware at the corresponding performance inflection point is determined comprehensively based on preset knowledge base rules and middleware dependency graphs, including:
[0014] Generate at least one single-point strategy for the middleware at the corresponding performance inflection point based on preset knowledge base rules;
[0015] Based on the middleware dependency graph, the impact of the single-point strategy on upstream and downstream middleware is analyzed, and at least one linkage strategy is generated, wherein the linkage strategy is used to optimize the upstream and downstream middleware.
[0016] The at least one single-point strategy and the at least one linkage strategy are combined into a candidate strategy set;
[0017] The expected return of each candidate strategy in the candidate strategy set is evaluated by reinforcement learning, and the candidate strategy with the highest return is selected as the optimization strategy, wherein the expected return is used to indicate the extent of improvement in response time or the degree of resource consumption of the candidate strategy.
[0018] Optionally, determining the dynamic threshold includes:
[0019] If the prediction result is a normally distributed index, then the 3σ principle is used to determine the alarm threshold of the prediction result;
[0020] If the prediction result is a non-normal indicator, the alarm threshold of the prediction result is determined by the quantile method;
[0021] Through online learning, the alarm threshold is recalculated at fixed time intervals based on the latest business data.
[0022] Optionally, optimizing the corresponding middleware based on the optimization strategy includes:
[0023] For different types of middleware, the following optimization strategies are adopted:
[0024] Based on optimization strategies, memory management, defragmentation, and connection optimization are performed on the Redis middleware.
[0025] Based on optimization strategies, we perform index optimization, parameter tuning, and table structure optimization on MySQL middleware.
[0026] Based on optimization strategies, queue backlog handling, memory level control, and dead letter queue handling are implemented in the MQ middleware.
[0027] Optionally, after optimizing the corresponding middleware based on the optimization strategy, the method further includes:
[0028] Continuously collect performance metrics within the preset verification period and compare key metrics before and after optimization;
[0029] If the optimization fails, a rollback will be performed based on the configuration snapshot before optimization.
[0030] If the optimization is successful but fails to achieve the expected results, the optimization strategy is regenerated based on the latest performance metrics.
[0031] Optionally, the method further includes:
[0032] The visualization platform displays the global topology map, performance trend chart, historical dashboard, and resource heat map in real time.
[0033] The global topology map is used to display the cluster status and dependencies of multiple middlewares; the performance trend map is used to display the year-on-year and month-on-month analysis results of multi-dimensional performance indicators; the historical dashboard is used to record the time, optimization strategy and optimization effect of each optimization; and the resource heat map is used to display the distribution of resource usage according to cluster or data center.
[0034] Secondly, this application provides a middleware performance optimization apparatus, the apparatus comprising:
[0035] The data acquisition module is used to collect performance metrics of each middleware during runtime through a distributed probe component.
[0036] The prediction module is used to predict the performance indicators through a prediction model and determine the performance inflection point when the prediction result reaches a dynamic threshold, wherein the dynamic threshold is calculated by learning the latest business data online.
[0037] The determination module is used to comprehensively determine the optimization strategy of each middleware at the corresponding performance inflection point based on the preset knowledge base rules and the middleware dependency graph. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of the change of the current middleware on the upstream and downstream middleware.
[0038] An optimization module is used to optimize the corresponding middleware based on the optimization strategy so that the middleware achieves optimal performance.
[0039] Thirdly, this application provides an electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.
[0040] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for executing the performance optimization method of the middleware described in any of the preceding claims of this application.
[0041] Compared with existing technologies, the technical solution provided in this application has the following advantages: It comprehensively collects performance indicators of multiple middleware through a distributed probe component, ensuring the real-time nature and integrity of the data. Then, it uses a predictive model combined with dynamic thresholds to analyze these performance indicators, identifying potential performance inflection points in advance. The dynamic thresholds are updated online by learning the latest business data to adapt to business fluctuations, thereby ensuring the accuracy of performance inflection point prediction. When the prediction results show that a certain middleware has reached its performance threshold, the system determines the strategy for that middleware based on rules provided by the expert knowledge base, and evaluates the impact of the middleware on upstream and downstream systems in conjunction with the middleware dependency graph, generating a global optimization strategy. Finally, the middleware is adjusted based on the determined optimization strategy to improve its performance stability. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0045] Figure 1 A flowchart of a middleware performance optimization method provided in this application embodiment;
[0046] Figure 2 A schematic diagram of the data acquisition process provided for embodiments of this application;
[0047] Figure 3 This is a schematic diagram of the system architecture provided for an embodiment of this application;
[0048] Figure 4 A flowchart illustrating the overall performance optimization process of middleware provided in this application embodiment;
[0049] Figure 5 A schematic diagram of the structure of a middleware performance optimization device provided in an embodiment of this application;
[0050] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0052] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0053] The following will describe in detail a middleware performance optimization method provided in this application embodiment, with reference to specific implementation methods. Figure 1 As shown, the specific steps are as follows:
[0054] Step 101: Collect performance metrics for each middleware during runtime using a distributed probe component;
[0055] Step 102: Predict performance indicators using a prediction model and determine the performance inflection point when the prediction results reach a dynamic threshold, where the dynamic threshold is calculated by learning the latest business data online.
[0056] Step 103: Based on the preset knowledge base rules and middleware dependency graph, determine the optimization strategy for each middleware at the corresponding performance inflection point. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of changes to the current middleware on upstream and downstream middleware.
[0057] Step 104: Optimize the corresponding middleware based on the optimization strategy to achieve optimal performance.
[0058] The following is an explanation of the terms used in the embodiments of this application, including the following content.
[0059] Distributed probe component: A lightweight monitoring tool deployed in a distributed system for real-time and accurate collection of various performance metrics during middleware runtime.
[0060] Predictive Model: Utilizing machine learning algorithms and built upon historical performance data, this model can predict the future performance trends of middleware metrics, helping to identify potential performance issues in advance.
[0061] Dynamic thresholds: Alarm thresholds that are dynamically adjusted based on business characteristics are used to determine whether performance indicators are in an abnormal state. Compared with fixed thresholds, they are more flexible and accurate.
[0062] Performance inflection point: The critical point at which the performance indicators of middleware change from a normal state to an abnormal state. When the prediction result reaches the dynamic threshold, it can be considered that the performance inflection point has been reached.
[0063] Knowledge base rules: Optimization rules derived from expert experience and historical optimization cases provide guidance and suggestions for middleware performance optimization.
[0064] Middleware dependency graph: A graphical representation of the dependencies between middleware, used to analyze the impact of performance changes in a single middleware on upstream and downstream middleware.
[0065] Optimization strategy: Specific optimization measures for middleware performance issues, including configuration adjustments and parameter optimization, aiming to achieve optimal middleware performance.
[0066] In step 101, middleware such as Redis, MySQL and MQ play a crucial role in modern backend systems. To ensure that these middleware can run efficiently and stably, it is especially important to monitor their running status in real time. Figure 2 The diagram illustrates the data acquisition process. First, lightweight probe programs are deployed on each middleware node. A data acquisition module, adapted to different middlewares (such as Redis, MySQL, and RabbitMQ), is developed based on the Prometheus Exporter architecture. This module collects 32 basic metrics, including CPU utilization, memory usage, connection count, and QPS (Queries Per Second), at 500-millisecond intervals via interfaces such as the Redis INFO command, MySQL Performance Schema, and MQ Admin API. The collected data is then serialized using Protobuf and transmitted to the data aggregation node via a Kafka message queue. This step not only achieves comprehensive coverage of middleware performance but also ensures the real-time nature and accuracy of the data.
[0067] In step 102, after obtaining a large number of middleware performance metrics, the next step is to analyze and predict them using advanced machine learning algorithms. A prediction model combining LSTM (Long Short-Term Memory) and random forest is employed. By learning from feature sequences of past historical periods (e.g., within 2 hours), it predicts the changing trends of key middleware performance metrics for a future time period (e.g., within 30 minutes), such as Redis memory usage or the number of MySQL query requests. This prediction model forms a middleware behavior feature library through training on historical performance data, enabling it to predict potential performance inflection points several time periods in advance. When the prediction results approach or reach dynamically set thresholds, the system automatically issues an early warning signal, alerting operations personnel to impending performance bottlenecks. This proactive performance monitoring mechanism significantly improves system response speed and stability, reducing the risk of service interruptions due to sudden failures.
[0068] To ensure the sensitivity and accuracy of the early warning system, the dynamic threshold used in this application is not fixed, but is automatically adjusted every fixed interval (e.g., 24 hours) based on the latest business data to adapt to the natural fluctuations in business volume. The dynamic threshold reduces the false alarm rate caused by business fluctuations, so that an alarm is only triggered when the actual performance indicators are truly close to or exceed the safe operating range.
[0069] In step 103, once a performance inflection point is detected, the system immediately activates the decision engine. It combines optimization rules from the expert knowledge base to generate a targeted single-point strategy for the current middleware. For example, if Redis memory usage exceeds 80% and fragmentation rate is higher than 40%, a memory defragmentation operation will be triggered. For a surge in slow MySQL queries, an index optimization algorithm will be automatically executed. Simultaneously, considering that a problem with a single middleware might have a cascading effect on its upstream and downstream components, a graph database (such as Neo4j) is used to construct a middleware dependency graph. This graph is used to assess the impact of a single middleware failure on its upstream and downstream components, generating association strategies for these components. Finally, the single-point strategy and association strategies are combined to determine the optimal optimization strategy. By comprehensively considering knowledge base rules and middleware dependency graph elements, an optimization strategy can be developed that addresses the current middleware performance issue without adversely affecting upstream and downstream middleware.
[0070] In step 104, based on the previously determined optimization strategy, specific adjustments are implemented to the corresponding middleware using automated tools. For example, in a Redis scenario, when memory usage exceeds 75%, the system automatically switches to the allkeys-lru eviction policy and adjusts parameters such as hash-max-ziplist-value; for MySQL, when an increase in slow queries is detected, the system automatically analyzes and creates the optimal index combination. All these optimization operations are performed non-intrusively through the middleware's API interfaces. These automated optimization methods help the middleware maintain optimal performance, improving the stability and efficiency of the entire system.
[0071] For example, in an e-commerce system, Redis and MySQL act as caching middleware, handling a large number of product information query tasks. During peak e-commerce promotions, a distributed probe component collects real-time performance metrics such as response time and throughput for Redis and MySQL. Based on historical data, the predictive model predicts that both Redis response time and MySQL throughput will continue to rise, potentially exceeding preset dynamic thresholds. Therefore, it is determined that both Redis and MySQL have reached performance inflection points. At this point, an optimization strategy is formulated based on knowledge base rules, including adjusting Redis's memory allocation strategy and optimizing MySQL's data structure. Simultaneously, considering the impact of Redis on downstream message queue (MQ) query performance, MQ is also optimized based on the middleware dependency graph. Finally, Redis, MySQL, and MQ are optimized according to the optimization strategy.
[0072] In this application, a distributed probe component is used to comprehensively collect performance metrics from multiple middleware components, ensuring the real-time nature and integrity of the data. Then, a predictive model combined with dynamic thresholds is used to analyze these performance metrics, identifying potential performance inflection points in advance. The dynamic thresholds are updated online by learning the latest business data to adapt to business fluctuations, thereby ensuring the accuracy of performance inflection point predictions. When the prediction results indicate that a middleware has reached its performance threshold, the system determines the strategy for that middleware based on rules provided by an expert knowledge base, and evaluates the impact of the middleware on upstream and downstream systems using a middleware dependency graph, generating a global optimization strategy. Finally, the middleware is adjusted based on the determined optimization strategy to improve its performance stability.
[0073] As an optional implementation, after collecting the performance metrics of each middleware runtime through the distributed probe component, the performance metrics need to be cleaned and feature extracted, including the following:
[0074] Step S11: Data Cleaning. After receiving middleware performance metric data from the Kafka message queue, the data aggregation node performs initial data cleaning using the Flink stream processing engine. The main purpose of this step is to remove invalid data points, such as outliers caused by network jitter. By filtering out this noisy data, it ensures that subsequent analysis is based on accurate and reliable information.
[0075] Step S12: Statistical Calculation. The system uses a sliding window algorithm (set to a 10-minute window size) to process the cleaned data and calculate key statistics within each window, including mean and variance. This time series analysis method helps to capture the trends and fluctuations of middleware performance metrics over a period of time.
[0076] Step S13: Feature Engineering. After calculating the basic statistical data, the feature engineering module uses PCA (Principal Component Analysis) dimensionality reduction technology to compress the original 32-dimensional features into 12 core features. In addition, it generates some derivative features with business significance, such as memory growth rate and slow query ratio. These features can more intuitively reflect the actual operating status of the middleware.
[0077] Step S14: Standardization. To ensure that features at different scales can be compared and analyzed under the same standard, all feature values will be mapped to the [0,1] interval using the range standardization method. This process improves data consistency.
[0078] This application improves the accuracy and efficiency of middleware performance data analysis through the aforementioned data preprocessing steps. First, data cleaning effectively removes noise data that may interfere with the analysis results, ensuring the quality of the input data. Second, the application of the sliding window algorithm enables the system to dynamically track the changing trends of performance indicators and promptly identify potential problems. Furthermore, through PCA dimensionality reduction and the construction of derived features, not only are data dimensionality reduced and computational complexity lowered, but the model's ability to capture key information is also enhanced. Finally, range standardization unifies the feature scale.
[0079] As an optional implementation, in step 102, the performance index is predicted using a prediction model, and the performance inflection point at which the prediction result reaches a dynamic threshold is determined, including the following:
[0080] Step S21: Input the performance metrics of the middleware into the prediction model to obtain the performance prediction values within a set future time period. The Long Short-Term Memory network in the prediction model is used to capture the time-series dependencies of the performance metrics, and the Random Forest in the prediction model is used to handle nonlinear feature interactions.
[0081] Step S22: Compare the performance prediction value with the corresponding dynamic threshold, wherein the dynamic threshold is calculated by learning the latest business data online;
[0082] Step S23: If the performance prediction value reaches the corresponding dynamic threshold, then the middleware is determined to have reached the performance inflection point.
[0083] In step S21, the system uses collected middleware performance metrics (such as Redis memory usage, MySQL QPS, etc.) as input data and feeds them into a pre-trained prediction model for analysis to obtain performance trend prediction results for a future period. This prediction model employs a two-layer LSTM network (each layer containing 128 neurons) to capture the time-dimensional dependencies and long-term trends of performance metrics; it also combines a random forest algorithm (e.g., 50 trees) to identify and process non-linear interactions between features. For example, the model's input is a feature sequence from the past two hours; the 12-dimensional core features and derived features, after feature engineering, constitute the input vector; and the output is the specific predicted value of the target performance metric for the next 30 minutes.
[0084] During model training, the Adam optimizer was used for parameter updates, and the loss function employed a weighted combination of MAE (Mean Absolute Error) and MSE (Mean Squared Error) to balance prediction accuracy and stability. To ensure good generalization ability, 5-fold cross-validation was used during training, and the model was thoroughly validated on historical fault datasets (such as CPU overload, memory leaks, and connection pool exhaustion). Ultimately, the model's prediction error rate was controlled below 8% under various abnormal scenarios, improving prediction accuracy.
[0085] By combining deep learning and machine learning in its modeling approach, the system can effectively identify time-series patterns and complex feature relationships of middleware performance changes, and predict potential performance problems in advance.
[0086] In step S22, after obtaining the future predicted values of performance indicators, the system compares them with the currently set dynamic thresholds. These dynamic thresholds are calculated in real-time based on the latest business data from the most recent time period, adapting to periodic fluctuations and sudden changes in business load. By dynamically adjusting the threshold mechanism, the system can more accurately identify performance inflection points, reducing the false alarm rate from the original 25% to below 5%, thus improving the robustness and practicality of anomaly detection.
[0087] The calculation of dynamic thresholds is a crucial step, used to determine whether middleware performance metrics are approaching or have reached a performance inflection point. This process employs different threshold calculation methods based on the statistical distribution characteristics of the prediction results to improve the accuracy and adaptability of anomaly detection.
[0088] If the predicted result is a normally distributed indicator (such as interface response time, QPS volatility, etc.), the 3σ principle is used to determine its alarm threshold. This involves calculating the mean and standard deviation of historical data over the past 24 hours, and setting the threshold as the mean ± 3 times the standard deviation as the upper and lower limits. Based on the statistical characteristics of normal distribution, this method can cover 99.7% of the normal data range, thus effectively identifying abnormal fluctuations.
[0089] If the predicted results are non-normally distributed indicators (such as the number of connections, the number of slow queries, the amount of message queues backlog, etc.), the threshold setting is based on the quantile method because the data distribution is irregular, skewed, or has a long tail. Typically, the 95th percentile is selected as the upper limit threshold, which filters out extreme outliers while retaining sensitivity to business fluctuations.
[0090] To ensure that the thresholds remain consistent with the current business status, the system automatically updates the thresholds of all indicators at fixed intervals (e.g., 24 hours) through an online learning mechanism. This dynamic adjustment method solves the problem that traditional fixed thresholds cannot adapt to periodic changes in business or sudden traffic surges, improving the accuracy and stability of the early warning system.
[0091] In step S23, when the predicted value of a key performance indicator of a middleware reaches or exceeds its dynamic threshold, the system determines that the middleware is about to enter a performance inflection point state. For example, if the prediction finds that Redis memory will exceed the threshold in 15 minutes, or that MySQL's QPS will continue to exceed its maximum capacity, a performance inflection point event is triggered. At this time, the system issues an alarm and initiates the subsequent optimization decision-making process to generate optimization strategies for the current middleware and its upstream and downstream components.
[0092] Alarms are categorized into four levels: Emergency Alarms, Warning Alarms, Alert Alarms, and Notification Alarms. Emergency alarms (such as Redis OOM) are notified via both SMS and phone; Warning alarms (such as a surge in MySQL slow queries) are pushed via WeChat; Alert alarms are displayed in the management backend; Notification alarms are not displayed and are only used for logging and subsequent analysis. Alarm content includes fault location information (such as specific middleware nodes and abnormal indicators) and suggested solutions (from an expert knowledge base), with an average alarm location time of less than 3 minutes.
[0093] By identifying performance inflection points in advance and responding quickly, the system can take intervention measures before performance bottlenecks actually occur, preventing service interruptions, request delays, and other problems, thereby improving the system's stability and fault tolerance.
[0094] This application achieves proactive identification of middleware performance inflection points by constructing a high-precision prediction model (LSTM + Random Forest) and combining it with a dynamic threshold mechanism. Compared with traditional monitoring methods based on static thresholds, this solution not only has higher prediction accuracy but also adapts to business fluctuations, reducing false alarm rates. This closed-loop mechanism of prediction, judgment, and response enables the system to intervene in advance before performance problems manifest, greatly improving the stability and resource utilization of middleware operation, thereby ensuring the high availability and quality of service of the overall system.
[0095] As an optional implementation, in step 203, the optimization strategy for each middleware at the corresponding performance inflection point is determined based on preset knowledge base rules and middleware dependency graphs, including the following:
[0096] Step S31: Generate at least one single-point strategy for the middleware at the corresponding performance inflection point based on preset knowledge base rules;
[0097] Step S32: Analyze the impact of single-point strategies on upstream and downstream middleware based on the middleware dependency graph, and generate at least one linkage strategy, wherein the linkage strategy is used to optimize upstream and downstream middleware;
[0098] Step S33: Combine at least one single-point strategy and at least one linkage strategy into a candidate strategy set;
[0099] Step S34: Evaluate the expected return of each candidate strategy in the candidate strategy set through reinforcement learning, and select the candidate strategy with the highest return as the optimization strategy. The expected return is used to indicate the extent of improvement in response time or the degree of resource consumption of the candidate strategy.
[0100] In step S31, the decision engine is activated when the predictive model detects that a specific middleware has reached or is about to reach a performance inflection point (e.g., Redis memory usage exceeds 80% and fragmentation rate is greater than 40%, or MySQL slow queries exceed 50 per minute). The system first accesses a knowledge base consisting of 200+ optimization rules, stored in JSON format and editable and maintainable via a web management interface. Each rule in the knowledge base is carefully designed to provide corresponding solutions for different types of performance problems. For example, when Redis memory usage > 80% and fragmentation rate > 40%, memory defragmentation is triggered; when MySQL slow queries exceed 50 per minute, index analysis is initiated. Furthermore, the system can automatically recommend rules for similar scenarios based on historical optimization cases, forming a self-evolving knowledge system.
[0101] This application utilizes a knowledge base built from expert experience and historical data to quickly identify performance bottlenecks and propose effective optimization suggestions, thereby reducing manual intervention while improving processing efficiency.
[0102] In step S32, to ensure that optimization measures do not cause new problems, the system uses a middleware dependency graph built with the Neo4j graph database to evaluate the impact of the proposed single-point strategy on its upstream and downstream components. For example, in the case of message queue backlog, the system automatically checks the connection pool status of the downstream MySQL and increases the number of connections in advance to cope with potential pressure. At the same time, the system uses the Apriori algorithm to mine performance correlation rules between middleware. For example, when the Redis slow query rate is >10%, the MySQL response time has an 80% probability of increasing, realizing a leap from single-point monitoring to global linkage optimization, which significantly reduces the risk of downstream system cascading failure.
[0103] By conducting in-depth analysis of middleware dependencies and applying cross-component performance correlation rules, the overall coordination and forward-looking nature of the optimization strategy are ensured, effectively preventing chain reactions caused by local optimization.
[0104] In step S33, after identifying possible single-point strategies and their impact on upstream and downstream components, the system combines these strategies to form multiple candidate strategy sets. Each set contains a complete set of optimization action plans, covering all operational instructions from the target middleware to its upstream and downstream components, from solving direct problems to preventing potential risks. This process not only considers technical feasibility but also takes into account resource allocation and scheduling in the actual operating environment.
[0105] Strategy combinations offer a variety of optimization path options, enhancing the system's adaptability and flexibility, and helping to find the solution best suited to the current situation.
[0106] In step S34, to select the optimal solution from numerous candidate policies, the system applies a reinforcement learning algorithm (DQN, Deep Q-Network) to calculate and evaluate the expected return of each candidate policy. The expected return reflects the policy's potential to improve response time and reduce resource consumption. During training, the system continuously adjusts and optimizes the scoring mechanism using historical performance data and simulated environments. Ultimately, the system can complete policy selection and instruction generation in less than 200ms, ensuring real-time performance and efficiency.
[0107] By using intelligent algorithms to accurately select the best strategy, the optimization effect is maximized, and the resource utilization is optimized, which greatly improves the system's autonomous decision-making ability and operation and maintenance efficiency.
[0108] As an optional implementation, in step 104, different optimization strategies are adopted for different types of middleware.
[0109] 1. Redis middleware.
[0110] 1.1 Memory management strategy optimization.
[0111] When Redis memory usage exceeds 75%, the system automatically switches the eviction policy to allkeys-lru, prioritizing the retention of recently accessed data. Simultaneously, parameters such as hash-max-ziplist-value are dynamically adjusted based on the type and size of the currently stored data to improve memory utilization efficiency.
[0112] 1.2 Memory defragmentation mechanism.
[0113] To address potential memory fragmentation issues that may arise after prolonged operation, the system asynchronously executes the ACTIVATE DEFrag command during idle periods when CPU load is below 30% to defragment memory. This operation does not block the main process, ensuring service continuity.
[0114] 1.3 Connection and network parameter optimization.
[0115] In scenarios with sudden surges in connection requests (such as flash sales), the system dynamically adjusts the tcp-backlog parameter based on the trend of connection count changes to increase the capacity of the connection queue and prevent connection loss or denial of service.
[0116] 2. MySQL middleware.
[0117] 2.1 Slow query index recommendations and automatic creation.
[0118] The pt-query-digest tool was used to analyze slow query logs and identify high-frequency, inefficient queries. A greedy algorithm was combined to evaluate field selectivity, recommend the optimal index combination, and automatically create indexes using the ALTER TABLE statement to improve query efficiency.
[0119] 2.2 Dynamic optimization of database parameters.
[0120] Based on the current system workload characteristics (such as read / write ratio, number of concurrent connections, etc.), more than 30 core parameters, including innodb_buffer_pool_size, sync_binlog, and max_connections, are dynamically adjusted to adapt to different business scenarios and improve resource utilization.
[0121] 2.3 Table structure optimization and table partitioning suggestions.
[0122] When a full table scan rate exceeding 50% is detected for a table, it indicates that the index may be ineffective or poorly designed. In this case, the system generates table partitioning suggestions and automatically generates data migration scripts for operation and maintenance personnel to review and execute, thus avoiding performance degradation caused by large tables.
[0123] 3. MQ middleware.
[0124] 3.1 Intelligent scaling of consumer parallelism.
[0125] When the system detects that the number of messages piling up in a queue exceeds a set threshold (e.g., 10,000), it automatically increases the number of consumer instances to improve consumption speed. Simultaneously, by setting an appropriate `prefetch_count` through the `channel.basic_qos` interface provided by RabbitMQ, load balancing is achieved to prevent overload of any single consumer.
[0126] 3.2 Broker memory level monitoring and optimization.
[0127] When Broker's memory usage exceeds 70%, the system triggers a lazy paging mechanism to reduce memory pressure; at the same time, it dynamically adjusts parameters such as page_cache_size to improve memory reclamation efficiency and prevent service interruptions due to insufficient memory.
[0128] 3.3 Cohort health status assessment and early warning.
[0129] The system continuously monitors indicators such as production rate, consumption rate, and backlog trend of each queue, and establishes a queue health scoring model. When the score falls below the threshold, an early warning is issued and a preset emergency handling procedure is initiated, such as restarting some consumers or temporarily expanding capacity.
[0130] As an optional implementation, after optimizing the corresponding middleware based on the optimization strategy, the method further includes:
[0131] Step S51: Continuously collect performance indicators within the preset verification period and compare the key indicators before and after optimization;
[0132] Step S52: If the optimization fails by comparison, roll back based on the configuration snapshot before optimization;
[0133] Step S53: If the optimization is successful but the expected effect is not achieved through comparison, the optimization strategy is regenerated based on the latest performance indicators.
[0134] In step S51, after the optimization operation is completed, the system immediately starts a preset verification period (e.g., 10 minutes). During this period, key performance indicators of the target middleware are continuously collected, including but not limited to response time, QPS (queries per second), slow queries, connections, and memory usage. Simultaneously, the system compares and analyzes this optimized data with the historical baseline before optimization. Furthermore, the system uses statistical methods such as t-tests to determine whether the performance changes are significant, avoiding misjudgments of the optimization effect due to random fluctuations.
[0135] In step S52, to ensure service stability, a configuration snapshot is generated before each optimization operation, recording all key parameter configurations of the current middleware. Examples include Redis's maxmemory, eviction policy, and hash compression threshold; MySQL's innodb_buffer_pool_size and log flushing policy; and MQ's consumer concurrency, prefetch quantity, and pagination cache size.
[0136] If, after the verification period ends, it is found that: key performance indicators have not improved significantly (e.g., response time decreases by <5%); or abnormal fluctuations occur (e.g., CPU spikes, connection interruptions); then the system determines that the optimization has failed and automatically triggers the rollback mechanism to restore the configuration state before optimization.
[0137] The rollback process is managed through a transaction mechanism to ensure atomicity and consistency. For parameters that can be dynamically modified via API (such as Redis's CONFIG SET and MySQL's SET GLOBAL), the interface is directly called remotely to restore the changes. For changes that require a restart to take effect (such as MySQL data directory migration), a rolling upgrade strategy is adopted. The new configuration is first deployed on a standby node and verified before switching traffic to ensure uninterrupted service.
[0138] In step S53, even if the optimization operation brings performance improvement, if the improvement is less than the expected target (e.g., response time only improves by 8%, while the target is 15%), the system will consider the optimization part effective, but further adjustments are still needed. At this point, the system will regenerate a better optimization strategy based on the latest performance metrics, combined with historical knowledge base rules and reinforcement learning strategies. This process achieves continuous optimization and strategy iteration, avoiding the limitation of ending after a single optimization, and improving the overall system's adaptability and self-healing capabilities.
[0139] This invention achieves closed-loop control of middleware performance governance by constructing a complete optimization verification and feedback mechanism, with the following beneficial effects: 1. Improved optimization success rate: Real-time data collection and comparative verification enable rapid determination of optimization effectiveness, avoiding blind changes. 2. Enhanced system robustness: Transaction mechanisms and configuration snapshots ensure safe rollback of any changes, greatly reducing risks. 3. Support for continuous optimization evolution: Even if the initial optimization fails to meet expectations, iterative iterations based on the latest data can continue, gradually approaching the optimal solution. 4. Increased automation: From optimization execution to effect evaluation to policy updates, the entire process requires no manual intervention, truly achieving intelligent operation and maintenance. 5. Reduced invalid configuration legacy: For optimization actions that fail to meet standards, timely rollback or replacement prevents configuration redundancy from affecting subsequent decisions.
[0140] As an optional implementation, the method further includes: displaying a global topology map, performance trend map, historical dashboard, and resource heat map in real time on a visualization platform; wherein, the global topology map is used to display the cluster status and dependencies of multiple middlewares; the performance trend map is used to display the year-on-year and month-on-month analysis results of multi-dimensional performance indicators; the historical dashboard is used to record the time, optimization strategy, and optimization effect of each optimization; and the resource heat map is used to display the distribution of resource usage according to cluster or data center.
[0141] The system builds a unified visual operation and maintenance platform, which centrally displays key information such as global topology map, performance trend map, historical dashboard and resource heat map, providing operation and maintenance personnel with real-time monitoring capabilities from a global perspective.
[0142] A global topology diagram clearly and graphically displays the call chains and dependencies between multiple middleware components. Each node represents a middleware instance or cluster, and the node color and size dynamically change based on the current health status (such as CPU utilization, memory usage, and response time), intuitively reflecting the system's operational status. The global topology diagram helps operations and maintenance personnel quickly locate fault points and their impact range, improving their understanding and control over complex system architectures.
[0143] The performance trend chart can simultaneously display the change curves of more than 10 core indicators such as QPS, response time, number of connections, number of slow queries, and message backlog. Users can choose year-on-year or month-on-month analysis. Through multi-dimensional trend comparison, it can help determine whether performance fluctuations are caused by business growth, configuration changes, or abnormal behavior, improving the efficiency of problem diagnosis.
[0144] The historical dashboard is used to record the complete lifecycle information of all automated optimization operations, including the optimization trigger time, the basis for identifying inflection points, the specific strategies executed, changes in key indicators before and after optimization, and status indicators such as whether it was successful, whether it was rolled back, and whether it triggered a secondary optimization.
[0145] The resource heatmap module aggregates and displays resource usage across different dimensions (such as cluster, data center, and availability zone). Examples include: memory usage heatmaps for each Redis cluster, CPU load distribution maps for MySQL nodes, and message throughput heatmaps for the MQ Broker. This module helps operations personnel quickly identify resource bottlenecks, aiding in capacity planning, scheduling decisions, and cross-data center disaster recovery design.
[0146] As an optional implementation method, the method also includes: achieving regular system self-optimization and upgrades by collecting and analyzing historical optimization data (including but not limited to strategies, effects and their impact on business), with specific measures as follows.
[0147] Periodic updates to the predictive model: The predictive model is retrained quarterly to incorporate the latest business scenario data. This step ensures that the model can adapt to constantly changing business needs and technological environments, thereby improving predictive accuracy and response efficiency.
[0148] Performance evaluation and adjustment of optimization strategies: Success rate analysis of implemented optimization strategies is conducted to identify and eliminate inefficient strategies that are underperforming or no longer applicable. This process not only helps to eliminate ineffective operations but also provides direction for the development of new strategies, promoting the dynamic evolution of the strategy library.
[0149] User feedback-based interaction flow optimization: Closely monitor user feedback and make necessary adjustments and optimizations to the system's interaction flow based on problems encountered and suggestions made during use. This can effectively improve the user experience and increase the system's usability and user satisfaction.
[0150] Version iteration is managed using GitFlow workflow: GitFlow branch management is used to standardize the development process for version iteration. This approach facilitates team collaboration, ensures code quality, and also makes it easier to manage and track project progress.
[0151] Docker containerization ensures service continuity: By leveraging Docker technology to containerize applications, systems can be updated and maintained without affecting service. This non-disruptive service model is crucial for maintaining business continuity and stability.
[0152] Figure 3 The system architecture diagram shows that the system is mainly divided into four layers: data acquisition layer, AI processing layer, visualization layer and automation control layer. The functions and processes of each layer will be described in detail below.
[0153] In the data acquisition layer, probe components are deployed on various middleware (MySQL middleware, MQ middleware, Redis middleware), responsible for collecting the performance metrics of these middleware in real time and transmitting the data to the data aggregation node.
[0154] In the AI processing layer, the feature engineering module cleans, transforms, and extracts features from the data acquisition layer. The prediction model cluster uses the feature vectors generated by the feature engineering module to perform model predictions. The policy generator generates specific optimization strategies based on the prediction results of the prediction models, the expert knowledge base, and the middleware dependency graph.
[0155] In the visualization layer, the global performance dashboard displays various performance indicators and system topology information provided by the data acquisition layer and AI processing layer in real time. The alarm center automatically triggers alarms based on the prediction results of the AI processing layer and dynamic threshold settings, and notifies relevant personnel via email, SMS and other means.
[0156] In the automation control layer, the execution engine makes configuration changes to the middleware through API calls or other means based on the received optimization strategy. The configuration management center records the operation logs and effect evaluations of each configuration change for subsequent backtracking and analysis.
[0157] The entire system acquires real-time data through a data acquisition layer, undergoes intelligent analysis and decision-making by an AI processing layer, generates optimization strategies, implements them through an automated control layer, and finally displays system status and alarm information through a visualization layer, forming a closed-loop automated operation and maintenance system. This architecture not only improves the system's operational efficiency and stability but also enables it to continuously learn and adapt to new business scenarios, achieving continuous optimization.
[0158] This application provides an overall flowchart for middleware performance optimization, such as... Figure 4 As shown, the steps include the following.
[0159] Step 401: Performance Metrics Collection and Performance Trend Prediction. Analyze the real-time performance metrics of the middleware using the established prediction model to predict performance trends over a future period (e.g., 30 minutes).
[0160] Step 402: Compare and determine the performance inflection point. Compare the predicted performance metrics with dynamic thresholds. If the threshold is reached or exceeded, it is identified as a performance inflection point.
[0161] Step 403: Generate a candidate strategy set. Based on the detected performance inflection point, and combining expert knowledge base rules with the dependency graph between middleware, a candidate strategy set is generated.
[0162] Step 404: Evaluate and select the optimal strategy. The expected benefit of each strategy in the candidate strategy set is evaluated using a reinforcement learning algorithm. The strategy that maximizes the improvement in response time or minimizes resource consumption is selected as the final execution plan.
[0163] Step 405: Perform optimization operations. Following the selected optimization strategy, use automated tools to configure and adjust the relevant middleware. All changes are managed through transaction mechanisms to ensure safe rollback in case of exceptions.
[0164] Step 406: Effect Verification. Initiate a short-term (e.g., 10-minute) effect verification cycle, continuously collect performance metrics and compare them with the data before optimization to assess whether the optimization effect has achieved the expected goals.
[0165] Step 407: Secondary optimization (if necessary). If the initial optimization fails to significantly improve performance or does not achieve the expected goals, the optimization strategy is regenerated based on the latest performance data, and the optimization operation is repeated until the requirements are met.
[0166] Step 408: Visualization and Monitoring. A visualization platform is used to display real-time information such as the global topology map, performance trend charts, historical dashboards, and resource heatmaps, enabling operations and maintenance personnel to fully understand the system status and take timely measures. Step 408 is integral to the entire process from steps 401 to 407.
[0167] The core idea of this application is to construct an automated middleware optimization and performance monitoring system that integrates AI algorithms, breaking through the traditional operation and maintenance model that relies on manual experience, and realizing intelligent perception and adaptive tuning of middleware such as Redis, MySQL, and MQ. Its technical principle follows a closed-loop control process of data perception—intelligent analysis—automatic decision-making—closed-loop execution: First, a distributed probe component collects more than 20 core performance indicators, including CPU utilization, memory usage, IO throughput, and connection count, in real time at each middleware node, forming a highly timely data input stream. Then, the collected data is input into a prediction model based on a fusion of LSTM and random forest. This model, through learning from historical performance data, can identify the behavioral characteristics of the middleware and predict potential performance inflection points (such as sudden increases in Redis memory usage, rising MySQL response latency, etc.) 30 minutes in advance. When the system detects performance anomalies or predicted values approaching a preset risk range, the decision engine will match optimization rules from the expert knowledge base with the current middleware status and upstream and downstream middleware to generate targeted optimization strategies. For example, it might automatically trigger index rebuilding for slow MySQL queries or dynamically adjust key parameters such as `innodb_buffer_pool_size` to improve caching efficiency. Finally, the execution module completes the configuration update operation through the middleware's native API, ensuring a non-intrusive and low-latency change process. Simultaneously, the system continuously monitors the optimized performance, verifies the optimization effect through a feedback mechanism, and automatically triggers secondary optimization if expectations are not met, thus forming a complete closed-loop iterative optimization chain.
[0168] This application can achieve the following beneficial effects:
[0169] 1. Reduce maintenance costs: Automation solutions reduce middleware maintenance manpower costs by about 70%, and shorten the average fault recovery time from 30 minutes to less than 5 minutes, saving enterprises maintenance costs.
[0170] 2. Enhance product competitiveness: After the middleware performance and stability are improved, the QPS of core business systems can be increased by about 50% and the response time can be reduced by 50%, helping enterprises maintain a competitive advantage in high-concurrency scenarios.
[0171] 3. Building a technological barrier: This automated optimization system forms a reusable technology platform, supporting the rapid deployment of high-performance middleware environments for the company's new businesses. At the same time, the related technologies can be packaged into SaaS services for external output.
[0172] Based on the same technical concept, this application provides a middleware performance optimization device, such as... Figure 5 As shown, the device includes:
[0173] The acquisition module 501 is used to collect performance metrics of each middleware during runtime through a distributed probe component;
[0174] The prediction module 502 is used to predict performance indicators through a prediction model and determine the performance inflection point when the prediction result reaches a dynamic threshold, wherein the dynamic threshold is calculated by learning the latest business data online.
[0175] The determination module 503 is used to comprehensively determine the optimization strategy of each middleware at the corresponding performance inflection point based on the preset knowledge base rules and the middleware dependency graph. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of the change of the current middleware on the upstream and downstream middleware.
[0176] The optimization module 504 is used to optimize the corresponding middleware based on the optimization strategy so that the middleware achieves optimal performance.
[0177] Optionally, the prediction module 502 is used for:
[0178] The performance metrics of the middleware are input into the prediction model to obtain the performance prediction values within a set future time period. The long short-term memory network in the prediction model is used to capture the time series dependencies of the performance metrics, and the random forest in the prediction model is used to handle nonlinear feature interactions.
[0179] Compare the predicted performance values with the corresponding dynamic thresholds;
[0180] If the predicted performance value reaches the corresponding dynamic threshold, then the middleware is determined to have reached a performance inflection point.
[0181] Optionally, the determining module 503 is used for:
[0182] Generate at least one single-point strategy for the middleware at the corresponding performance inflection point based on preset knowledge base rules;
[0183] Based on the middleware dependency graph analysis, the impact of single-point strategies on upstream and downstream middleware is analyzed, and at least one linkage strategy is generated, wherein the linkage strategy is used to optimize upstream and downstream middleware.
[0184] Combine at least one single-point strategy and at least one linkage strategy into a candidate strategy set;
[0185] The expected return of each candidate policy in the candidate policy set is evaluated by reinforcement learning, and the candidate policy with the highest return is selected as the optimization policy. The expected return is used to indicate the extent of improvement in response time or the degree of resource consumption of the candidate policy.
[0186] Optionally, the device is also used for:
[0187] If the prediction result is a normally distributed index, the 3σ principle is used to determine the alarm threshold of the prediction result;
[0188] If the prediction result is a non-normal indicator, the quantile method is used to determine the alarm threshold of the prediction result;
[0189] Through online learning, alarm thresholds are recalculated at fixed time intervals based on the latest business data.
[0190] Optionally, the optimization module 504 is used for:
[0191] For different types of middleware, the following optimization strategies are adopted:
[0192] Based on optimization strategies, memory management, defragmentation, and connection optimization are performed on the Redis middleware.
[0193] Based on optimization strategies, we perform index optimization, parameter tuning, and table structure optimization on MySQL middleware.
[0194] Based on optimization strategies, queue backlog handling, memory level control, and dead letter queue handling are implemented in the MQ middleware.
[0195] Optionally, the device is also used for:
[0196] Continuously collect performance metrics within the preset verification period and compare key metrics before and after optimization;
[0197] If the optimization fails, a rollback will be performed based on the configuration snapshot before optimization.
[0198] If the optimization is successful but fails to achieve the expected results, the optimization strategy is regenerated based on the latest performance metrics.
[0199] Optionally, the device is also used for:
[0200] The visualization platform displays the global topology map, performance trend chart, historical dashboard, and resource heat map in real time.
[0201] Among them, the global topology map is used to display the cluster status and dependencies of multiple middlewares; the performance trend map is used to display the year-on-year and month-on-month analysis results of multi-dimensional performance indicators; the historical dashboard is used to record the time, optimization strategy and optimization effect of each optimization; and the resource heat map is used to display the distribution of resource usage according to cluster or data center.
[0202] like Figure 6 As shown, this application provides an electronic device including a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0203] Memory 603 is used to store computer programs.
[0204] In one embodiment of this application, when the processor 601 executes the program stored in the memory 603, it implements the middleware performance optimization method provided in any of the foregoing method embodiments.
[0205] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the middleware performance optimization method provided in any of the foregoing method embodiments.
[0206] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0208] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0209] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A middleware performance optimization method, characterized in that, The method includes: The performance metrics of each middleware runtime are collected through a distributed probe component. The performance indicators are predicted using a predictive model, and the performance inflection point when the prediction result reaches a dynamic threshold is determined, wherein the dynamic threshold is calculated by learning the latest business data online. Based on the preset knowledge base rules and middleware dependency graph, the optimization strategy for each middleware at the corresponding performance inflection point is determined. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of changes to the current middleware on upstream and downstream middleware. The middleware is optimized based on the optimization strategy to achieve optimal performance.
2. The method according to claim 1, characterized in that, The performance indicators are predicted using a predictive model, and the performance inflection point at which the prediction result reaches a dynamic threshold is determined, including: The performance metrics of the middleware are input into the prediction model to obtain the performance prediction values within a future set time period. The Long Short-Term Memory network in the prediction model is used to capture the time-series dependencies of the performance metrics, and the Random Forest in the prediction model is used to handle nonlinear feature interactions. The predicted performance value is compared with the corresponding dynamic threshold; If the predicted performance value reaches the corresponding dynamic threshold, then the middleware is determined to have reached a performance inflection point.
3. The method according to claim 1, characterized in that, Based on preset knowledge base rules and middleware dependency graphs, the optimization strategies for each middleware at its corresponding performance inflection point are determined, including: Generate at least one single-point strategy for the middleware at the corresponding performance inflection point based on preset knowledge base rules; Based on the middleware dependency graph, the impact of the single-point strategy on upstream and downstream middleware is analyzed, and at least one linkage strategy is generated, wherein the linkage strategy is used to optimize the upstream and downstream middleware. The at least one single-point strategy and the at least one linkage strategy are combined into a candidate strategy set; The expected return of each candidate strategy in the candidate strategy set is evaluated by reinforcement learning, and the candidate strategy with the highest return is selected as the optimization strategy, wherein the expected return is used to indicate the extent of improvement in response time or the degree of resource consumption of the candidate strategy.
4. The method according to claim 1, characterized in that, Determining the dynamic threshold includes: If the prediction result is a normally distributed index, then the 3σ principle is used to determine the alarm threshold of the prediction result; If the prediction result is a non-normal indicator, the alarm threshold of the prediction result is determined by the quantile method; Through online learning, the alarm threshold is recalculated at fixed time intervals based on the latest business data.
5. The method according to claim 1, characterized in that, Optimizing the corresponding middleware based on the aforementioned optimization strategy includes: For different types of middleware, the following optimization strategies are adopted: Based on optimization strategies, memory management, defragmentation, and connection optimization are performed on the Redis middleware. Based on optimization strategies, we perform index optimization, parameter tuning, and table structure optimization on MySQL middleware. Based on optimization strategies, queue backlog handling, memory level control, and dead letter queue handling are implemented in the MQ middleware.
6. The method according to claim 1, characterized in that, After optimizing the corresponding middleware based on the optimization strategy, the method further includes: Continuously collect performance metrics within the preset verification period and compare key metrics before and after optimization; If the optimization fails, a rollback will be performed based on the configuration snapshot before optimization. If the optimization is successful but fails to achieve the expected results, the optimization strategy is regenerated based on the latest performance metrics.
7. The method according to claim 1, characterized in that, The method further includes: The visualization platform displays the global topology map, performance trend chart, historical dashboard, and resource heat map in real time. The global topology map is used to display the cluster status and dependencies of multiple middlewares; the performance trend map is used to display the year-on-year and month-on-month analysis results of multi-dimensional performance indicators; the historical dashboard is used to record the time, optimization strategy and optimization effect of each optimization; and the resource heat map is used to display the distribution of resource usage according to cluster or data center.
8. A middleware performance optimization device, characterized in that, The device includes: The data acquisition module is used to collect performance metrics of each middleware during runtime through a distributed probe component. The prediction module is used to predict the performance indicators through a prediction model and determine the performance inflection point when the prediction result reaches a dynamic threshold, wherein the dynamic threshold is calculated by learning the latest business data online. The determination module is used to comprehensively determine the optimization strategy of each middleware at the corresponding performance inflection point based on the preset knowledge base rules and the middleware dependency graph. The knowledge base rules are used to determine the strategy corresponding to the current middleware, and the middleware dependency graph is used to determine the impact of the change of the current middleware on the upstream and downstream middleware. An optimization module is used to optimize the corresponding middleware based on the optimization strategy so that the middleware achieves optimal performance.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.