An Adaptive Tuning Method and System for Cloud Distributed Cache Databases
By constructing a network causal model in a cross-multi-subject cloud environment, and automatically determining the tuning strategy of the distributed cache database, the problem of low manual adjustment efficiency in the existing technology is solved, and efficient business indicator optimization is achieved.
Patent Information
- Application Number
- CN202411047939.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-07-31
AI Technical Summary
In a multi-main cloud environment, it is difficult to efficiently optimize business indicators for system tuning of distributed cache databases, especially due to the large number of configuration items and low manual adjustment efficiency, resulting in low business indicator optimization efficiency.
Causal AI technology is used to build a network causal model, and by analyzing the causal relationship between configuration items and business indicators, automatically determine the tuning strategy and adjust the configuration parameters to achieve adaptive tuning.
It improves the optimization efficiency of business indicators, reduces manual intervention time, and improves the adaptive tuning capability of the system.
Smart Images

Figure CN119066095B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of cloud computing technology, and in particular, to a method and system for adaptive tuning of an inter-cloud distributed cache database. Background Art
[0002] A cache database deployed distributively in a multi-cloud environment across multiple entities needs to ensure technical characteristics of high reliability, high availability, high throughput, and low latency. However, due to the fact that the usage scenarios often have diverse business requests, high-concurrency access, require real-time data return, and high complexity of the deployment architecture, etc., it is difficult to optimize the system around the business goals. With the increase in the cluster scale and the retrieval concurrency of the business system, the difficulty of manually optimizing the parameters of the core business-side business metrics increases rapidly.
[0003] The main means to solve this problem currently is that a database administrator with rich professional experience needs to configure some configuration items of the cluster system manually according to the business load requirements and runtime characteristics, and then continuously try and error to optimize the business metrics on the core business side.
[0004] However, since the number of configuration items of the cluster system is large, it takes a long time to determine which configuration item to adjust and its adjustment parameters manually. Therefore, the efficiency of optimizing the business metrics through the existing technology is low. Summary of the Invention
[0005] Embodiments of the present disclosure provide a method and system for adaptive tuning of an inter-cloud distributed cache database, which can improve the optimization efficiency and quality of business metrics.
[0006] In a first aspect, embodiments of the present disclosure provide a method for adaptive tuning of an inter-cloud distributed cache database, including:
[0007] In response to a tuning request for the inter-cloud distributed cache database, obtaining an expected value of at least one business metric carried in the tuning request;
[0008] According to the expected value of the at least one business metric and the network causal model corresponding to the inter-cloud distributed cache database, determining a tuning strategy corresponding to the inter-cloud distributed cache database, where the tuning strategy includes at least one target configuration item and an adjustment parameter for each target configuration item, and the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the business metric based on the causal relationship from the configuration item to the business metric variable constructed by analyzing multiple types of data;
[0009] According to the tuning strategy corresponding to the inter-cloud distributed cache database, adjusting the configuration parameters of each target configuration item of the inter-cloud distributed cache database to their respective corresponding adjustment parameters.
[0010] Second aspect, embodiments of the present disclosure provide a cloud-based distributed cache database adaptive tuning system, including:
[0011] An acquisition unit, configured to obtain an expected value of at least one service metric carried in the tuning request in response to a tuning request for the cloud-based distributed cache database;
[0012] A determination unit, configured to determine a tuning strategy corresponding to the cloud-based distributed cache database according to the expected value of the at least one service metric and the network causal model corresponding to the cloud-based distributed cache database, where the tuning strategy includes at least one target configuration item and an adjustment parameter for each target configuration item, and the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the service metric based on the causal relationship from the configuration item to the service metric variable constructed by analyzing multiple types of data;
[0013] An adjustment unit, configured to adjust the configuration parameters of each target configuration item of the cloud-based distributed cache database to their respective corresponding adjustment parameters according to the tuning strategy corresponding to the cloud-based distributed cache database.
[0014] Third aspect, embodiments of the present disclosure provide an electronic device, including: a processor and a memory;
[0015] The memory stores computer execution instructions;
[0016] The processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the cloud-based distributed cache database adaptive tuning method as described in the first aspect and various possible designs of the first aspect above.
[0017] Fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, in which computer execution instructions are stored, and when the processor executes the computer execution instructions, the cloud-based distributed cache database adaptive tuning method as described in the first aspect and various possible designs of the first aspect above is implemented.
[0018] Fifth aspect, embodiments of the present disclosure provide a computer program product, including a computer program, and when the computer program is executed by the processor, the cloud-based distributed cache database adaptive tuning method as described in the first aspect and various possible designs of the first aspect above is implemented.
[0019] The cloud distributed cache database adaptive tuning method and system provided in this embodiment, the method includes: in response to a tuning request for the cloud distributed cache database, obtaining the expected values of at least one service metric carried in the tuning request; according to the expected values of at least one service metric and the network causal model corresponding to the cloud distributed cache database, determining the tuning strategy corresponding to the cloud distributed cache database, the tuning strategy includes at least one target configuration item and the adjustment parameter of each target configuration item, wherein the network causal model is used to construct the causal relationship from the configuration item to the service metric variable based on various types of data analysis, and predict and analyze the configuration parameters of the configuration item for the expected value of the service metric; according to the tuning strategy corresponding to the cloud distributed cache database, adjusting the configuration parameters of each target configuration item of the cloud distributed cache database to their respective corresponding adjustment parameters. In this technical solution, since the network causal model is used to construct the causal relationship from the configuration item to the service metric variable based on various types of data analysis, and predict and analyze the configuration parameters of the configuration item for the expected value of the service metric, the network causal model can determine the adjustment parameters of at least one target configuration item corresponding to the expected value of the service metric, realize counterfactual reasoning to generate the tuning strategy of the cloud distributed cache database, realize the adaptive tuning of the cloud distributed cache database, and therefore improve the optimization efficiency of the service metric. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic diagram of the application scenario of the cloud distributed cache database adaptive tuning method provided in the embodiment of the present disclosure;
[0022] Figure 2 It is the flow of the cloud distributed cache database adaptive tuning method provided in the embodiment of the present disclosure Figure 1 ;
[0023] Figure 3 It is the schematic of the cloud distributed cache database adaptive tuning method provided in the embodiment of the present disclosure Figure 1 ;
[0024] Figure 4 It is the flow of the method for generating the network causal model provided in the embodiment of the present disclosure Figure 1 ;
[0025] Figure 5 It is the update schematic of the network causal model provided in the embodiment of the present disclosureFigure 1 ;
[0026] Figure 6 It is a schematic structural diagram of the cloud distributed cache database adaptive tuning system provided by the embodiments of the present disclosure;
[0027] Figure 7 It is a schematic structural diagram of the electronic device provided by the embodiments of the present disclosure. Specific embodiments
[0028] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0029] First, the technical terms and technical background related to the embodiments of the present disclosure will be described:
[0030] Apdex (Application Performance Index): An index for measuring application performance, which evaluates the performance of an application or service based on user satisfaction.
[0031] P99 Latency (99Percentile Latency): A service metric used to measure the time required for 99% of all requests to complete.
[0032] QPS (Queries Per Second): A service metric used to measure the number of requests that a system can process per unit time, usually used to measure the performance of a database, server, or network service.
[0033] Availability: An index for measuring the normal operation ability of a system, service, or component. It describes the degree to which a system can provide services when needed. High availability means that the system rarely or almost never fails and can continuously provide services to users.
[0034] Configuration item: usually used to store various parameters and settings that a program or system needs to use during operation. In the disclosed embodiment, the configuration item of the cache database can be any adjustable system parameter. For example, the configuration item can be active defrag cycle max (maximum number of cycles for memory defragmentation), or it can be Append-only (append only), or it can be Slave Read Only (read only from the server). Among them, Append-only means storing and processing data in append-only mode. Among them, Slave Read Only is a setting in the database replication architecture, which is used to indicate that the slave server (slave) is only used for read operations, not for write operations.
[0035] Business indicator: In an optimization problem, a business indicator is an indicator that the optimization algorithm attempts to maximize or minimize. In the embodiment of the present disclosure, the business indicator of the cache database may include one of the following: Apdex, P99 Latency, QPS, Availability.
[0036] The distributed cache database deployed in a multi-agent cloud environment needs to ensure high reliability, high availability, high throughput and low latency. However, since the usage scenarios often have diverse business requests, high concurrent access, real-time data return requirements and high deployment architecture complexity, it is difficult to optimize the system around business goals. With the increase in cluster scale and business system retrieval concurrency, the difficulty of manual parameter optimization around core business indicators such as Apdex, P99 Latency, QPS, and Availability increases rapidly.
[0037] Currently, the main means to solve this problem is to require professional and experienced database administrators to manually configure some configuration items of the cluster system according to business load requirements and operating characteristics, and then continuously try and error to optimize the business indicators on the core business side.
[0038] However, since there are a large number of configuration items in the cluster system, it takes a long time to manually determine which configuration item to adjust and its adjustment parameters, so the efficiency of optimizing business indicators using existing technologies is low.
[0039] It can be seen that how to quickly determine which configuration item of the cache database to adjust and how much to adjust in order to improve the efficiency of optimizing business indicators is a technical problem that needs to be solved urgently.
[0040] In view of the technical problems in the prior art, the technical concept of the inventor is as follows: adopting Causal AI (Causal Artificial Intelligence) technology to discover the causal relationship between configuration items and business metrics from distributed cache data, and realizing the adaptive tuning of the cache database with distributed deployment in the cloud environment.
[0041] Correspondingly, the specific steps may include: First, in response to a tuning request for the cloud distributed cache database, obtain the expected values of at least one business metric carried in the tuning request. Then, according to the expected values of at least one business metric and the network causal model corresponding to the cloud distributed cache database, determine the tuning strategy corresponding to the cloud distributed cache database. The tuning strategy includes at least one target configuration item and the adjustment parameter of each target configuration item, where the network causal model is used to build the causal relationship from configuration items to business metric variables based on various types of data analysis, and predict and analyze the configuration parameters of configuration items for the expected values of business metrics. Finally, according to the tuning strategy corresponding to the cloud distributed cache database, adjust the configuration parameters of each target configuration item of the cloud distributed cache database to their respective corresponding adjustment parameters.
[0042] In this technical solution, since the network causal model is used to build the causal relationship from configuration items to business metric variables based on various types of data analysis, and predict and analyze the configuration parameters of configuration items for the expected values of business metrics, the network causal model can determine the adjustment parameters of at least one target configuration item corresponding to the expected value of the business metric, realize counterfactual reasoning to generate the tuning strategy of the cloud distributed cache database, and realize the adaptive tuning of the cloud distributed cache database, thus improving the optimization efficiency of business metrics.
[0043] Next, the application scenarios of the embodiments of the present disclosure will be explained:
[0044] The method for adaptive tuning of a cloud distributed cache database provided by the embodiments of the present disclosure can be applied to various scenarios for tuning the cloud distributed cache database. Figure 1 FIG. is a schematic diagram of the application scenario of a method for adaptive tuning of a cloud distributed cache database provided by an embodiment of the present disclosure. As Figure 1 shown, the terminal 101 transmits a tuning request for the cloud distributed cache database to the server 102 through a wireless network. The server 102 receives the tuning request, determines a tuning strategy including at least one target configuration item and the adjustment parameter of each target configuration item through the method for adaptive tuning of the cloud distributed cache database provided by the embodiments of the present disclosure, and according to the tuning strategy, adjusts each target configuration item of the cloud distributed cache database to its respective corresponding adjustment parameter.
[0045] The following uses a specific scenario as an example to illustrate: During the operation of the Cloud Distributed Cache Database, the database administrator discovers from the business side that the Apdex metric of the cache database is lower than 0.5. If, under the load conditions in the recent period of time, we want to know which configuration variables need to be adjusted and to what values in order to achieve the Apdex metric level of 0.8. At this time, the target variable value of Apdex = 0.8 can be input into the causal inference engine through the configuration change manager. The causal inference engine can find the observed metric covariates linearly related to it through the SEM causal model and calculate their value ranges. With the prior knowledge of the values of some observed metric variables, the maximum posterior probability of other metrics can be inferred using the CBN causal network model, and then the maximum probability values of each configuration variable can be obtained, thus obtaining the tuning strategy for this distributed cache database.
[0046] Figure 2 Flow of the adaptive tuning method for the Cloud Distributed Cache Database provided by the embodiments of the present disclosure Figure 1 In the embodiments of the present disclosure, the execution subject of this adaptive tuning method can be a terminal or a server. As Figure 2 shown, this adaptive tuning method may include:
[0047] S201. In response to a tuning request for the Cloud Distributed Cache Database, obtain the expected values of at least one business metric carried in the tuning request.
[0048] In the embodiments of the present disclosure, the business metric can be any metric for the Cloud Distributed Cache Database. Optionally, the business metric can be a business metric, a stability metric, a security metric, etc. of the Cloud Distributed Cache Database. Exemplarily, the business metric may include: Apdex metric, P99 Latency metric, QPS metric, Availability metric, RTT (Round Trip Time, request response time) metric, RPS (Requests Per Second, requests processed per second) metric.
[0049] Exemplarily, as Figure 3 shown, the system tuning controller can receive the tuning request, obtain the expected values of at least one business metric carried in the tuning request, and generate a vector set to send to the configuration change strategy generator.
[0050] S202. Determine the tuning strategy corresponding to the cloud distributed cache database according to the expected values of at least one business metric and the network causal model corresponding to the cloud distributed cache database. The tuning strategy includes at least one target configuration item and the adjustment parameters of each target configuration item. The network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the business metric based on the causal relationship from the configuration item to the business metric variable constructed from various types of data analysis.
[0051] In some embodiments, the network causal model includes a multi-layer heterogeneous causal model constructed based on various types of data, where different causal models are constructed and generated based on different types of data; among them, the various types of data include verification data, cloud observation data, empirical data, etc.
[0052] In some embodiments, determining the tuning strategy corresponding to the cloud distributed cache database according to the expected values of at least one business metric and the network causal model corresponding to the cloud distributed cache database includes: calling each layer of the causal model in sequence according to the call order of the multi-layer heterogeneous causal model in the network causal model from the bottom layer to the top layer, using the output of the current layer of the causal model as the input of the upper layer of the causal model connected to it, and obtaining at least one target configuration item output by the top layer of the causal model and the adjustment parameters of each target configuration item as the tuning strategy corresponding to the cloud distributed cache database.
[0053] Optionally, the network causal model includes a three-layer heterogeneous causal model constructed based on various types of data; among them, the first-layer causal model of the three-layer heterogeneous causal model is constructed based on the verification data among various types of data and is used to determine the causal relationship between the configuration item and the running metric. The second-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data among various types of data and is used to determine the causal relationship between the running metrics. The third-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data and empirical data among various types of data and is used to determine the causal relationship between the running metric and the business metric.
[0054] Accordingly, a tuning strategy corresponding to the cloud-edge distributed cache database is determined according to the expected values of at least one service metric and the network causal model corresponding to the cloud-edge distributed cache database, including: based on the expected values of at least one service metric and the third-layer causal model, searching for at least one first operating metric causally related and the value range of each first operating metric; according to at least one first operating metric, the value range of each first operating metric, and the second-layer causal model, determining at least one second operating metric causally related to at least one first operating metric and the value range of each second operating metric; according to at least one first operating metric, the value range of each first operating metric, at least one second operating metric, the value range of each second operating metric, and the first-layer causal model, determining at least one target configuration item and the adjustment parameter of each target configuration item.
[0055] Optionally, the network causal model includes a three-layer heterogeneous causal model of network communication from the bottom layer to the top layer. Among them, the first-layer causal model includes the Rubin Causal Model (RCM). The second-layer causal model includes the Causal Bayesian Network (CBN), and the third-layer causal model includes the Structure Equation Model (SEM).
[0056] Exemplarily, as Figure 3 shown, the configuration change policy generator can receive the vector set of the expected values of at least one service metric sent by the system tuning controller, send it to the causal inference engine to obtain the vector set containing the configuration items to be changed, and generate a configuration change plan (i.e., the tuning strategy) and send it to the configuration change manager.
[0057] S203. According to the tuning strategy corresponding to the cloud-edge distributed cache database, adjust the configuration parameters of each target configuration item of the cloud-edge distributed cache database to their respective corresponding adjustment parameters.
[0058] Exemplarily, as Figure 3 shown, the configuration variable manager can perform configuration changes on the distributed cache database deployed in the cloud-edge according to the received tuning strategy.
[0059] An embodiment of the present disclosure provides a method for adaptive tuning of an inter-cloud distributed cache database: in response to a tuning request for the inter-cloud distributed cache database, obtain the expected values of at least one service metric carried in the tuning request; according to the expected values of the at least one service metric and the network causal model corresponding to the inter-cloud distributed cache database, determine the tuning strategy corresponding to the inter-cloud distributed cache database, where the tuning strategy includes at least one target configuration item and the adjustment parameter of each target configuration item, and the network causal model is used to construct the causal relationship from the configuration item to the service metric variable based on various types of data analysis, and predict and analyze the configuration parameters of the configuration item for the expected value of the service metric; according to the tuning strategy corresponding to the inter-cloud distributed cache database, adjust the configuration parameters of each target configuration item of the inter-cloud distributed cache database to their respective corresponding adjustment parameters. In this technical solution, since the network causal model is used to construct the causal relationship from the configuration item to the service metric variable based on various types of data analysis, and predict and analyze the configuration parameters of the configuration item for the expected value of the service metric, the network causal model can determine the adjustment parameters of at least one target configuration item corresponding to the expected value of the service metric, realize counterfactual reasoning to generate the tuning strategy of the inter-cloud distributed cache database, and realize the adaptive tuning of the inter-cloud distributed cache database, so the optimization efficiency of the service metric is improved.
[0060] It should be noted that in the embodiment of the present disclosure, we first analyze and model three different types of data obtained during the test period and the operation period of the distributed cache database respectively, and then construct a network causal model through an innovative three-layer heterogeneous causal model fusion method.
[0061] In some embodiments, as Figure 4 shown, before obtaining the expected values of at least one service metric carried in the tuning request in response to the tuning request for the inter-cloud distributed cache database, heterogeneous causal models can be constructed respectively by using different modeling methods through the three types of data collected, and then the multiple causal models constructed are further fused to generate a network causal model. Among them, the method for generating the network causal model includes:
[0062] S401. Obtain the monitoring metric data set of the inter-cloud distributed cache database, where the monitoring metric data set includes verification data participating in the system test of the inter-cloud distributed cache database and cloud observation data and empirical data during the operation of the inter-cloud distributed cache database.
[0063] In this step, the three types of data include verification data, cloud observation data, and empirical data.
[0064] Among them, the verification data can be the verification data collected through a randomized controlled experiment by grouping, or can be obtained from cloud observation data through methods to eliminate bias, such as the do-operator (operator overloading method). Exemplarily, the impact of enabling or disabling memory fragmentation in the cloud distributed cache database (Redis) on operating metrics such as the CPU usage rate, memory usage rate, and cache hit rate of Redis service nodes. The cache hit rate refers to the ratio of the number of times the requested data is found in the cache to the total number of requests. Among them, the configuration item is memory fragmentation, and the test parameters of the configuration item include: enabling or disabling memory fragmentation. Among them, the operating metrics to be verified include the CPU usage rate, memory usage rate, and cache hit rate of the service node. This type of data has a small volume and is easy to remove the influence of bias and confounding variables to obtain relatively stable information, which is of greater value for generating safe and reliable causal relationships.
[0065] Among them, the cloud observation data can be the operating metric data directly observed during the system operation phase. Optionally, the cloud observation data includes multiple operating metrics. This type of data has a large volume, a fast change speed, and records the real-time detailed information of the internal performance and stability fluctuations of the system. However, the cloud observation data does not include the data caused by intervention actions such as configuration changes and instance migrations. Due to the influence of bias and confounding variables introduced by external interventions (such as changes in load and cloud infrastructure configuration), it is difficult for us to directly generate stable and reliable causal relationships from the cloud observation data.
[0066] Among them, the empirical data can be the hidden variable data that cannot be directly observed. Optionally, the empirical data includes multiple business metrics. For example: customer experience Apdex, system availability, etc. It is necessary to define a mathematical model by combining historical statistical data, industry standards, and business requirements. This type of data is what we call empirical data. Business operation empirical data is usually used to define the mapping relationship between the observed metric variables and the hidden target metric variables, and obtain the quantitative value of the target hidden variable through determining mathematical formulas or non-deterministic probabilistic inferences (such as Bayesian inference, etc.) for decision-making reference.
[0067] Exemplarily, such as Figure 3As shown in the figure, the monitoring metric data collector can be used to collect the metric data during the system testing phase and the running phase from the cloud environment. The time series event data extracted from the cloud environment log data can be used to obtain the monitoring metric dataset of the cloud distributed cache database, and the monitoring metric dataset can be stored in the monitoring database. The monitoring structure data collector can be used to collect the deployment relationships of the various components of the cache database deployed in the cloud environment from monitoring systems such as CMDB and APM, and the mutual call relationships during the execution of queries, so as to initially correlate the running metrics of different components and store them in the monitoring database. Among them, the monitoring database is used to centrally store the monitoring data of the cloud distributed database during the testing period and the running period, so as to drive the subsequent Causal AI causal discovery algorithm to generate a causal model.
[0068] S402. Generate a network causal model according to the monitoring metric dataset.
[0069] In some embodiments, the cloud observation data includes multiple running metrics, and the empirical data includes multiple business metrics. The network causal model includes a three-layer heterogeneous causal model; among them, the first-layer causal model of the three-layer heterogeneous causal model is used to determine the causal relationship between the configuration items and the running metrics, the second-layer causal model of the three-layer heterogeneous causal model is used to determine the causal relationship between the running metrics, and the third-layer causal model of the three-layer heterogeneous causal model is used to determine the causal relationship between the running metrics and the business metrics; correspondingly, this step may include the following steps (1) to (2):
[0070] (1) Generate the first-layer causal model corresponding to the cloud distributed cache database according to the intervention effects of the test parameters of the configuration items and the test values of at least one running metric to be verified in each verification data, and generate the second-layer causal model corresponding to the cloud distributed cache database according to the dependency relationships and subordination relationships among multiple running metrics, and generate the third-layer causal model corresponding to the cloud distributed cache database according to multiple business metrics, multiple running metrics, and a preset functional relationship.
[0071] In the implementation of the present disclosure, as Figure 3 shown, through the RCM causal analysis model generator, based on the set of configuration item (knob) variables provided by the distributed cache database and the set of monitoring metric variables collected in the monitoring database, the RCM causal analysis method is used to fit the historical monitoring data to find the RCM model with significant causal relationships.
[0072] In some embodiments, the first-layer causal model includes a potential causal RCM model. Accordingly, a first-layer causal model corresponding to the cloud distributed cache database is generated according to the intervention effects of the test parameters of the configuration items and the test values of at least one running metric to be verified in each validation data, including: for the test parameters of the configuration items and the test values of at least one running metric to be verified in each validation data, determining the intervention effect of the test parameters of the configuration items on the test values of at least one running metric to be verified; according to the intervention effect, determining the corresponding relationship between at least one configuration item, the adjustment parameter of each configuration item, at least one running metric, and the value range of each running metric, and generating an RCM model corresponding to the cloud distributed cache database.
[0073] Optionally, the intervention effect includes an average intervention effect and / or a conditional intervention effect.
[0074] It should be noted that the validation data collection needs to be data obtained from load testing according to a specific validation environment during the development and testing phase of the cache database. During the load testing process, since there are potential unobservable variables on the customer side and the running environment dependency side that are uncontrollable during the production phase, we design test cases and analyze the average intervention effect (Average Treatment Effect, ATE) and / or the conditional average intervention effect (Conditional Average Treatment Effect, CATE) based on the Potential Outcome framework to obtain a relatively stable RCM causal relationship model between the configuration items corresponding to the control variables and the observed variables.
[0075] In some embodiments, the second-layer causal model includes a Bayesian network causal CBN model. Accordingly, a second-layer causal model corresponding to the cloud distributed cache database is generated according to the dependency relationship and subordination relationship between multiple running metrics, including: according to the Application Performance Management (APM) and the Configuration Management Database (CMDB), determining the dependency relationship and subordination relationship between multiple running metrics corresponding to the target event, where the target event includes a scheduling event or an upgrade event; according to the dependency relationship and subordination relationship between multiple running metrics corresponding to the target event and the Conditional Probability Table (CPT), generating a CBN model corresponding to the cloud distributed cache database.
[0076] Exemplarily, as Figure 3 shown, through an SCM (Structural Causal Model) generator: based on the set of monitoring metric variables collected in the monitoring database, a CBN model is generated by fitting historical monitoring data using the SCM algorithm.
[0077] It should be noted that cloud observation data are observed variables extracted from metrics, logs and other information collected by monitoring tools during the operation stage of the cache database. Due to the fast change speed and the influence of potential unobservable variable changes in the business load and cloud operation environment, we cannot adopt the RCM method to construct a causal model. To address the above technical problems, we propose to first collect event logs including scheduling, upgrading, etc. from the environment operation logs; then extract the events in the logs as an observed variable reflecting the change of the environment event state, and merge it with the limited observed variables obtained from the monitoring metrics; finally, use the Probability Graph Model (PGM) modeling method to generate a causal Bayesian network graph model (a directed graph model). During the generation process, the deployment architecture collected from data sources such as APM (Application Performance Management) and CMDB (Configuration Management Database) is used as the infrastructure for the association relationship between metrics, and the causal relationship is generated on this basis.
[0078] In some embodiments, the third-layer causal model includes a Structural Equation Model (SEM). The preset functional relationships include a first functional relationship and a second functional relationship. Correspondingly, according to multiple business metrics, multiple operation metrics and the preset functional relationships, a third-layer causal model corresponding to the cloud-interconnected distributed cache database is generated, including: for each business metric, constructing a first functional relationship between the value of the business metric and the value range of at least one first target operation metric; determining at least one second target operation metric linearly related to the first target operation metric from multiple operation metrics, and determining a second functional relationship between the first target operation metric and the second target operation metric; according to the first functional relationship and the second functional relationship, establishing an association relationship between the value range of at least one second target operation metric and the value of the business metric, and generating an SEM corresponding to the cloud-interconnected distributed cache database.
[0079] The empirical data comes from target metrics that cannot be directly observed on the business system side, such as Apdex user experience, RPS (requests per second) average number of requests processed per second, etc. These metrics cannot be directly observed and need to be calculated from the observable metrics on the business side according to the established calculation logic. For example, the calculation formula for Apdex is:
[0080] Apdex index = (number of requests with expected responses + number of tolerable request responses / 2) / total number of request samples
[0081] During the operation period, since the public network with unobservable operating status exists between the cloud cache database and the customer-side business system, the network latency, congestion, and jitter conditions are unknown. However, since the customer-side request latency y in the Apdex calculation formula can be calculated, and the execution latency x in the database can be observed, and they conform to a linear correlation relationship, we can construct a causal influence relationship based on the collected cloud observation data through linear regression: Y = ax + e; where a is the correlation coefficient, and e is the interference caused by factors such as the public network connection status.
[0082] According to the influence relationship between the known distributed cache database business target variables and the observed variables, combined with the collected monitoring metric data, a Structural Equation Model (SEM model) is generated to construct a linear causal association relationship between the distributed cache database operation period monitoring metric variables and the business target variables in the form of a path diagram.
[0083] Exemplarily, as Figure 3 shown, an SEM model can be generated through an SEM structural equation model generator: based on a predefined set of business target variables (such as Apdex, Availability, etc.), the most relevant observed variables associated in the CBM based on business experience are used, and a regression algorithm is used to fit a linear function to construct the SEM model.
[0084] (2) Integrate the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate a network causal model.
[0085] Optionally, this step may include: integrating the first-layer causal model and the second-layer causal model with multiple third operating metrics as the combination points, and integrating the second-layer causal model and the third-layer causal model with multiple fourth operating metrics as the combination points to generate a network causal model. Among them, the third operating metric is an operating metric that is causally related to the test parameters of the configuration item based on the first-layer causal model; the third operating metric is an operating metric that is causally related to the expected value of the business metric based on the third-layer causal model.
[0086] In some embodiments, the first-layer causal model includes a potential causal RCM model, the second-layer causal model includes a Bayesian network causal CBN model, and the third-layer causal model includes a structural equation SEM model. Exemplarily, in the process of generating the network causal model, we first start from the configuration items, find the observed variables with the highest causal significance associated with the configuration items, and use this set of observed variables as the junction point to splice the RCM model and the CBN model. Secondly, starting from the business metrics, search for the set of observed variables that are causally related, and use this set as the junction point to splice the SEM model and the CBN model. By fusing the RCM model, the CBN model, and the SEM model, a causal inference model (i.e., the network causal model) that can go from the target effect to the intervention cause and from the intervention cause to the target effect is obtained.
[0087] Exemplarily, as Figure 3 shown, the above-mentioned causal inference model is generated by the causal inference engine fusing the already constructed RCM causal model, CBN causal model, and SEM causal model respectively. An inference service call interface is provided in the form of a service in the loaded inference engine.
[0088] The generation process of the NeoCG network causal model is described below through specific examples.
[0089] (1) Collect the monitoring data during the operation period of the distributed cache data and the environmental log data. Define the database configurable knob as the treatment variable, the monitoring metric collection data as the observable variable, and the unobservable environmental state and the customer-side experience-related state as the latent variable.
[0090] (2) Find the average treatment effect (ATE) between the database configuration item and the observable variable by performing potential outcome analysis on the configuration variables during the database operation period, and determine the causal relationship with the target observable variable.
[0091] (3) Based on the physical deployment architecture read out through systems such as CMDB and APM, associate the relationships between the observed variables. At the same time, extract the events in the operation environment log as an observable metric to reflect the events occurring in the environmental latent variable. On this basis, combined with the monitoring data, find the probabilistic correlation and causal relationship between the observed variables through the causal architecture discovery algorithm and the Bayesian network generation algorithm.
[0092] (4) Construct a structural equation between the observed variables and the state target variables related to the customer-side experience based on the Structure Equation Model (SEM), and generate the corresponding parameters through cloud observation data fitting to construct the relationship between the observed variables and the target hidden variables;
[0093] (5) Integrate the models separately modeled by potential outcome analysis, structural causal network, and structural equation model above into a causal graph model with a three-layer hybrid structure, which we call the NeoCG network causal model (analogous to the six-layer neural network hybrid structure of the neocortex in the prefrontal cortex of the human brain, so it is named Neo Causal Graph).
[0094] It should be noted that after the state of the latent variable in the operating environment changes, the network causal model can also be automatically updated.
[0095] In some embodiments, after the configuration change manager receives the tuning policy, it performs configuration changes on the target cache database, and observes through the monitoring system whether the tuning policy enables the target variable to achieve the expectation; if the expectation is not achieved, it calls to update the monitoring metric dataset and regenerates the model. Correspondingly, after adjusting each target configuration item of the cloud distributed cache database to its corresponding adjustment parameter according to the tuning policy corresponding to the cloud distributed cache database, it further includes: if the empirical data corresponding to the business metric does not meet the expected value of the business metric, update the monitoring metric dataset; generate a new network causal model through the updated monitoring metric dataset. Exemplarily, the expected value of the business metric Apdex is 0.8. If the empirical data corresponding to the business metric Apdex is 0.7, which does not meet the expected value of the business metric at this time, the monitoring metric dataset is updated.
[0096] In this step, the amount of data in the updated monitoring metric dataset is larger than that in the original monitoring metric dataset, and thus the network causal model can be incrementally updated.
[0097] Exemplarily, as Figure 5 shown, the steps for automatically updating the network causal model include:
[0098] (1) Define the target variable expectation vector according to the metric expectations of the data administrator for the customer experience of the customer-side cache database (such as: round-trip request response time RTT, requests processed per second throughput RPT).
[0099] (2) After the system tuning controller receives the expectation vector, it sends it to the configuration change policy generator and then forwards it to the NeoCG inference engine.
[0100] (3) The NeoCG inference engine infers the set of variable configuration Knob items and encapsulates the values into a vector and returns it to the configuration change policy generator.
[0101] (4) The configuration change policy generator obtains the inference result and then generates a tuning policy for the target cache database.
[0102] (5) After receiving the tuning policy, the configuration change manager performs variable configuration on the target cache database and observes through the monitoring system whether the tuning policy enables the target variable to achieve the expectation.
[0103] (6) If the expectation is not achieved, the updated monitoring metric data set is called to regenerate the CBN model.
[0104] (7) After the NeoCG inference engine obtains the updated CBN model, it incrementally updates the NeoCG model again and continues to execute step (3).
[0105] In the embodiment of the present disclosure, since after the database tuning policy is generated, we add a process for judging the execution effect of the tuning policy, and can automatically update the mechanism of the network causal model after the state of the potential variables in the running environment changes, the robustness of the system is improved.
[0106] Figure 6 The structural schematic diagram of the cloud distributed cache database adaptive tuning system provided by the embodiment of the present disclosure is as Figure 6 shown. The cloud distributed cache database adaptive tuning system includes:
[0107] An acquisition unit 601, configured to respond to a tuning request for the cloud distributed cache database and acquire the expected values of at least one service metric carried in the tuning request;
[0108] A determination unit 602, configured to determine a tuning policy corresponding to the cloud distributed cache database according to the expected values of the at least one service metric and the network causal model corresponding to the cloud distributed cache database, where the tuning policy includes at least one target configuration item and the adjustment parameter of each target configuration item, and the network causal model is used to construct the causal relationship from the configuration item to the service metric variable based on various types of data analysis, and predict and analyze the configuration parameters of the configuration item for the expected value of the service metric;
[0109] An adjustment unit 603, configured to adjust the configuration parameters of each target configuration item of the cloud distributed cache database to their respective corresponding adjustment parameters according to the tuning policy corresponding to the cloud distributed cache database.
[0110] According to one or more embodiments of the present disclosure, the network causal model includes a multi-layer heterogeneous causal model constructed based on multiple types of data, where different causal models are constructed and generated based on different types of data; among them, the multiple types of data include verification data, cloud observation data, empirical data, etc.
[0111] According to one or more embodiments of the present disclosure, the network causal model includes a three-layer heterogeneous causal model constructed based on multiple types of data; among them, the first-layer causal model of the three-layer heterogeneous causal model is constructed based on the verification data among the multiple types of data and is used to determine the causal relationship between configuration items and operation metrics. The second-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data among the multiple types of data and is used to determine the causal relationship between operation metrics. The third-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data and empirical data among the multiple types of data and is used to determine the causal relationship between operation metrics and business metrics.
[0112] According to one or more embodiments of the present disclosure, the determining unit 602 determines the tuning strategy corresponding to the cloud-interconnected distributed cache database according to the expected value of the at least one business metric and the network causal model corresponding to the cloud-interconnected distributed cache database, which specifically includes: calling each layer of causal model in turn according to the calling order of the multi-layer heterogeneous causal model in the network causal model from the bottom layer to the top layer, using the output of the current layer of causal model as the input of the upper layer of causal model connected to it, and obtaining at least one target configuration item output by the top-layer causal model and the adjustment parameter of each target configuration item as the tuning strategy corresponding to the cloud-interconnected distributed cache database.
[0113] According to one or more embodiments of the present disclosure, the network causal model includes a three-layer heterogeneous causal model that communicates layer by layer from the bottom layer to the top layer through the network, where the first-layer causal model includes a potential causal RCM model, the second-layer causal model includes a Bayesian network causal CBN model, and the third-layer causal model includes a structural equation SEM model.
[0114] According to one or more embodiments of the present disclosure, the cloud-interconnected distributed cache database adaptive tuning system further includes: a model generation unit; the model generation unit is used to obtain the monitoring metric dataset of the cloud-interconnected distributed cache database, and the monitoring metric dataset includes verification data participating in the cloud-interconnected distributed cache database system test and cloud observation data and empirical data during the operation of the cloud-interconnected distributed cache database; and generate the network causal model according to the monitoring metric dataset.
[0115] According to one or more embodiments of the present disclosure, the cloud-edge distributed cache database adaptive tuning system further includes: an update unit; the update unit is configured to update the monitoring metric dataset if the empirical data corresponding to the service metric does not meet the expected value of the service metric; and generate a new network causal model through the updated monitoring metric dataset.
[0116] According to one or more embodiments of the present disclosure, the model generation unit generates the network causal model according to the monitoring metric dataset, including: generating a first-layer causal model corresponding to the cloud-edge distributed cache database according to the intervention effect of the test parameters of the configuration item and the test values of at least one running metric to be verified in each validation data, and generating a second-layer causal model corresponding to the cloud-edge distributed cache database according to the dependency relationship and subordination relationship among the multiple running metrics, and generating a third-layer causal model corresponding to the cloud-edge distributed cache database according to the multiple service metrics, the multiple running metrics, and a preset functional relationship; fusing the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model.
[0117] According to one or more embodiments of the present disclosure, the model generation unit fuses the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model, including: fusing the first-layer causal model and the second-layer causal model with multiple third running metrics as the combination points, and fusing the second-layer causal model and the third-layer causal model with multiple fourth running metrics as the combination points to generate the network causal model; wherein the third running metric is a running metric that is causally related to the test parameters of the configuration item determined based on the first-layer causal model; the third running metric is a running metric that is causally related to the expected value of the service metric determined based on the third-layer causal model.
[0118] According to one or more embodiments of the present disclosure, the first-layer causal model includes a potential causal RCM model; correspondingly, the model generation unit generates the first-layer causal model corresponding to the cloud-edge distributed cache database according to the intervention effect of the test parameters of the configuration item and the test values of at least one running metric to be verified in each validation data, including: for the test parameters of the configuration item and the test values of at least one running metric to be verified in each validation data, determining the intervention effect of the test parameters of the configuration item on the test values of the at least one running metric to be verified; and determining the corresponding relationship between at least one configuration item, the adjustment parameter of each configuration item, at least one running metric, and the value range of each running metric according to the intervention effect to generate the potential causal model corresponding to the cloud-edge distributed cache database.
[0119] According to one or more embodiments of the present disclosure, the second-layer causal model includes a Bayesian network causal (CBN) model; correspondingly, the model generation unit generates the second-layer causal model corresponding to the cloud-edge distributed cache database according to the dependency relationship and subordination relationship among the multiple operation metrics, including: determining the dependency relationship and subordination relationship among the multiple operation metrics corresponding to a target event according to application performance management and configuration management databases, where the target event includes a scheduling event or an upgrade event; generating the Bayesian network causal model corresponding to the cloud-edge distributed cache database according to the dependency relationship, subordination relationship, and conditional probability table among the multiple operation metrics corresponding to the target event.
[0120] According to one or more embodiments of the present disclosure, the third-layer causal model includes a structural equation (SEM) model; correspondingly, the model generation unit generates the third-layer causal model corresponding to the cloud-edge distributed cache database according to the multiple service metrics, the multiple operation metrics, and a preset functional relationship, including: constructing a first functional relationship between the value of a service metric and the value range of at least one first target operation metric for each service metric; determining at least one second target operation metric linearly related to the first target operation metric from the multiple operation metrics, and determining a second functional relationship between the first target operation metric and the second target operation metric; establishing an association relationship between the value range of the at least one second target operation metric and the value of the service metric according to the first functional relationship and the second functional relationship, and generating the structural equation model corresponding to the cloud-edge distributed cache database.
[0121] Reference Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The electronic device 700 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0122] As Figure 7As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0123] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0124] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0125] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0126] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0127] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0128] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0130] The units described in the embodiments of this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0131] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a method for adaptively tuning an inter-cloud distributed cache database, including:
[0134] In response to a tuning request for the inter-cloud distributed cache database, obtaining an expected value of at least one service metric carried in the tuning request;
[0135] According to the expected value of the at least one service metric and the network causal model corresponding to the inter-cloud distributed cache database, determining a tuning strategy corresponding to the inter-cloud distributed cache database, where the tuning strategy includes at least one target configuration item and an adjustment parameter for each target configuration item, and the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the service metric based on the causal relationship from the configuration item to the service metric variable constructed from various types of data analysis;
[0136] According to the tuning strategy corresponding to the inter-cloud distributed cache database, adjusting the configuration parameters of each target configuration item of the inter-cloud distributed cache database to their respective corresponding adjustment parameters.
[0137] According to one or more embodiments of the present disclosure, the network causal model includes a multi-layer heterogeneous causal model constructed based on various types of data, where different causal models are constructed and generated based on different types of data; and the various types of data include verification data, cloud observation data, empirical data, and the like.
[0138] According to one or more embodiments of the present disclosure, the network causal model includes a three - layer heterogeneous causal model constructed based on multiple types of data; wherein, the first - layer causal model of the three - layer heterogeneous causal model is constructed based on the verification data among the multiple types of data and is used to determine the causal relationship between configuration items and operation metrics, the second - layer causal model of the three - layer heterogeneous causal model is constructed based on the cloud observation data among the multiple types of data and is used to determine the causal relationship between operation metrics, and the third - layer causal model of the three - layer heterogeneous causal model is constructed based on the cloud observation data and empirical data among the multiple types of data and is used to determine the causal relationship between operation metrics and business metrics.
[0139] According to one or more embodiments of the present disclosure, determining the tuning strategy corresponding to the cloud - edge distributed cache database according to the expected value of the at least one business metric and the network causal model corresponding to the cloud - edge distributed cache database includes: calling each layer of the causal model in sequence according to the calling order of the multi - layer heterogeneous causal model in the network causal model from the bottom layer to the top layer, using the output of the current - layer causal model as the input of the upper - layer causal model connected thereto, and obtaining at least one target configuration item output by the top - layer causal model and the adjustment parameter of each target configuration item as the tuning strategy corresponding to the cloud - edge distributed cache database.
[0140] According to one or more embodiments of the present disclosure, the network causal model includes a three - layer heterogeneous causal model that communicates layer by layer from the bottom layer to the top layer in the network, wherein the first - layer causal model includes a potential causal model, the second - layer causal model includes a Bayesian network causal model, and the third - layer causal model includes a structural equation model.
[0141] According to one or more embodiments of the present disclosure, before obtaining the expected value of the at least one business metric carried in the tuning request in response to a tuning request for the cloud - edge distributed cache database, the method further includes: obtaining a monitoring metric data set of the cloud - edge distributed cache database, where the monitoring metric data set includes verification data participating in the cloud - edge distributed cache database system test and cloud observation data and empirical data during the operation of the cloud - edge distributed cache database; and generating the network causal model according to the monitoring metric data set.
[0142] According to one or more embodiments of the present disclosure, after adjusting each target configuration item of the cloud - edge distributed cache database to its corresponding adjustment parameter according to the tuning strategy corresponding to the cloud - edge distributed cache database, it further includes: if the empirical data corresponding to the business metric does not meet the expected value of the business metric, updating the monitoring metric data set; and generating a new network causal model through the updated monitoring metric data set.
[0143] According to one or more embodiments of the present disclosure, generating the network causal model based on the monitoring metric dataset includes: generating a first-layer causal model corresponding to the cloud distributed cache database according to the intervention effect of the test parameters of the configuration items and the test values of at least one running metric to be verified in each verification data; generating a second-layer causal model corresponding to the cloud distributed cache database according to the dependency and subordination relationships among the multiple running metrics; and generating a third-layer causal model corresponding to the cloud distributed cache database according to the multiple service metrics, the multiple running metrics, and a preset functional relationship; fusing the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model.
[0144] According to one or more embodiments of the present disclosure, fusing the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model includes: fusing the first-layer causal model and the second-layer causal model with multiple third running metrics as the combination points, and fusing the second-layer causal model and the third-layer causal model with multiple fourth running metrics as the combination points to generate the network causal model; wherein the third running metric is a running metric that is causally related to the test parameters of the configuration item determined based on the first-layer causal model; and the third running metric is a running metric that is causally related to the expected value of the service metric determined based on the third-layer causal model.
[0145] According to one or more embodiments of the present disclosure, the first-layer causal model includes a potential causal model; correspondingly, generating the first-layer causal model corresponding to the cloud distributed cache database according to the intervention effect of the test parameters of the configuration items and the test values of at least one running metric to be verified in each verification data includes: for the test parameters of the configuration items and the test values of at least one running metric to be verified in each verification data, determining the intervention effect of the test parameters of the configuration item on the test values of the at least one running metric to be verified; and according to the intervention effect, determining the corresponding relationship between at least one configuration item, the adjustment parameter of each configuration item, at least one running metric, and the value range of each running metric to generate the potential causal model corresponding to the cloud distributed cache database.
[0146] According to one or more embodiments of the present disclosure, the second-layer causal model includes a Bayesian network causal model; correspondingly, generating the second-layer causal model corresponding to the cloud-edge distributed cache database according to the dependency and subordination relationships among the multiple operation metrics includes: determining the dependency and subordination relationships among the multiple operation metrics corresponding to a target event according to the application performance management and configuration management database, where the target event includes a scheduling event or an upgrade event; generating the Bayesian network causal model corresponding to the cloud-edge distributed cache database according to the dependency and subordination relationships among the multiple operation metrics corresponding to the target event and the conditional probability table.
[0147] According to one or more embodiments of the present disclosure, the third-layer causal model includes a structural equation model; correspondingly, generating the third-layer causal model corresponding to the cloud-edge distributed cache database according to the multiple service metrics, the multiple operation metrics, and the preset functional relationship includes: for each service metric, constructing a first functional relationship between the value of the service metric and the value range of at least one first target operation metric; determining at least one second target operation metric linearly related to the first target operation metric from the multiple operation metrics, and determining a second functional relationship between the first target operation metric and the second target operation metric; establishing an association relationship between the value range of the at least one second target operation metric and the value of the service metric according to the first functional relationship and the second functional relationship, and generating the structural equation model corresponding to the cloud-edge distributed cache database.
[0148] Second, according to one or more embodiments of the present disclosure, a cloud-edge distributed cache database adaptive tuning system is provided, including:
[0149] An acquisition unit, configured to acquire the expected value of at least one service metric carried in the tuning request in response to a tuning request for the cloud-edge distributed cache database;
[0150] A determination unit, configured to determine a tuning strategy corresponding to the cloud-edge distributed cache database according to the expected value of the at least one service metric and the network causal model corresponding to the cloud-edge distributed cache database, where the tuning strategy includes at least one target configuration item and the adjustment parameter of each target configuration item, and the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the service metric based on the causal relationship from the configuration item to the service metric variable constructed by analyzing multiple types of data;
[0151] An adjustment unit, configured to adjust the configuration parameters of each target configuration item of the cloud-edge distributed cache database to their respective corresponding adjustment parameters according to the tuning strategy corresponding to the cloud-edge distributed cache database.
[0152] According to one or more embodiments of the present disclosure, the network causal model includes a multi-layer heterogeneous causal model constructed based on multiple types of data, wherein different causal models are constructed and generated based on different types of data; wherein the multiple types of data include verification data, cloud observation data, empirical data, etc.
[0153] According to one or more embodiments of the present disclosure, the network causal model includes a three-layer heterogeneous causal model constructed based on multiple types of data; wherein the first-layer causal model of the three-layer heterogeneous causal model is constructed based on the verification data among the multiple types of data and is used to determine the causal relationship between configuration items and running metrics, the second-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data among the multiple types of data and is used to determine the causal relationship between running metrics, and the third-layer causal model of the three-layer heterogeneous causal model is constructed based on the cloud observation data and empirical data among the multiple types of data and is used to determine the causal relationship between running metrics and business metrics.
[0154] According to one or more embodiments of the present disclosure, the determining unit determines the tuning strategy corresponding to the cloud-interconnected distributed cache database according to the expected value of the at least one business metric and the network causal model corresponding to the cloud-interconnected distributed cache database, which specifically includes: sequentially calling each layer of causal model according to the calling order of the multi-layer heterogeneous causal model in the network causal model from the bottom layer to the top layer, using the output of the current layer of causal model as the input of the upper layer of causal model connected thereto, and obtaining at least one target configuration item output by the top-layer causal model and the adjustment parameter of each target configuration item as the tuning strategy corresponding to the cloud-interconnected distributed cache database.
[0155] According to one or more embodiments of the present disclosure, the network causal model includes a three-layer heterogeneous causal model that communicates layer by layer from the bottom layer to the top layer through the network, wherein the first-layer causal model includes a latent causal model, the second-layer causal model includes a Bayesian network causal model, and the third-layer causal model includes a structural equation model.
[0156] According to one or more embodiments of the present disclosure, the cloud-interconnected distributed cache database adaptive tuning system further includes: a model generation unit; the model generation unit is used to obtain the monitoring metric data set of the cloud-interconnected distributed cache database, and the monitoring metric data set includes verification data participating in the cloud-interconnected distributed cache database system test and cloud observation data and empirical data during the operation of the cloud-interconnected distributed cache database; and generate the network causal model according to the monitoring metric data set.
[0157] According to one or more embodiments of the present disclosure, the cloud distributed cache database adaptive tuning system further includes: an update unit; the update unit is configured to update the monitoring metric dataset if the empirical data corresponding to the service metric does not meet the expected value of the service metric; and generate a new network causal model through the updated monitoring metric dataset.
[0158] According to one or more embodiments of the present disclosure, the model generation unit generates the network causal model according to the monitoring metric dataset, including: generating a first-layer causal model corresponding to the cloud distributed cache database according to the intervention effect of the test parameter of the configuration item and the test value of at least one running metric to be verified in each verification data, and generating a second-layer causal model corresponding to the cloud distributed cache database according to the dependency relationship and subordination relationship among the multiple running metrics, and generating a third-layer causal model corresponding to the cloud distributed cache database according to the multiple service metrics, the multiple running metrics, and a preset functional relationship; fusing the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model.
[0159] According to one or more embodiments of the present disclosure, the model generation unit fuses the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model, including: fusing the first-layer causal model and the second-layer causal model with multiple third running metrics as the combination points, and fusing the second-layer causal model and the third-layer causal model with multiple fourth running metrics as the combination points to generate the network causal model; wherein, the third running metric is a running metric causally related to the test parameter of the configuration item determined based on the first-layer causal model; the third running metric is a running metric causally related to the expected value of the service metric determined based on the third-layer causal model.
[0160] According to one or more embodiments of the present disclosure, the first-layer causal model includes a potential causal RCM model; correspondingly, the model generation unit generates the first-layer causal model corresponding to the cloud distributed cache database according to the intervention effect of the test parameter of the configuration item and the test value of at least one running metric to be verified in each verification data, including: determining the intervention effect of the test parameter of the configuration item on the test value of the at least one running metric to be verified for the test parameter of the configuration item and the test value of at least one running metric to be verified in each verification data; and determining the corresponding relationship between at least one configuration item, the adjustment parameter of each configuration item, at least one running metric, and the value range of each running metric according to the intervention effect to generate the potential causal model corresponding to the cloud distributed cache database.
[0161] According to one or more embodiments of the present disclosure, the second - layer causal model includes a Bayesian network causal (CBN) model; correspondingly, the model generation unit generates the second - layer causal model corresponding to the cloud - edge distributed cache database according to the dependency relationship and subordination relationship among the multiple operation metrics, including: determining the dependency relationship and subordination relationship among the multiple operation metrics corresponding to a target event according to application performance management and configuration management databases, where the target event includes a scheduling event or an upgrade event; generating the Bayesian network causal model corresponding to the cloud - edge distributed cache database according to the dependency relationship, subordination relationship and conditional probability table among the multiple operation metrics corresponding to the target event.
[0162] According to one or more embodiments of the present disclosure, the third - layer causal model includes a structural equation (SEM) model; correspondingly, the model generation unit generates the third - layer causal model corresponding to the cloud - edge distributed cache database according to the multiple service metrics, the multiple operation metrics and a preset functional relationship, including: for each service metric, constructing a first functional relationship between the value of the service metric and the value range of at least one first target operation metric; determining at least one second target operation metric linearly related to the first target operation metric from the multiple operation metrics, and determining a second functional relationship between the first target operation metric and the second target operation metric; establishing an association relationship between the value range of the at least one second target operation metric and the value of the service metric according to the first functional relationship and the second functional relationship, and generating the structural equation model corresponding to the cloud - edge distributed cache database.
[0163] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;
[0164] The memory stores computer - executable instructions;
[0165] The at least one processor executes the computer - executable instructions stored in the memory, so that the at least one processor executes the cloud - edge distributed cache database adaptive tuning method as described in the first aspect and various possible designs of the first aspect above.
[0166] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer - readable storage medium, in which computer - executable instructions are stored, and when the processor executes the computer - executable instructions, the cloud - edge distributed cache database adaptive tuning method as described in the first aspect and various possible designs of the first aspect above is implemented.
[0167] Fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the cloud-based distributed cache database adaptive tuning method as described in the first aspect above and various possible designs of the first aspect.
[0168] The above description is only for the preferred embodiments of the present disclosure and the illustration of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0169] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0170] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An adaptive tuning method for an inter-cloud distributed cache database, characterized in that Including: In response to a tuning request for an inter-cloud distributed cache database, obtaining the expected values of at least one service metric carried in the tuning request; According to the expected values of the at least one service metric and the network causal model corresponding to the inter-cloud distributed cache database, determining a tuning strategy corresponding to the inter-cloud distributed cache database, the tuning strategy including at least one target configuration item and adjustment parameters for each target configuration item, where the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected values of the service metric based on the causal relationship from the configuration item to the service metric variable constructed based on various types of data analysis; According to the tuning strategy corresponding to the inter-cloud distributed cache database, adjusting the configuration parameters of each target configuration item of the inter-cloud distributed cache database to their respective corresponding adjustment parameters; Among them, before the step of in response to a tuning request for an inter-cloud distributed cache database, obtaining the expected values of at least one service metric carried in the tuning request, further including: obtaining a monitoring metric data set of the inter-cloud distributed cache database, the monitoring metric data set including verification data participating in the system test of the inter-cloud distributed cache database and cloud observation data and empirical data during the operation of the inter-cloud distributed cache database, the cloud observation data including multiple operation metrics, and the empirical data including multiple service metrics; generating a first-layer causal model corresponding to the inter-cloud distributed cache database according to the intervention effect between the test parameters of the configuration item in each verification data and the test values of at least one operation metric to be verified, and generating a second-layer causal model corresponding to the inter-cloud distributed cache database according to the dependence relationship and subordination relationship between the multiple operation metrics, and generating a third-layer causal model corresponding to the inter-cloud distributed cache database according to the multiple service metrics, the multiple operation metrics and a preset function relationship; fusing the first-layer causal model, the second-layer causal model and the third-layer causal model to generate the network causal model.
2. The method according to claim 1, characterized in that The network causal model includes a multi-layer heterogeneous causal model constructed based on various types of data, where different causal models are constructed and generated based on different types of data; among them, the various types of data include multiple types such as verification data, cloud observation data, and empirical data.
3. The method according to claim 1, characterized in that The network causal model includes a three-layer heterogeneous causal model constructed based on various types of data; among them, the first-layer causal model of the three-layer heterogeneous causal model is constructed based on verification data among various types of data and is used to determine the causal relationship between the configuration item and the operation metric, the second-layer causal model of the three-layer heterogeneous causal model is constructed based on cloud observation data among various types of data and is used to determine the causal relationship between the operation metrics, and the third-layer causal model of the three-layer heterogeneous causal model is constructed based on cloud observation data and empirical data among various types of data and is used to determine the causal relationship between the operation metric and the service metric.
4. The method according to any one of claims 1 to 3, characterized in that, Determining the tuning strategy corresponding to the cloud-edge distributed cache database according to the expected value of the at least one service metric and the network causal model corresponding to the cloud-edge distributed cache database includes: According to the calling order of the multi-layer heterogeneous causal models in the network causal model from the bottom layer to the top layer, call each layer of causal models in turn, use the output of the current layer of causal model as the input of the upper layer of causal model connected to it, and obtain at least one target configuration item output by the top layer of causal model and the adjustment parameter of each target configuration item as the tuning strategy corresponding to the cloud-edge distributed cache database.
5. The method according to any one of claims 1 to 3, characterized in that, The network causal model includes a three-layer heterogeneous causal model with network communication layer by layer from the bottom layer to the top layer. Among them, the first layer of causal model includes a latent causal model, the second layer of causal model includes a Bayesian network causal model, and the third layer of causal model includes a structural equation model.
6. The method according to claim 1, characterized in that, After adjusting the configuration parameters of each target configuration item of the cloud-edge distributed cache database to their respective corresponding adjustment parameters according to the tuning strategy corresponding to the cloud-edge distributed cache database, it further includes: If the empirical data corresponding to the service metric does not meet the expected value of the service metric, update the monitoring metric dataset; Generate a new network causal model through the updated monitoring metric dataset.
7. The method according to claim 1, characterized in that Fusing the first layer of causal model, the second layer of causal model and the third layer of causal model to generate the network causal model includes: Fuse the first layer of causal model and the second layer of causal model with multiple third running metrics as the combination points, and fuse the second layer of causal model and the third layer of causal model with multiple fourth running metrics as the combination points to generate the network causal model; Among them, the third running metric is a running metric that is causally related to the test parameter of the configuration item determined based on the first layer of causal model; the fourth running metric is a running metric that is causally related to the expected value of the service metric determined based on the third layer of causal model.
8. The method according to claim 1, wherein The first layer of causal model includes a latent causal model; Correspondingly, generating the first layer of causal model corresponding to the cloud-edge distributed cache database according to the intervention effect of the test parameter of the configuration item in each verification data and the test value of at least one running metric to be verified includes: For the test parameter of the configuration item and the test value of at least one running metric to be verified in each verification data, determine the intervention effect of the test parameter of the configuration item on the test value of the at least one running metric to be verified; According to the intervention effect, determine the corresponding relationship between at least one configuration item, the adjustment parameter of each configuration item, at least one running metric, and the value range of each running metric, and generate the latent causal model corresponding to the cloud-edge distributed cache database.
9. The method according to claim 1, wherein The second layer of causal model includes a Bayesian network causal model; Correspondingly, generating the second layer of causal model corresponding to the cloud-edge distributed cache database according to the dependency relationship and subordination relationship between the multiple running metrics includes: Determine the dependency relationship and subordination relationship among multiple running metrics corresponding to the target event according to the application performance management and the configuration management database, where the target event includes a scheduling event or an escalation event; Generate a Bayesian network causal model corresponding to the cloud distributed cache database according to the dependency relationship, subordination relationship, and conditional probability table among the multiple running metrics corresponding to the target event.
10. The method according to claim 1, characterized in that, The third-layer causal model includes a structural equation model; Correspondingly, the generating the third-layer causal model corresponding to the cloud distributed cache database according to the multiple business metrics, the multiple running metrics, and the preset functional relationship includes: For each business metric, construct a first functional relationship between the value of the business metric and the value range of at least one first target running metric; Determine at least one second target running metric that is linearly related to the first target running metric from the multiple running metrics, and determine a second functional relationship between the first target running metric and the second target running metric; According to the first functional relationship and the second functional relationship, establish an association relationship between the value range of the at least one second target running metric and the value of the business metric, and generate a structural equation model corresponding to the cloud distributed cache database.
11. A cloud-based distributed cache database adaptive tuning system, characterized in that, Include: An acquisition unit, configured to acquire the expected value of at least one business metric carried in the tuning request in response to a tuning request for the cloud distributed cache database; A determination unit, configured to determine a tuning strategy corresponding to the cloud distributed cache database according to the expected value of the at least one business metric and the network causal model corresponding to the cloud distributed cache database, where the tuning strategy includes at least one target configuration item and an adjustment parameter for each target configuration item, and the network causal model is used to predict and analyze the configuration parameters of the configuration item for the expected value of the business metric based on the causal relationship from the configuration item to the business metric variable constructed by analyzing various types of data; An adjustment unit, configured to adjust the configuration parameters of each target configuration item of the cloud distributed cache database to their respective corresponding adjustment parameters according to the tuning strategy corresponding to the cloud distributed cache database; Among them, the cloud distributed cache database adaptive tuning system further includes: a generation unit; the generation unit is configured to obtain a monitoring metric dataset of the cloud distributed cache database, where the monitoring metric dataset includes verification data participating in the cloud distributed cache database system test, as well as cloud observation data and empirical data during the operation of the cloud distributed cache database. The cloud observation data includes multiple operation metrics, and the empirical data includes multiple business metrics; generate a first-layer causal model corresponding to the cloud distributed cache database according to the intervention effect of the test parameters of the configuration items in each verification data and the test values of at least one operation metric to be verified, and, generate a second-layer causal model corresponding to the cloud distributed cache database according to the dependency relationship and subordination relationship among the multiple operation metrics, and, generate a third-layer causal model corresponding to the cloud distributed cache database according to the multiple business metrics, the multiple operation metrics, and a preset function relationship; fuse the first-layer causal model, the second-layer causal model, and the third-layer causal model to generate the network causal model.
12. An electronic device, characterized in that, including: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the cloud distributed cache database adaptive tuning method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the cloud distributed cache database adaptive tuning method according to any one of claims 1 to 10 is implemented.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the cloud distributed cache database adaptive tuning method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Database parameter adjusting method and device
CN114546990A
Parameter optimization method and device for production process
CN116050607A