Baseline threshold adjustment method and device, electronic equipment and storage medium
By dynamically evaluating the resource metric weights of database nodes and adjusting the baseline thresholds, the problem of fixed thresholds in load balancing strategies in distributed database systems is solved, improving system stability and resource utilization, and optimizing user experience.
Patent Information
- Application Number
- CN202511158684.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-21
AI Technical Summary
In existing distributed database systems, load balancing strategies use a fixed threshold mechanism, which may lead to performance degradation or service interruption under sudden load or resource contention scenarios, and cannot dynamically adapt to load changes.
By periodically acquiring operational data from database nodes, the weights of resource metrics are dynamically evaluated, and baseline thresholds are adjusted to adapt to load changes, ensuring the effective triggering of load balancing strategies and resource utilization.
This improved the stability and resource utilization of the database service, optimized the user experience for end users, and avoided issues such as node overload or unresponsiveness.
Smart Images

Figure CN120994393A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technology, and in particular to a baseline threshold adjustment method, apparatus, electronic device, and storage medium. Background Technology
[0002] In distributed database systems, load balancing strategies play a crucial role in ensuring service stability. Current mainstream solutions employ fixed threshold triggering mechanisms. However, due to the lack of dynamic adaptability in threshold settings, under sudden loads or resource contention scenarios, the threshold may be too high to trigger the load balancing strategy, leading to actual database node overload, performance degradation, or even service interruption. Conversely, if the threshold is too low, frequent false triggers may occur, resulting in policy suppression and unresponsiveness under real high loads, severely impacting user experience. Summary of the Invention
[0003] In view of this, embodiments of this application provide a baseline threshold adjustment method, apparatus, electronic device, and storage medium, which can effectively alleviate the above-mentioned technical problems.
[0004] In a first aspect, embodiments of this application provide a baseline threshold adjustment method, the method comprising:
[0005] Periodically acquire the first running data of the database node under test when it is running under non-steady-state test load;
[0006] The execution status of the load balancing strategy of the data node under test is determined based on the first running data.
[0007] In the event that the strategy execution status fails, a first resource weight is determined for each resource indicator in the data node under test based on the first running data; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of impact of the resource indicators in the database node under test on the overall performance bottleneck;
[0008] Adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
[0009] Optionally, as described above, determining the policy execution status of the load balancing strategy for the data node under test based on the first running data includes:
[0010] Obtain the operation log data from the first operation data;
[0011] The load balancing conditions are verified based on the runtime log data; wherein, the load balancing conditions include at least one of the following: post-migration status, data migration integrity, connection redirection correctness, and transaction interruption rate.
[0012] If the load balancing condition verification is passed, the load balancing strategy execution status of the data node under test is determined to be successful.
[0013] If the load balancing condition verification fails, the triggering status of the load balancing strategy of the data node under test is determined to be failed.
[0014] Optionally, as described above, determining the first resource weight of each resource indicator in the data node to be tested based on the first operational data includes:
[0015] Obtain the total number of SLA requests and the number of SLA violations for each of the resource metrics from the first running data;
[0016] For each of the resource metrics, an SLA default rate is determined based on the number of SLA violations and the total number of SLA requests for the resource metrics, and a first resource weight for the resource metrics is determined based on the SLA default rate.
[0017] Optionally, as described above, determining the first resource weight of the resource indicator based on the SLA default rate includes:
[0018] Obtain the second resource weight and the smoothing coefficient; wherein, the second resource weight is the resource weight of the resource indicator determined in the previous period;
[0019] The first resource weight of the resource indicator is determined based on the second resource weight, the smoothing coefficient, and the SLA default rate.
[0020] Optionally, as described above, adjusting the baseline threshold corresponding to the first resource indicator with the highest first resource weight includes:
[0021] Obtain the forgetting factor, the first resource data corresponding to the first resource indicator, and the first baseline threshold; wherein, the first baseline threshold is the baseline threshold of the first resource indicator determined in the previous period;
[0022] A second baseline threshold is determined based on the forgetting factor, the first resource data, and the first baseline threshold; wherein, the second baseline threshold is a baseline threshold adjusted from the first baseline threshold;
[0023] Obtain the second resource data, third baseline threshold, and first resource weight corresponding to each second resource indicator; wherein, the second resource indicator is not the first resource indicator;
[0024] The isolation baseline score is determined based on the second resource data, third baseline threshold and first resource weight corresponding to each of the second resource indicators, and the first resource data, second baseline threshold and first resource weight corresponding to the first resource indicator.
[0025] Detect whether the isolation baseline score reaches a preset scoring threshold;
[0026] If the isolation baseline score is detected to reach the preset score threshold, the second baseline threshold is determined as the adjusted baseline threshold of the first resource indicator;
[0027] If the isolation baseline score is found to be below the preset score threshold, the second baseline threshold is adjusted until the isolation baseline score reaches the preset score threshold.
[0028] Optionally, as described above, before periodically acquiring the first running data of the database node under test during non-steady-state test load operation, the method further includes:
[0029] Load steady-state test loads from the pre-set load template library;
[0030] The second running data of the database node under test is periodically acquired during the steady-state test load operation.
[0031] The second running data is input into a pre-trained weight model, and the weight model outputs the third resource weight and load type corresponding to each resource indicator in the database node under test;
[0032] Based on the load type, the priority resource indicator is determined from among the multiple resource indicators;
[0033] The load parameters of the corresponding resource indicators in the steady-state test load are adjusted upward based on the third resource weight corresponding to the emphasized resource indicators to obtain the unsteady-state test load.
[0034] Optionally, as described above, the method further includes:
[0035] Query the node faults corresponding to the load type from the fault injection database; wherein the fault injection database pre-stores the correspondence between load types and node faults;
[0036] The node faults are injected into the database node under test, forming a multi-dimensional composite test scenario with the non-steady-state test load.
[0037] Secondly, embodiments of this application provide a baseline threshold adjustment device, characterized in that the device comprises:
[0038] The acquisition module is used to periodically acquire the first running data of the database node under test when it is running under non-steady-state test load;
[0039] The first determining module is used to determine the triggering state of the load balancing strategy of the data node under test based on the first running data.
[0040] The second determining module is used to determine the first resource weight of each resource indicator in the data node under test based on the first running data when the triggering state is an untriggered state; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of impact of the resource indicators in the database node under test on the overall performance bottleneck;
[0041] The adjustment module is used to adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
[0042] Thirdly, embodiments of this application provide an electronic device, comprising: a processor and a memory, wherein the processor is configured to execute a baseline threshold adjustment program stored in the memory to implement the baseline threshold adjustment method described above.
[0043] Fourthly, embodiments of this application provide a storage medium storing one or more programs that can be executed by one or more processors to implement the baseline threshold adjustment method described above.
[0044] The baseline threshold adjustment method, apparatus, electronic device, and storage medium provided in this application include: periodically acquiring first operating data of a database node under test during non-steady-state test load operation; determining the policy execution status of the load balancing strategy of the data node under test based on the first operating data; in the case of policy execution failure, determining the first resource weight of each resource indicator in the data node under test based on the first operating data; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of influence of resource indicators in the database node under test on the overall performance bottleneck; and adjusting the baseline threshold corresponding to the first resource indicator with the highest first resource weight. This technical solution continuously monitors the real-time operating data of the database node, dynamically analyzes the influence weight of each resource indicator on system performance, accurately locates the most critical bottleneck resource, and adjusts the baseline threshold corresponding to the high-weight resource, ensuring that the baseline threshold setting is always synchronized with the actual load demand. This dynamic adaptation mechanism effectively solves the problem of node overload or unresponsiveness that may be caused by traditional fixed threshold schemes under sudden loads, significantly improves the stability and resource utilization of database services, and ultimately optimizes the user experience for end users. Attached Figure Description
[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 A flowchart illustrating an embodiment of a baseline threshold adjustment method provided in this application;
[0048] Figure 2 A flowchart illustrating an embodiment of another baseline threshold adjustment method provided in this application;
[0049] Figure 3 A flowchart illustrating an embodiment of another baseline threshold adjustment method provided in this application;
[0050] Figure 4 A flowchart illustrating an embodiment of another baseline threshold adjustment method provided in this application;
[0051] Figure 5 A flowchart illustrating an embodiment of another baseline threshold adjustment method provided in this application;
[0052] Figure 6 A block diagram illustrating an embodiment of a baseline threshold adjustment device provided in this application;
[0053] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0057] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0058] Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0059] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0060] To facilitate understanding of the embodiments of this application, the following will provide further explanation and description with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.
[0061] This application provides a baseline threshold adjustment method. See also... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of a baseline threshold adjustment method provided in this application. Figure 1 The process shown may include the following steps:
[0062] Step 101: Periodically acquire the first running data of the database node under test when it is running under non-steady-state test load;
[0063] A monitoring mechanism is periodically triggered by a timer to collect the first running data of the database node under test (including resource indicators such as CPU (Central Processing Unit), memory, and I / O (Input / Output)) under non-steady-state test loads such as simulated burst traffic or resource contention in real time (e.g., every 5 seconds). Since non-steady-state test loads can simulate the fluctuating loads in real business scenarios, this first running data can reflect the resource fluctuation characteristics under real business pressure, providing a data basis for subsequent dynamic adjustment of baseline thresholds.
[0064] Step 102: Determine the policy execution status of the load balancing strategy for the data node under test based on the first running data;
[0065] The core purpose of step 102 above, determining the strategy execution status, is to verify whether the load balancing strategy of the database node under test is correctly activated under non-steady-state test load. When the strategy execution status shows success, it indicates that the current baseline threshold setting is reasonable and the load balancing strategy has been successfully invoked (such as automatic load migration between nodes), and there is no need to adjust the threshold parameters at this time. When the strategy execution status shows failure, it indicates that the existing baseline threshold has failed to effectively identify the actual performance bottleneck (such as a surge in query latency but resource indicators not reaching the threshold), and at this time, it is necessary to start the dynamic threshold adjustment mechanism to optimize and calibrate the baseline threshold of resource indicators.
[0066] Step 103: If the strategy execution status is failed, determine the first resource weight of each resource indicator in the data node to be tested based on the first running data.
[0067] The aforementioned first resource weight is a quantitative parameter used to dynamically assess the impact of resource metrics in the tested database node on the overall performance bottleneck. It can be understood as a parameter that dynamically assigns values to resource metrics such as CPU, memory, and I / O through quantitative analysis, accurately reflecting the contribution of each resource metric to the current performance bottleneck. A higher first resource weight indicates a more significant impact of that resource metric on the current performance bottleneck. For example, a first resource weight of 90% for CPU indicates that CPU contention is the primary factor causing query latency.
[0068] As described above, the first resource weight of each resource indicator can directly reflect the reason for the failure of the load balancing strategy. For example, when the first resource weight corresponding to I / O is 0.8 while the first resource weight corresponding to CPU is only 0.2, it indicates that the existing threshold does not cover the storage performance bottleneck, resulting in the strategy not being activated. In the subsequent baseline threshold adjustment, the determination of the first resource weight is the core basis, and its role is reflected in the following aspects: First, the order of the first resource weight determines the priority of adjusting the baseline threshold of each resource indicator (such as resource indicators with a weight > 0.7 triggering the recalculation of the baseline threshold first). Second, the first resource weight is related to the baseline adjustment (which will be detailed later).
[0069] Step 104: Adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
[0070] When the load balancing strategy fails, the adjustment direction of the baseline threshold of the resource metric with the highest first resource weight needs to be dynamically determined based on the performance bottleneck type and optimization objective.
[0071] Specifically, when the first resource indicator corresponding to the highest first resource weight is an indicator that represents resource overload, such as CPU or memory, and the current non-steady-state test load is continuously close to but has not reached the baseline threshold (e.g., CPU is at 85% for a long time, and the baseline threshold is 90%), it is necessary to adjust the baseline threshold upward (e.g., from 90% to 92%) to allow the database node under test to bear a higher load and avoid unnecessary balancing actions.
[0072] When the primary resource indicator corresponding to the highest primary resource weight is an indicator representing resource idleness such as disk idle rate or network bandwidth remaining, and the current resource utilization rate is consistently lower than the baseline threshold (e.g., I / O utilization rate is only 30%, while the baseline threshold is 40%), if the baseline threshold is too high, it will be impossible to detect resource idleness, leading to concentrated load. In this case, it is necessary to adjust the threshold downward (e.g., from 40% to 35%) to prompt the load balancing strategy to intervene earlier and improve resource utilization.
[0073] In practical applications, when the first resource indicator corresponding to the highest first resource weight is adjusted, other resource indicators (such as I / O) may become new bottlenecks due to load changes or resource competition. They can be promoted to the new highest first resource weight to trigger the baseline threshold adjustment process.
[0074] The method described in this embodiment continuously monitors the real-time operational data of database nodes, dynamically analyzes the impact weight of various resource indicators on system performance, accurately identifies the most critical bottleneck resources, and adjusts the baseline thresholds corresponding to high-weight resources accordingly, ensuring that the baseline threshold settings are always synchronized with actual load demands. This dynamic adaptation mechanism effectively solves the problem of node overload or unresponsiveness that may occur under sudden loads in traditional fixed-threshold solutions, significantly improving the stability and resource utilization of database services, and ultimately optimizing the user experience for end users.
[0075] like Figure 2 As shown, as an optional implementation, the method described above, step 102 of determining the load balancing strategy execution status of the data node under test based on the first running data includes the following steps:
[0076] Step 201: Obtain runtime log data from the first runtime data;
[0077] To verify the following equilibrium conditions, the following key runtime log data needs to be extracted from the first runtime data: resource contention log (including lock waiting and thread pool status), metadata verification log (including table structure change records), session sticky log (including IP (Internet Protocol) mapping and TCP (Transmission Control Protocol) state transitions), and distributed transaction log (including 2PC (Two-Phase Commit Protocol) phase markers and compensation records).
[0078] Step 202: Verify the load balancing conditions based on the runtime log data;
[0079] The aforementioned load balancing conditions include at least one of the following: post-migration status, data migration integrity, connection redirection correctness, and transaction interruption rate. Specifically, post-migration status checks whether the target node has successfully received data and is functioning normally (e.g., whether the Region Server is readable after migration); data migration integrity verifies whether data was lost or corrupted during the migration process (e.g., verifying hash values); connection redirection correctness confirms whether client requests are correctly routed to the new database node; and the transaction interruption rate is used to statistically analyze the proportion of failed transactions during the migration, which must be below a preset threshold (e.g., <0.1%). These three criteria—post-migration status, data migration integrity, connection redirection correctness, and transaction interruption rate—are all verification items used to measure the success of the load balancing strategy's execution. Failure in any one of these will result in the load balancing strategy's execution status being deemed a failure.
[0080] Step 203: If the load balancing condition verification is passed, determine that the load balancing strategy execution status of the data node under test is successful.
[0081] If the equilibrium condition verification is passed, it means that all the verification items required above have been successfully verified, and there is no need to carry out the process of determining resource weights and adjusting baseline thresholds.
[0082] Step 204: If the load balancing condition verification fails, determine that the load balancing strategy execution status of the data node under test is failed.
[0083] If the equilibrium condition verification fails, it indicates that at least one of the verification items required above has failed. In this case, the process of determining resource weights and adjusting baseline thresholds is required to eliminate the resource contention bottleneck.
[0084] like Figure 3 As shown, as an optional implementation, the method described above, step 103, which determines the first resource weight of each resource indicator in the data node to be tested based on the first operating data, includes the following steps:
[0085] Step 301: Obtain the total number of SLA requests and the number of SLA violations for each resource metric from the first running data;
[0086] The total number of SLA (Service Level Agreement) requests is the total number of all valid requests (including successful and failed requests) within the statistical period, excluding interfering data such as test traffic; while the number of SLA violations for each resource metric is the total number of all failed requests within the statistical period.
[0087] Step 302: For each resource indicator, determine the SLA default rate based on the number of SLA violations and the total number of SLA requests for the resource indicator, and determine the first resource weight of the resource indicator based on the SLA default rate.
[0088] Divide the number of SLA violations for each resource metric by the total number of SLA requests to obtain the SLA default rate for each resource metric. Example: If the CPU has 50 SLA violations / 100,000 total SLA requests, the default rate is 0.05%.
[0089] The process of determining the first resource weight of each resource indicator based on the SLA default rate is as follows: obtaining the second resource weight and the smoothing coefficient; wherein, the second resource weight is the resource weight of the resource indicator determined in the previous period; and determining the first resource weight of the resource indicator based on the second resource weight, the smoothing coefficient and the SLA default rate.
[0090] It should be noted that the smoothing coefficients for all the above resource indicators are consistent. Using the same smoothing coefficient can avoid comparability issues caused by differences in algorithms. However, the second resource weights for each resource indicator may vary due to differences in business scenarios or resource characteristics.
[0091] The first resource weight for each resource indicator can be calculated using the following formula:
[0092] W i (t)=aW i (t-1)+(1-a)×b;
[0093] Among them, W i (t) represents the first resource weight of resource index i, W i (t-1) represents the second resource weight of resource indicator i, a represents the smoothing coefficient (a∈[0,1], default 0.7), and b represents the SLA default rate.
[0094] The above method accurately reflects the actual default rate of each resource indicator by dividing the number of SLA violations by the total number of SLA requests, avoiding the misleading nature of absolute values. Resources with high default rates are automatically given higher weights, and the baseline threshold of the resource indicator with the highest resource weight is adjusted first, which can quickly eliminate resource competition bottlenecks and ensure database node services.
[0095] like Figure 4 As shown, as an optional implementation, the method described above, step 104 of adjusting the baseline threshold corresponding to the first resource indicator with the highest first resource weight includes the following steps:
[0096] Step 401: Obtain the forgetting factor, the first resource data corresponding to the first resource indicator, and the first baseline threshold;
[0097] Among them, the first baseline threshold is the baseline threshold of the first resource indicator determined in the previous period, that is, the historical threshold benchmark calculated by the same process in the previous period; the forgetting factor is used to control the influence weight of historical data on the current baseline threshold (value range 0-1, usually 0.9). The smaller the value, the faster the historical reference value decays. The first resource data is the original data of the first resource indicator (such as CPU) collected in the current time period.
[0098] Step 402: Determine a second baseline threshold based on the forgetting factor, the first resource data, and the first baseline threshold; wherein the second baseline threshold is a baseline threshold adjusted from the first baseline threshold.
[0099] The second baseline threshold can be calculated using the following formula:
[0100] E i (t)=c×Ei (t-1)+(1-c)×A i (t);
[0101] Among them, E i (t) represents the second baseline threshold, E i (t-1) represents the first baseline threshold, c represents the forgetting factor, and A i (t) represents the first resource data.
[0102] Step 403: Obtain the second resource data, third baseline threshold and first resource weight corresponding to each second resource indicator;
[0103] Among them, the second resource indicator is not the first resource indicator; the third baseline threshold can be understood as the historical threshold benchmark calculated by each second resource indicator in the previous cycle through the same process; the second resource data is the raw data of the second resource indicator (such as memory, I / O) collected in the current time period.
[0104] Step 404: Determine the isolation baseline score based on the second resource data, third baseline threshold and first resource weight corresponding to each second resource indicator, and the first resource data, second baseline threshold and first resource weight corresponding to the first resource indicator.
[0105] The isolation baseline score can be calculated using the following formula:
[0106]
[0107] Where score(t) represents the isolation baseline score, and n represents the total number of resource metrics. It should be noted that E i (t) represents the second baseline threshold for the first resource indicator, and the third baseline threshold for the second resource indicator. i (t) represents the first resource data for the first resource indicator and the second resource data for the second resource indicator.
[0108] Step 405: Check whether the isolation baseline score has reached the preset score threshold;
[0109] The preset scoring threshold is the critical value at which the representation isolation capability set according to the business SLA meets the SLA requirements. The higher the score, the better the resource isolation effect.
[0110] When the isolation baseline score is detected to have reached the preset score threshold, the current baseline threshold adjustment is deemed valid, and step 406 is executed directly to confirm the baseline threshold. If the isolation baseline score does not reach the preset score threshold, the adjustment is deemed unsatisfactory, and the iterative optimization process in step 407 is automatically triggered until the isolation baseline score meets the requirements. This design ensures the reliability of the baseline threshold adjustment through a closed-loop verification mechanism.
[0111] Step 406: Determine the second baseline threshold as the adjusted baseline threshold of the first resource indicator;
[0112] Step 407: Continue to adjust the second baseline threshold until the isolation baseline score reaches the preset score threshold.
[0113] This feedback loop, formed through the scoring threshold verification process, ensures that each baseline threshold adjustment meets the overall health requirements of the node, effectively avoiding over-adjustment.
[0114] In practical applications, the generation of the aforementioned unsteady-state test load can be found in [reference needed]. Figure 5 As shown, as an optional implementation, the following steps are included:
[0115] Step 501: Load the steady-state test load from the preset load template library;
[0116] The aforementioned pre-built load template library includes steady-state test loads for standard benchmark templates such as TPC-C (Transaction Processing Performance Council-C), TPC-H (Transaction Processing Performance Council-H), and YCSB (Yahoo! Cloud Serving Benchmark).
[0117] Step 502: Periodically acquire the second running data of the database node under test during steady-state test load operation;
[0118] Similarly, the monitoring mechanism is periodically triggered by a timer. The second set of running data collected in real time (e.g., every 5 seconds) includes, but is not limited to, data such as QPS (Queries Per Second), slow queries, lock waits, and SQL (Structured Query Language) characteristics collected using Prometheus + custom exporter; CPU utilization and CGroup (Control Groups) limits collected using cAdvisor + node exporter; and bandwidth and packet loss rate collected using eBPF (extended Berkeley Packet Filter) traffic analysis.
[0119] Step 503: Input the second running data into the pre-trained weight model. The weight model outputs the third resource weight and load type corresponding to each resource indicator in the database node to be tested.
[0120] The definition of the third resource weight is the same as that of the first resource weight, and will not be repeated here. The process for determining the third resource weight can also refer to the process for determining the first resource weight in the above embodiments. Similarly, the process for determining the first resource weight can be determined by a weight model, which is not limited here. The weight model is obtained by training a machine learning model using training samples of historical resource weights and load types.
[0121] Step 504: Determine the priority resource indicator from multiple resource indicators based on the load type;
[0122] In practical applications, different workload types emphasize different resource metrics. For example, when the workload type is OLTP (Online Transaction Processing), the focus is on I / O, so I / O resource metrics are determined as the key resource metrics; when the workload type is OLAP (Online Analytical Processing), the focus is on memory, so memory resource metrics are determined as the key resource metrics.
[0123] Step 505: Adjust the load parameters of the corresponding resource indicators in the steady-state test load upward based on the third resource weight corresponding to the resource indicator to obtain the unsteady-state test load.
[0124] In steady-state test loads, resource bottlenecks can be artificially created by selectively increasing the load parameters of specific resource metrics (such as I / O, CPU, or memory) to simulate sudden traffic spikes or resource contention scenarios caused by business fluctuations in a production environment. This type of non-steady-state test load can effectively verify whether the system's load balancing strategy is triggered correctly during resource contention.
[0125] Specific example: If the test plan sets the I / O resource weight to 70%, and the current steady-state load I / O parameter is 1500 IOPS (Input / Output Operations Per Second), then the I / O load parameter needs to be adjusted proportionally until its actual resource consumption reaches the preset 70% weight target. This process exposes the system's coordination capability under critical conditions through stress testing.
[0126] In practical applications, node faults corresponding to the load type can be injected into the nodes of the database under test to increase the complexity of the test scenario. Specifically, node faults corresponding to the load type are queried from the fault injection database; the node faults are injected into the nodes of the database under test to form a multi-dimensional composite test scenario with the non-steady-state test load.
[0127] The fault injection library pre-stores the correspondence between load types and node faults. This means that different load types correspond to different node faults. For example, the node fault corresponding to the OLTP load type is disk bad sectors, while the node fault corresponding to the OLAP load type is transparent large page splits. It should be noted that the above-listed correspondence between load types and node faults is only an example. The specific correspondence between load types and node faults can be set according to actual needs and is not limited here.
[0128] This testing solution employs a collaborative implementation mechanism combining standard load models and chaos engineering methods. Through predefined load patterns and systemic fault injection, it accurately reproduces typical failure scenarios caused by traffic surges and resource contention in the production environment. The implementation goal of this method focuses on establishing a multi-dimensional baseline threshold system by dynamically adjusting the load from steady-state to non-steady-state test loads. This provides a scientific reference for operations personnel, enabling the operations team to more accurately set resource thresholds and optimize node performance monitoring and alarm strategies.
[0129] See Figure 6 This is a block diagram illustrating an embodiment of a baseline threshold adjustment device provided in this application. Figure 6 As shown, the device includes:
[0130] The acquisition module 600 is used to periodically acquire the first running data of the database node under test when it is running under non-steady-state test load.
[0131] The first determining module 601 is used to determine the triggering state of the load balancing strategy of the data node under test based on the first running data.
[0132] The second determining module 602 is used to determine the first resource weight of each resource indicator in the data node under test based on the first running data when the triggering state is an untriggered state; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of influence of the resource indicators in the database node under test on the overall performance bottleneck;
[0133] The adjustment module 603 is used to adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
[0134] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0135] As an optional implementation, the first determining module 601 is further configured to:
[0136] Obtain the operation log data from the first operation data;
[0137] The load balancing conditions are verified based on the runtime log data; wherein, the load balancing conditions include at least one of the following: post-migration status, data migration integrity, connection redirection correctness, and transaction interruption rate.
[0138] If the load balancing condition verification is passed, the load balancing strategy execution status of the data node under test is determined to be successful.
[0139] If the load balancing condition verification fails, the load balancing strategy execution status of the data node under test is determined to be failed.
[0140] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0141] As an optional implementation, the second determining module 602 further includes:
[0142] The data acquisition module is used to obtain the total number of SLA requests and the number of SLA violations for each of the resource metrics from the first running data;
[0143] The weight determination module is used to determine the SLA default rate for each resource indicator based on the number of SLA violations and the total number of SLA requests, and to determine the first resource weight of the resource indicator based on the SLA default rate.
[0144] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0145] As an optional implementation, the weight determination module described above is also used for:
[0146] Obtain the second resource weight and the smoothing coefficient; wherein, the second resource weight is the resource weight of the resource indicator determined in the previous period;
[0147] The first resource weight of the resource indicator is determined based on the second resource weight, the smoothing coefficient, and the SLA default rate.
[0148] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0149] As an optional implementation, the adjustment module 603 is further configured to:
[0150] Obtain the forgetting factor, the first resource data corresponding to the first resource indicator, and the first baseline threshold; wherein, the first baseline threshold is the baseline threshold of the first resource indicator determined in the previous period;
[0151] A second baseline threshold is determined based on the forgetting factor, the first resource data, and the first baseline threshold; wherein, the second baseline threshold is a baseline threshold adjusted from the first baseline threshold;
[0152] Obtain the second resource data, third baseline threshold, and first resource weight corresponding to each second resource indicator; wherein, the second resource indicator is not the first resource indicator;
[0153] The isolation baseline score is determined based on the second resource data, third baseline threshold and first resource weight corresponding to each of the second resource indicators, and the first resource data, second baseline threshold and first resource weight corresponding to the first resource indicator.
[0154] Detect whether the isolation baseline score reaches a preset scoring threshold;
[0155] If the isolation baseline score is detected to reach the preset score threshold, the second baseline threshold is determined as the adjusted baseline threshold of the first resource indicator;
[0156] If the isolation baseline score is found to be below the preset score threshold, the second baseline threshold is adjusted until the isolation baseline score reaches the preset score threshold.
[0157] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0158] As an optional implementation, the above-described apparatus further includes:
[0159] The loading module is used to load steady-state test loads from a pre-built load template library;
[0160] The timed acquisition module is used to periodically acquire the second running data of the database node under test during the steady-state test load.
[0161] The model output module is used to input the second running data into the pre-trained weight model, and the weight model outputs the third resource weight and load type corresponding to each resource indicator in the database node to be tested;
[0162] The third determining module is used to determine the priority resource indicator from multiple resource indicators based on the load type.
[0163] The adjustment module is used to adjust the load parameters of the corresponding resource indicators in the steady-state test load based on the third resource weight corresponding to the emphasized resource indicator, so as to obtain the unsteady-state test load.
[0164] Specifically, the detailed process by which each module in the device of this invention implements its function can be found in the relevant description in the method embodiment, and will not be repeated here.
[0165] As an optional implementation, the above-described apparatus further includes:
[0166] The query module is used to query the node faults corresponding to the load type from the fault injection database; wherein the fault injection database stores the correspondence between load types and node faults in advance.
[0167] The injection module is used to inject the node fault into the database node under test, forming a multi-dimensional composite test scenario with the non-steady-state test load.
[0168] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 The illustrated electronic device 1200 includes at least one processor 1201, a memory 1202, at least one network interface 1204, and other user interfaces 1203. The various components in the electronic device 1200 are coupled together via a bus system 1205. It is understood that the bus system 1205 is used to implement communication between these components. In addition to a data bus, the bus system 1205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 7 The general labeled all buses as Bus System 1205.
[0169] The user interface 1203 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).
[0170] It is understood that the memory 1202 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 1202 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0171] In some implementations, memory 1202 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 12021 and application program 12022.
[0172] The operating system 12021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 12022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this application embodiment can be included in the application program 12022.
[0173] In this embodiment of the application, the processor 1201 executes the method steps provided by each method embodiment by calling the program or instructions stored in the memory 1202, specifically the program or instructions stored in the application program 12022.
[0174] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1201. The processor 1201 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the processor 1201. The processor 1201 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1202. Processor 1201 reads the information in memory 1202 and completes the steps of the above method in conjunction with its hardware.
[0175] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.
[0176] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0177] The electronic device provided in this embodiment may be as follows: Figure 7 The electronic device shown can perform the following: Figure 1-5 All steps of the intermediate-generation baseline threshold adjustment method, thereby achieving Figure 1-5 For details on the technical effects of the baseline threshold adjustment method shown, please refer to [link / reference]. Figure 1-5 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0178] This application also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.
[0179] One or more programs in the storage medium can be executed by one or more processors to implement the baseline threshold adjustment method described above.
[0180] The processor is used to execute a baseline threshold adjustment program stored in memory to implement the steps of the baseline threshold adjustment method.
[0181] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0182] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0183] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A baseline threshold adjustment method, characterized in that, The method includes: Periodically acquire the first running data of the database node under test when it is running under non-steady-state test load; The execution status of the load balancing strategy of the data node under test is determined based on the first running data. In the event that the strategy execution status fails, a first resource weight is determined for each resource indicator in the data node under test based on the first running data; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of impact of the resource indicators in the database node under test on the overall performance bottleneck; Adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
2. The method according to claim 1, characterized in that, The step of determining the load balancing strategy execution status of the data node under test based on the first running data includes: Obtain the operation log data from the first operation data; The load balancing conditions are verified based on the runtime log data; wherein, the load balancing conditions include at least one of the following: post-migration status, data migration integrity, connection redirection correctness, and transaction interruption rate. If the load balancing condition verification is passed, the load balancing strategy execution status of the data node under test is determined to be successful. If the load balancing condition verification fails, the load balancing strategy execution status of the data node under test is determined to be failed.
3. The method according to claim 1, characterized in that, The step of determining the first resource weight of each resource indicator in the data node to be tested based on the first operating data includes: Obtain the total number of SLA requests and the number of SLA violations for each of the resource metrics from the first running data; For each of the resource metrics, an SLA default rate is determined based on the number of SLA violations and the total number of SLA requests for the resource metrics, and a first resource weight for the resource metrics is determined based on the SLA default rate.
4. The method according to claim 3, characterized in that, The step of determining the first resource weight of the resource indicator based on the SLA default rate includes: Obtain the second resource weight and the smoothing coefficient; wherein, the second resource weight is the resource weight of the resource indicator determined in the previous period; The first resource weight of the resource indicator is determined based on the second resource weight, the smoothing coefficient, and the SLA default rate.
5. The method according to claim 1, characterized in that, The threshold adjustment of the baseline threshold corresponding to the first resource indicator with the highest first resource weight includes: Obtain the forgetting factor, the first resource data corresponding to the first resource indicator, and the first baseline threshold; wherein, the first baseline threshold is the baseline threshold of the first resource indicator determined in the previous period; A second baseline threshold is determined based on the forgetting factor, the first resource data, and the first baseline threshold; wherein, the second baseline threshold is a baseline threshold adjusted from the first baseline threshold; Obtain the second resource data, third baseline threshold, and first resource weight corresponding to each second resource indicator; wherein, the second resource indicator is not the first resource indicator; The isolation baseline score is determined based on the second resource data, third baseline threshold and first resource weight corresponding to each of the second resource indicators, and the first resource data, second baseline threshold and first resource weight corresponding to the first resource indicator. Detect whether the isolation baseline score reaches a preset scoring threshold; If the isolation baseline score is detected to reach the preset score threshold, the second baseline threshold is determined as the adjusted baseline threshold of the first resource indicator; If the isolation baseline score is found to be below the preset score threshold, the second baseline threshold is adjusted until the isolation baseline score reaches the preset score threshold.
6. The method according to claim 1, characterized in that, Before periodically acquiring the first running data of the database node under test during non-steady-state test load operation, the method further includes: Load steady-state test loads from the pre-set load template library; The second running data of the database node under test is periodically acquired during the steady-state test load operation. The second running data is input into a pre-trained weight model, and the weight model outputs the third resource weight and load type corresponding to each resource indicator in the database node under test; Based on the load type, the priority resource indicator is determined from among the multiple resource indicators; The load parameters of the corresponding resource indicators in the steady-state test load are adjusted upward based on the third resource weight corresponding to the emphasized resource indicators to obtain the unsteady-state test load.
7. The method according to claim 6, characterized in that, The method further includes: Query the node faults corresponding to the load type from the fault injection database; wherein the fault injection database pre-stores the correspondence between load types and node faults; The node faults are injected into the database node under test, forming a multi-dimensional composite test scenario with the non-steady-state test load.
8. A baseline threshold adjustment device, characterized in that, The device includes: The acquisition module is used to periodically acquire the first running data of the database node under test when it is running under non-steady-state test load; The first determining module is used to determine the triggering state of the load balancing strategy of the data node under test based on the first running data. The second determining module is used to determine the first resource weight of each resource indicator in the data node under test based on the first running data when the triggering state is an untriggered state; wherein, the first resource weight is a quantitative parameter used to dynamically evaluate the degree of impact of the resource indicators in the database node under test on the overall performance bottleneck; The adjustment module is used to adjust the baseline threshold corresponding to the first resource indicator with the highest first resource weight.
9. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a baseline threshold adjustment program stored in the memory to implement the baseline threshold adjustment method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the baseline threshold adjustment method according to any one of claims 1 to 7.