Pressure measurement method and device of distributed system, electronic equipment, medium and program product

By obtaining the business scenario tags and operating indicators of the distributed system, generating and dynamically adjusting the stress testing strategy, the problems of test distortion and resource waste in the existing technology are solved, and efficient and stable distributed system performance testing is achieved.

CN120653526APending Publication Date: 2025-09-16INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510783873.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing distributed system performance testing technologies lack awareness of real-time node load and resource heterogeneity, and are unable to simulate the actual distribution of user behavior and multi-service linkage, resulting in test distortion and resource waste. In addition, monitoring and feedback are delayed, making it impossible to identify anomalies in real time.

Method used

By obtaining the business scenario labels and current operating indicators of the target distributed system, a target stress testing strategy is generated based on the stress testing strategy template. The strategy is adjusted in real time during the stress testing task, and dynamic resource adjustment is performed using fuzzy logic and causal reasoning models. The test process is optimized by combining red-black deployment strategies and preheating and capacity expansion.

Benefits of technology

It improves the authenticity and accuracy of stress testing, reduces resource consumption, enhances system stability and the convenience of the testing process, reduces the operation and maintenance burden, and achieves rapid response to sudden loads and performance degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653526A_ABST
    Figure CN120653526A_ABST
Patent Text Reader

Abstract

The invention provides a pressure measurement method for a distributed system. The pressure measurement method can be applied to the technical field of distribution, the technical field of artificial intelligence and the technical field of financial science and technology. The method comprises the following steps: acquiring a service scene label and a current operation index of a target distributed system; obtaining a pressure measurement strategy template based on the business scene label, and generating a target pressure measurement strategy based on the pressure measurement strategy template and the current operation index; according to the target pressure measurement strategy, a pressure measurement task is executed on a target distributed system, and in the execution process of the pressure measurement task, a real-time operation index of the target distributed system is obtained; and dynamically adjusting the target pressure measurement strategy in response to the real-time operation index reaching a preset adjustment condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of distributed technology and artificial intelligence technology, and more specifically to a stress testing method, apparatus, device, medium, and program product for a distributed system. Background Art

[0002] To support business requirements such as high concurrency, high availability, and high consistency, financial institutions widely adopt distributed architectures to carry complex business processing flows. In distributed system performance testing technology, traditional solutions typically include steps such as environment deployment, script development, load generation, and offline result analysis, and are widely used to verify the stability of systems such as microservice architectures and database clusters. However, existing technologies generally rely on static configurations and fixed processes, which have significant shortcomings: First, current load distribution strategies lack awareness of the real-time load and resource heterogeneity of each node, which can easily lead to hotspot aggregation and test distortion; second, existing test script models are rigid and cannot simulate complex scenarios such as the distribution of real user behavior and multi-service linkage; in addition, current data analysis lags and cannot identify and respond to anomalies that occur during the test in real time, resulting in insufficient overall test intelligence. Summary of the Invention

[0003] In view of the above problems, the present disclosure provides a stress testing method, apparatus, device, medium and program product for a distributed system.

[0004] According to a first aspect of the present disclosure, a stress testing method for a distributed system is provided, the method comprising: obtaining a business scenario label and a current operating indicator of a target distributed system; obtaining a stress testing policy template based on the business scenario label, and generating a target stress testing policy based on the stress testing policy template and the current operating indicator; and performing a stress testing task on the target distributed system according to the target stress testing policy, wherein, during the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained; and in response to the real-time operating indicators reaching a preset adjustment condition, the target stress testing policy is dynamically adjusted.

[0005] According to an embodiment of the present disclosure, obtaining a stress testing policy template based on the business scenario label and generating a target stress testing policy based on the stress testing policy template and the current operating indicators specifically include: based on the business scenario label, retrieving a stress testing policy template that matches the business scenario label in a preset stress testing policy template library; and fine-tuning the initial stress testing policy in the stress testing policy template based on the current operating indicators to obtain the target stress testing policy.

[0006] According to an embodiment of the present disclosure, the fine-tuning of the initial stress testing strategy in the stress testing strategy template based on the current operating indicators specifically includes: obtaining the initial stress testing strategy based on the initial preset parameters in the stress testing strategy template; adjusting the target parameters in the initial stress testing strategy based on the current operating indicators to obtain a dynamic stress testing strategy; and fusing the initial stress testing strategy and the dynamic stress testing strategy to form the target stress testing strategy.

[0007] According to an embodiment of the present disclosure, the dynamic adjustment of the target stress testing strategy specifically includes: in response to the real-time operation indicator reaching a preset adjustment condition, using fuzzy logic to judge the real-time operation indicator and obtain the resource urgency level; selecting a resource adjustment strategy based on the resource urgency level, and performing a resource adjustment operation based on the resource adjustment strategy.

[0008] According to an embodiment of the present disclosure, the resource adjustment strategy selected based on the resource urgency level specifically includes: in response to the resource urgency level being the first level, selecting a rapid capacity expansion operation; in response to the resource urgency level being the second level, selecting a gradual resource adjustment operation; and in response to the resource urgency level being the third level, selecting a preheating capacity expansion operation.

[0009] According to an embodiment of the present disclosure, in the process of performing resource adjustment operations based on the resource adjustment strategy, the red-black deployment strategy is used to switch the stress testing traffic; and / or a resource change cooling period is set so that the time interval between two dynamic adjustments is not less than a preset duration.

[0010] According to an embodiment of the present disclosure, the method also includes: before executing the resource adjustment operation based on the resource adjustment policy, predicting the system impact of this resource adjustment and obtaining an impact score; and in response to the impact score exceeding the target threshold, preventing the execution of this resource change operation.

[0011] According to an embodiment of the present disclosure, the method further includes: performing indicator anomaly detection on the real-time operation indicator based on a rule engine; and / or performing trend prediction on the real-time operation indicator based on a time series prediction model, and performing pattern anomaly detection based on the result of trend prediction.

[0012] According to an embodiment of the present disclosure, the method further includes: if indicator anomaly detection or pattern anomaly detection generates an abnormal result, obtaining a root cause analysis result based on a causal reasoning model, wherein the causal reasoning model includes a causal correlation relationship diagram between multiple operating indicators.

[0013] A second aspect of the present disclosure provides a stress testing device for a distributed system, the device comprising: a data acquisition module, configured to obtain business scenario labels and current operating indicators of a target distributed system; a target stress testing strategy generation module, configured to obtain a stress testing strategy template based on the business scenario labels, and generate a target stress testing strategy based on the stress testing strategy template and the current operating indicators; a dynamic adjustment module, configured to execute a stress testing task on the target distributed system according to the target stress testing strategy, wherein, during the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained; and in response to the real-time operating indicators reaching preset adjustment conditions, the target stress testing strategy is dynamically adjusted.

[0014] According to an embodiment of the present disclosure, the target stress testing strategy generation module can also be used to retrieve a stress testing strategy template that matches the business scenario label in a preset stress testing strategy template library based on the business scenario label; and fine-tune the initial stress testing strategy in the stress testing strategy template based on the current operating indicators to obtain the target stress testing strategy.

[0015] According to an embodiment of the present disclosure, the target stress testing strategy generation module can also be used to obtain an initial stress testing strategy based on the initial preset parameters in the stress testing strategy template; adjust the target parameters in the initial stress testing strategy based on the current operating indicators to obtain a dynamic stress testing strategy; and merge the initial stress testing strategy and the dynamic stress testing strategy to form the target stress testing strategy.

[0016] According to an embodiment of the present disclosure, the dynamic adjustment module can also be used to respond to the real-time operation indicator reaching a preset adjustment condition, use fuzzy logic to judge the real-time operation indicator, and obtain the resource urgency level; select a resource adjustment strategy based on the resource urgency level, and perform resource adjustment operations based on the resource adjustment strategy.

[0017] According to an embodiment of the present disclosure, the dynamic adjustment module can also be used to select a resource adjustment strategy based on the resource urgency level, specifically including: in response to the resource urgency level being the first level, selecting a rapid capacity expansion operation; in response to the resource urgency level being the second level, selecting a gradual resource adjustment operation; and in response to the resource urgency level being the third level, selecting a preheating capacity expansion operation.

[0018] According to an embodiment of the present disclosure, the dynamic adjustment module can also be used to switch stress testing traffic using a red-black deployment strategy during the process of executing resource adjustment operations based on the resource adjustment strategy; and / or set a resource change cooling period so that the time interval between two dynamic adjustments is not less than a preset duration.

[0019] According to an embodiment of the present disclosure, the stress testing device of the distributed system can also be used to predict the system impact of this resource adjustment and obtain an impact score before executing the resource adjustment operation based on the resource adjustment policy; and in response to the impact score exceeding the target threshold, prevent the execution of this resource change operation.

[0020] According to an embodiment of the present disclosure, the stress testing device of the distributed system can also be used to perform indicator anomaly detection on the real-time operation indicators based on a rule engine; and perform trend prediction on the real-time operation indicators based on a time series prediction model, and perform pattern anomaly detection based on the results of the trend prediction.

[0021] According to an embodiment of the present disclosure, the stress testing device of a distributed system can also be used to obtain root cause analysis results based on a causal reasoning model if indicator anomaly detection or pattern anomaly detection generates an abnormal result, wherein the causal reasoning model includes a causal correlation relationship diagram between multiple operating indicators.

[0022] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0023] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0024] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0025] According to the embodiments of the present disclosure, through the business tag-driven policy matching mechanism, it is possible to quickly select a stress testing solution that fits the system structure and operating characteristics, reducing unnecessary resource consumption; at the same time, the stress testing strategy is dynamically adjusted in combination with current and real-time operating indicators, so that the system can quickly respond and automatically adjust the stress testing intensity or resource allocation when facing resource bottlenecks, sudden loads or performance degradation, significantly improving the authenticity, accuracy and system stability of the stress testing. In addition, testers do not need to manually adjust scripts frequently or intervene in the stress testing process. The system can automatically generate, optimize and adjust the stress testing strategy based on business semantics and operating status, greatly reducing the stress testing configuration threshold and operation and maintenance burden, and improving the convenience, efficiency and reliability of the testing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0027] Figure 1 A diagram schematically illustrates an application scenario of a stress testing method, apparatus, device, medium, and program product for a distributed system according to an embodiment of the present disclosure;

[0028] Figure 2 The following schematically shows a flow chart of a stress testing method for a distributed system according to an embodiment of the present disclosure;

[0029] Figure 3 A flowchart of a method for generating a target stress testing strategy according to an embodiment of the present disclosure is schematically shown;

[0030] Figure 4 A flowchart of a method for dynamically adjusting a target stress testing strategy according to an embodiment of the present disclosure is schematically shown;

[0031] Figure 5 A block diagram schematically illustrates a structure of a stress testing device for a distributed system according to an embodiment of the present disclosure; and

[0032] Figure 6 A block diagram of an electronic device suitable for implementing a stress testing method for a distributed system according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0033] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0034] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0035] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0036] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0037] First, the technical terms described in this article are explained and illustrated as follows.

[0038] Load testing (abbreviated as "stress testing") refers to simulating a real or overloaded operating environment by constructing a large number of concurrent requests or high-frequency operations to evaluate the performance, resource carrying capacity, and stability of the target system under high-load conditions.

[0039] The Blue-Green Deployment (Red-Black Deployment) strategy is a version switching strategy during system changes. During resource expansion or policy adjustments, the system deploys the new version resources (red) in parallel with the current version resources (black). After the new version passes verification, traffic is switched to the new version, leaving the old version in a rollback state.

[0040] Fuzzy logic is a logical reasoning mechanism for dealing with uncertainty or fuzzy boundary problems, and is used to make flexible judgments when the indicator status does not have clear boundaries.

[0041] A causal graph, or causal inference model, is a graph structure used to express the causal relationships between variables within a system, such as a Bayesian network. This model can be used to model the dependency paths between operational metrics and trace the specific root causes when performance anomalies are detected.

[0042] Predictive warm-up means completing resource preparation operations (such as pre-scaling database read replicas) in advance of the actual system load based on load forecast results. This prevents response delays or resource shortages caused by a "cold start" of the system, thereby achieving a smoother stress testing process.

[0043] To support business needs such as high concurrency, high availability, and high consistency, financial institutions widely adopt distributed architectures to carry complex business processing flows. In this context, how to conduct high-intensity, high-realism performance testing on these distributed systems has become a prerequisite for ensuring the stable operation of financial systems and business continuity. In the field of distributed system performance testing, existing technologies generally adopt preset, batch-type load testing processes. Their goal is to assist development and operation and maintenance teams in evaluating the performance and bottlenecks of the system under high concurrency pressure by constructing an artificial load environment, running standardized test scripts, collecting response indicators, and conducting centralized analysis. This type of testing solution generally includes the following stages:

[0044] 1. Environment Preparation: Existing technologies typically rely on manual or semi-automated deployment methods to build the test infrastructure. For example, this includes hardware resource allocation, which involves pre-configuring physical machines or virtualized resources based on the test scale to form a fixed-size test cluster; system deployment, which involves deploying the test object into the test environment; and monitoring system deployment, which involves installing resource monitoring tools to collect server-level system metrics.

[0045] 2. Test Script Development: To simulate user behavior, traditional testing solutions generally rely on predefined scripts and tools for automated operations. Examples include: test case design, which abstracts representative scenarios (such as user login, order placement, and payment) based on business processes and converts them into static operation sequences; script implementation and parameterization, which uses tools to record or write test scripts and configure variable fields such as user IDs and access paths; and logical flow control, which uses structures like if / loop to achieve a certain degree of branching and looping control. However, the script's behavioral logic is relatively rigid and lacks runtime adaptability.

[0046] 3. Load Generation and Scheduling: The key to load testing lies in generating sufficiently realistic and continuous user requests. Existing technologies typically achieve this through methods such as virtual user construction, request scheduling strategies, and load control.

[0047] 4. Data collection and logging phase: To evaluate test results, existing systems generally deploy indicator collection and log collection mechanisms, which mainly include: performance indicator collection, periodic collection of system-level indicators (CPU usage, memory usage, load status) and business indicators; service log collection: using the log collection framework to aggregate the operation logs and exception information of each service for analyzing the root cause of bottlenecks; database indicator statistics, collecting data such as the number of database connections, slow queries, lock waits, etc., to assist in analyzing backend bottlenecks.

[0048] 5. Test report and result analysis phase After the test is completed, the system will provide visual test conclusions through summary analysis.

[0049] However, despite the relatively standard operating procedures in existing distributed system load testing solutions, which cover environment deployment, test script development, load generation, data collection, and report generation, they are still static, low in intelligence, and low in real-time performance. These shortcomings are exposed in practical applications, including the following:

[0050] On the one hand, the stress generation mechanism lacks intelligent scheduling capabilities. Existing test platforms generally use polling or simple random algorithms to distribute virtual user requests to different nodes. This ignores the real-time load differences between nodes in terms of resources such as CPU, memory, and I / O, and does not consider the heterogeneity of node hardware resources. This easily leads to the coexistence of "load hotspots" and "resource islands" during the stress distribution process.

[0051] On the other hand, test scenario modeling methods are overly rigid. Existing load testing solutions primarily rely on static scripts or parameterized requests to simulate business behavior, failing to reflect the dynamic fluctuations of real-world user traffic. For example, in e-commerce, finance, and other businesses, user access has distinct temporal distribution characteristics and sudden growth trends (such as flash sales, rush purchases, and centralized inquiries), and their behavior often exhibits nonlinear, clustered, or Poisson distribution patterns. Traditional scripts are unable to simulate the rhythm of traffic peaks, the variability of concurrent user behavior, or reproduce business processes involving multi-system collaboration (such as order-payment-refund). Furthermore, all virtual users use the same behavioral path by default, lacking profile differentiation and failing to simulate the complex interactions of multi-role collaboration.

[0052] Furthermore, resource allocation strategies lack flexibility and the ability to integrate with cloud environments. In traditional testing, resource configuration is generally statically set before the test begins, without the ability to dynamically expand or reclaim computing resources based on the test progress. In the early stages of testing, when the load has not yet reached its peak, a large amount of pre-allocated computing resources are idle, resulting in significant resource waste. During the surge in test pressure, the system is prone to request loss or forced degradation due to resource bottlenecks, making it impossible to stably execute the entire test process. Especially in hybrid cloud or edge computing environments, existing testing solutions cannot take into account the flexible scheduling of resources at the main center and edge nodes, have a high degree of resource coupling, and lack a unified cross-domain control mechanism.

[0053] In the monitoring and feedback phase, the collection of metrics and response control during the testing process are generally lagging behind. Existing solutions typically record system operating metrics such as response time, error rate, and system resource utilization at a sampling frequency of minutes, lacking high-frequency monitoring capabilities at the second or even millisecond level. During actual load testing, when the system experiences avalanche effects, cascading failures, or performance degradation, the test platform is often unable to detect and respond in a timely manner, missing the window for dynamic adjustment of test parameters.

[0054] Based on this, an embodiment of the present disclosure provides a stress testing method for a distributed system, including: obtaining a business scenario label and current operating indicators of a target distributed system; obtaining a stress testing policy template based on the business scenario label, generating a target stress testing policy based on the stress testing policy template and the current operating indicators; and executing a stress testing task on the target distributed system according to the target stress testing policy, wherein, during the execution of the stress testing task, the real-time operating indicators of the target distributed system are obtained; in response to the real-time operating indicators reaching preset adjustment conditions, the target stress testing policy is dynamically adjusted. Through the business label-driven policy matching mechanism, a stress testing solution that fits the system structure and operating characteristics can be quickly selected to reduce unnecessary resource consumption; at the same time, the stress testing policy is dynamically adjusted in combination with the current and real-time operating indicators, so that the system can quickly respond and automatically adjust the stress testing intensity or resource allocation when facing resource bottlenecks, sudden loads or performance degradation, significantly improving the authenticity, accuracy and system stability of the stress testing. In addition, testers no longer need to manually adjust scripts frequently or intervene in the stress testing process. The system can automatically generate, optimize, and adjust stress testing strategies based on business semantics and operating status, significantly reducing the stress testing configuration threshold and operation and maintenance burden, and improving the convenience, efficiency, and reliability of the testing process.

[0055] It should be noted that the stress testing methods, apparatuses, devices, media, and program products for distributed systems identified in this disclosure can be used in the fields of distributed technology, artificial intelligence technology, and financial technology, and can also be used in a variety of fields other than distributed technology, artificial intelligence technology, and financial technology. The application fields of the stress testing methods, apparatuses, devices, media, and program products for distributed systems provided in the embodiments of this disclosure are not limited.

[0056] In the technical solutions disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0057] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide users with corresponding operation portals for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge, and skills, and have reached a certain level of professionalism.

[0058] Figure 1 The application scenario diagram of the stress testing method, apparatus, device, medium and program product of the distributed system according to the embodiment of the present disclosure is schematically shown.

[0059] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0060] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0061] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0062] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0063] It should be noted that the stress testing method for the distributed system provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the stress testing device for the distributed system provided in the embodiment of the present disclosure can generally be set in the server 105. The stress testing method for the distributed system provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the stress testing device for the distributed system provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0064] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0065] The following will be based on Figure 1 The scene described by Figures 2 to 4 The stress testing method of the distributed system of the disclosed embodiment is described in detail.

[0066] Figure 2 The flowchart of the stress testing method for a distributed system according to an embodiment of the present disclosure is schematically shown.

[0067] like Figure 2 As shown, the stress testing method for a distributed system in this embodiment includes operations S210 to S230 , and the stress testing method for a distributed system can be executed by the server 105 .

[0068] In operation S210 , the business scenario tag and current operation index of the target distributed system are obtained.

[0069] In the embodiments of the present disclosure, the target distributed system may refer to any system composed of multiple computing nodes or service units that cooperate with each other, such as an e-commerce platform with a microservice architecture, a distributed database cluster, a big data processing platform, a real-time communication system, etc.

[0070] In an embodiment of the present disclosure, a business scenario label can be a feature description or classification of a target distributed system under a specific business activity. Exemplarily, business scenario labels can be defined and obtained from the following multiple dimensions: based on business processes, such as user registration and login, product browsing and searching, shopping cart operations, order payment, flash sales, data report generation, etc.; based on traffic characteristics, such as daily stable traffic, periodic peak traffic (such as daily noon peak and evening peak), sudden pulse traffic (such as the moment a marketing campaign starts), specific API high-frequency call scenarios, etc.; based on user behavior, such as novice user exploration paths, active user core function usage, dormant user awakening paths, etc.

[0071] In the embodiments of this disclosure, current operating metrics refer to the performance and status data of individual components or the entire target distributed system before the stress test task is executed. These metrics are used to assess the current health and resource usage of the system, providing a baseline for subsequent stress test strategy generation.

[0072] Exemplarily, current operating indicators may include: infrastructure layer indicators, such as CPU usage, average load, memory usage, available memory, network bandwidth, network latency, packet loss rate, etc. of each node; application / service layer indicators, such as QPS (queries per second), TPS (transactions per second), average service response time, error rate, thread pool, connection pool usage, etc. of key services; middleware indicators, such as the accumulation length of the message queue, production / consumption rate, cache system hit rate, number of database connections, etc.

[0073] In operation S220 , a stress testing policy template is obtained based on the business scenario tag, and a target stress testing policy is generated based on the stress testing policy template and the current operating indicator.

[0074] In the embodiments of this disclosure, a stress testing policy template is a set of stress testing frameworks or parameters preset for a specific business scenario. It defines the basic stress testing pattern and recommended values ​​or ranges for key parameters in that business scenario. For example, a stress testing policy template may include the target API / service for stress testing, a request parameter template, a stress testing traffic model, stress testing duration, and expected performance indicators.

[0075] For example, a template for a "flash sale" scenario might include: extremely high concurrent users, concentrated requests to the ordering API for specific products, short think times, short stress testing durations, but rapid traffic ramp-up. For example, a template for a "data report generation" scenario might include: low concurrency, but requests that consume significant backend resources (CPU, memory, I / O), and a high tolerance for response times. These templates can be stored in a template library and indexed and retrieved using business scenario tags.

[0076] In the embodiments of the present disclosure, the generation of a target stress testing strategy is a process of concretizing and personalizing a selected template. The current operating indicators are used to adjust and optimize the parameters in the template to make it more suitable for the current state of the system.

[0077] For example, if current system resource utilization (such as CPU and memory) is low, the initial concurrency or RPS defined in the template can be appropriately increased. Conversely, if certain key resources are already at high levels, it may be necessary to lower the initial load or adjust the traffic model to avoid components that are already under stress. For example, if the response time of certain services is currently approaching SLA (Service Level Agreement) thresholds, the generated target stress testing strategy may focus more on the stability of these services and set more sensitive monitoring alerts.

[0078] In operation S230, a stress testing task is performed on the target distributed system according to the target stress testing strategy, wherein, during the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained; in response to the real-time operating indicators reaching preset adjustment conditions, the target stress testing strategy is dynamically adjusted.

[0079] In the embodiments of the present disclosure, the preset adjustment conditions are rules or thresholds that trigger dynamic adjustments to the stress testing policy. For example, the preset adjustment conditions can be based on performance thresholds, such as the average response time of a critical service exceeding X milliseconds; resource bottlenecks, such as the CPU utilization of any service node consistently exceeding A% (e.g., 90%); or system stability, such as a large number of timed-out requests, or the failure or restart of a critical component.

[0080] During the execution of stress testing tasks, the system can determine the current load status based on real-time operating indicators, and dynamically adjust the target stress testing strategy when the preset adjustment conditions are met to achieve more intelligent and stable stress testing control.

[0081] For example, if all system indicators perform well and are far from reaching a bottleneck, you can gradually increase the number of concurrent users and request rate, or enter the next stress testing phase. If the system experiences a sharp drop in performance, a surge in error rates, or nears resource exhaustion, you should immediately reduce the load or even suspend stress testing to prevent the system from crashing, and record the system's performance under the current stress. If a specific API is found to be a bottleneck, you can try reducing the proportion of requests to that API and increasing requests to other APIs to detect other potential bottlenecks in the system. Adjust the think time of virtual users to change the intensity of requests, etc.

[0082] For example, during a stress test, if the CPU utilization of the order service reaches 95%, while the CPU utilization of the user service is only 30%, and the overall error rate begins to rise, the system may automatically trigger adjustments, temporarily maintaining the current pressure on the order service or slightly reducing it, while attempting to increase the pressure on the user service to more comprehensively assess the capacity of each part of the system. Alternatively, if the overall error rate exceeds 10%, the system may automatically reduce the current concurrency by 20% and observe whether the system returns to stability.

[0083] Figure 3 A flowchart of a method for generating a target stress testing strategy according to an embodiment of the present disclosure is schematically shown.

[0084] like Figure 3 As shown, the method for generating a target stress testing strategy in this embodiment may include operations S310 to S320.

[0085] In operation S310, based on the business scenario tag, a stress testing policy template matching the business scenario tag is searched in a preset stress testing policy template library. Upon receiving the business scenario tag, the system searches the library to find a stress testing policy template that matches or is most similar to the current business scenario tag. The selection can be based on an exact tag match, a partial tag match, or a predefined mapping relationship.

[0086] According to the embodiments of the present disclosure, once a stress testing policy template is successfully retrieved and selected from the stress testing policy template library, an initial stress testing policy can be directly generated based on the initial preset parameters contained in the template. The initial stress testing policy can be regarded as a universal, standardized stress testing solution for the business scenario.

[0087] In operation S320 , the initial stress testing strategy in the stress testing strategy template is fine-tuned based on the current operating indicator to obtain the target stress testing strategy.

[0088] According to embodiments of the present disclosure, some or all of the target parameters in the initial stress testing strategy can be analyzed and adjusted based on current operating indicators. The target parameters here refer to parameters in the initial stress testing strategy that are suitable for optimization based on the current state of the system, such as initial concurrency, load growth rate, weight or frequency of specific requests, etc.

[0089] Furthermore, the initial stress testing strategy can be integrated with the dynamic stress testing strategy adjustment factors. This integration can be done in a variety of ways, such as directly adding or multiplying the dynamic adjustment value to the corresponding parameter of the initial strategy, or using more complex rule-based logic to determine the final parameter value.

[0090] For example, if the initial policy recommends a concurrency of X, and the dynamic policy recommends increasing the concurrency by Y based on the current judgment that system resources are very idle, the combined target concurrency may be X + Y. If the initial policy weights a certain type of request at P, and the dynamic policy recommends reducing its weight by Q percentage points based on the current slow response of the backend service for this type of request, the combined weight may be P*(1-Q).

[0091] Through the fusion process, the target stress testing strategy inherits the experience and general settings for specific business scenarios in the stress testing strategy template, and also incorporates consideration and adaptive adjustments to the current actual operating status of the target distributed system, making it more targeted and effective.

[0092] Guided by the target stress testing strategy, the system initiates stress testing tasks on the target distributed system. During the stress test, the system continuously monitors and collects various operational metrics of the target distributed system in real time, such as CPU utilization, memory consumption, network traffic, disk I / O, service response time, transaction processing rate (TPS), queries per second (QPS), and error rate. When the combination or trend of real-time operational metrics reaches preset adjustment conditions—for example, if certain key metrics exceed warning thresholds or the system demonstrates potential performance inflection points or resource bottlenecks—the system triggers a dynamic adjustment mechanism for the current target stress testing strategy.

[0093] Figure 4 A flowchart of a method for dynamically adjusting a target stress testing strategy according to an embodiment of the present disclosure is schematically shown.

[0094] like Figure 4 As shown, the method for dynamically adjusting the target stress testing strategy in this embodiment may include operations S410 to S420.

[0095] In operation S410 , in response to the real-time operation indicator reaching a preset adjustment condition, the real-time operation indicator is judged using fuzzy logic to obtain a resource emergency level.

[0096] In some exemplary embodiments, the process of determining real-time operating indicators can utilize fuzzy logic theory. Fuzzy logic can handle imprecise and uncertain information, transforming real-time, multi-dimensional operating indicators (such as "high CPU usage" and "low error rate") into a comprehensive assessment of system resource status through a preset membership function and fuzzy rule base, and outputting a resource urgency level.

[0097] The resource urgency level can be a quantitative or qualitative indicator that indicates the severity of current system resource pressure or the urgency of the need for resource adjustment. For example, the level can be divided into high, medium, low, or more detailed levels, such as level 1, level 2, etc.

[0098] In operation S420 , a resource adjustment policy is selected based on the resource urgency level, and a resource adjustment operation is performed based on the resource adjustment policy.

[0099] In an embodiment of the present disclosure, the resource adjustment strategy includes adjustment measures for the computing, storage, network and other infrastructure resources occupied by the target distributed system itself, thereby influencing and coordinating the adjustment of the stress testing strategy.

[0100] For example, in response to the resource urgency level being level 1 (e.g., indicating extreme resource constraints and the system is on the verge of or has already experienced severe performance issues), the system selects to perform a rapid capacity expansion operation. Rapid capacity expansion can quickly and significantly increase the target system's available resources, such as significantly increasing the number of service instances or upgrading virtual machine specifications, in order to quickly alleviate system pressure and prevent system crashes.

[0101] Preferably, when the target distributed system enters a resource emergency state during the stress test, the system can trigger a second-level expansion operation to quickly improve the system's resource processing capabilities and prevent the stress test process from being interrupted or distorted due to performance bottlenecks. Specifically, the system can complete the startup and registration of a new container instance or service replica in about 5 seconds. After the expansion is completed, the stress test traffic will be automatically dispatched by the load balancing module or service grid to the newly expanded resource unit to form a balanced distribution. This "second-level expansion" mechanism can be widely used in scenarios with sudden loads, high concurrency shocks or rapid resource depletion, significantly enhancing the system's stress resistance and recovery capabilities during the stress test process, and is especially suitable for scenarios with elastic deployment capabilities under microservice architectures.

[0102] If the resource urgency level is level 2 (e.g., indicating high resource pressure and performance degradation, but not yet critical), you can choose to perform a gradual resource adjustment. Gradual resource adjustments are more gradual than rapid capacity expansion and can include adding resource instances in small increments, gradually increasing resource configurations, or optimizing existing resource allocations. This approach helps you observe the effects of each adjustment and avoid over-adjustments.

[0103] If the resource urgency level is Level 3 (for example, indicating that resources have some headroom, but based on stress testing trends or forecasts, there may be a shortage in the future), you can choose to perform a pre-emptive capacity expansion operation. This pre-emptive capacity expansion operation is proactive, meaning that resources are added and pre-empted (e.g., loading caches, establishing connection pools, etc.) before the expected peak load arrives, ensuring that the new resources can immediately handle traffic and function when needed.

[0104] In the embodiments of the present disclosure, in order to ensure the stability and security of the automated resource adjustment, a pre-evaluation step may be added before actually performing any of the above resource adjustment operations.

[0105] In this embodiment, the impact of a planned resource adjustment (such as capacity expansion, capacity reduction, or configuration change) on the entire target distributed system can be predicted. This prediction can be based on historical data, system models, or predefined rules, taking into account factors such as the magnitude of the adjustment, the components involved, and the current system stability, ultimately providing an impact score.

[0106] If the impact score exceeds a preset security threshold, indicating that the resource adjustment operation itself may bring significant risks (for example, it may cause temporary service interruption, high risk of data inconsistency, or cost exceeding the budget), the system will block the resource change operation. In this case, the stress testing strategy can be adjusted in other ways, such as adjusting the stress testing load parameters without changing system resources, or issuing an alert to operations personnel.

[0107] In the embodiments of the present disclosure, if the resource adjustment operation passes the impact assessment and is executed, to ensure a smooth process and reliable results, at least one of the following measures may be adopted:

[0108] 1. During resource adjustments (especially capacity expansions that involve adding new instances or replacing old ones), a "red-black" deployment strategy (or similar smooth transition strategies like blue-green deployment) can be used to seamlessly direct stress testing traffic to the updated system architecture. Specifically, the new resources are first configured as a "black" or "standby" environment. Once the environment is ready and passes preliminary verification, stress testing traffic is gradually switched from the "red" or "current" environment to the new environment, ensuring business continuity and testing effectiveness.

[0109] 2. To prevent system oscillation and instability caused by excessively frequent dynamic adjustments (including system resource adjustments and stress testing policy parameter adjustments), a resource change cooldown period can be set. This means that after a dynamic adjustment is completed, a minimum preset duration must elapse before the next dynamic adjustment can be initiated. This provides the necessary time window for the system to adapt to the new resource configuration and stress testing load, and for monitoring data to reflect actual performance.

[0110] Furthermore, during the execution of the stress testing task, in addition to using real-time operating indicators for the aforementioned dynamic adjustment decisions, continuous, multi-dimensional anomaly detection analysis can also be performed on these indicators in parallel.

[0111] In this embodiment, the system can have a built-in or integrated rule engine. The rule engine can pre-configure a series of detection rules for real-time operating indicators. These rules can be defined by experienced operation and maintenance or testing engineers, or automatically generated by learning from historical data.

[0112] For example, a threshold rule can be used: a single indicator exceeds or falls below a static or dynamically set threshold (for example, CPU usage is higher than 90% for 5 minutes, or the service error rate exceeds 5% instantaneously).

[0113] For example, a combination rule can be used: the logical combination of multiple indicators meets specific conditions (for example, when the TPS of service A drops by more than 20% and its average response time increases by more than 50%).

[0114] For example, a fluctuation rule can be used: the rate of change or fluctuation of the indicator within a certain time window exceeds the normal range.

[0115] When any one or a group of real-time operating indicators meets the abnormal conditions defined in the rule engine, the system will determine that an indicator abnormality has occurred and generate a corresponding abnormal event.

[0116] Optionally, in order to capture complex anomalies that do not necessarily violate fixed rules but deviate from normal behavior patterns, the system can also adopt a time series prediction model.

[0117] Specifically, time series prediction models (e.g., using ARIMA, LSTM, Prophet, etc.) can be built separately or jointly for key real-time operational metrics (such as QPS, response time, and resource utilization). These models are trained based on historical and current data and continuously predict the future short-term trends or expected values ​​of these metrics.

[0118] In this embodiment, the observed real-time operating indicator values ​​can be compared with the predicted values ​​provided by the time series forecasting model. When the deviation between the two (for example, absolute deviation, relative deviation, or statistically significant difference) exceeds a preset tolerance range or confidence interval, the system determines that a pattern anomaly has occurred. This anomaly typically indicates an unexpected and atypical change in system behavior, even if individual indicators may not have yet reached hard thresholds.

[0119] By combining the two anomaly detection mechanisms mentioned above, the system can more comprehensively monitor the health of the target distributed system, discovering both obvious indicator overruns and behavioral anomalies hidden in complex data patterns.

[0120] When any of the above anomaly detection mechanisms (i.e., indicator anomaly detection or pattern anomaly detection) generates an abnormal result, you can also initiate a root cause analysis process to help understand the nature of the problem and guide subsequent adjustments or repairs.

[0121] In one embodiment of the present disclosure, a causal reasoning model is used to perform root cause analysis. The core of the model is a causal relationship graph. The causal relationship graph can depict the known or learned causal dependencies between different components, different levels, and different operating indicators in the target distributed system. For example, the graph may include the following relationships: database server disk I / O saturation (cause) / increased database query latency (effect); increased database query latency (cause) / extended response time of the order service processing order interface (effect); a large number of messages are accumulated in the message queue (cause) / insufficient processing capacity of downstream consumer services (effect); memory leak in a specific microservice instance (cause) / frequent GC of the instance leads to increased CPU and slow response (effect).

[0122] When an anomaly is detected, the system uses this anomaly as a starting point to perform backtracking or correlation analysis on a causal relationship graph. By traversing the causal chain in the graph and combining it with the real-time indicator status of each node, the causal reasoning model can infer the most likely root cause or a set of candidate causes that led to the initial anomaly, known as the root cause analysis result.

[0123] For example, if the order service interface is detected to be responding slowly, the root cause analysis model may trace it back to the increased latency of the database query it relies on, which may be due to the database server disk I / O reaching a bottleneck.

[0124] The root cause analysis results can provide more precise guidance for dynamically adjusting the target stress testing strategy in operation S230. For example, if the root cause is a resource bottleneck in a specific component, then the stress testing strategy adjustment or system resource adjustment can be more targeted to that component.

[0125] By integrating the above-mentioned indicator anomaly detection, pattern anomaly detection, and root cause analysis capabilities based on causal reasoning, the disclosed method can not only dynamically execute and adjust stress testing, but also gain in-depth insights into the system's internal state and behavioral logic during the stress testing process, thereby achieving a higher level of automation and intelligence, and providing more powerful technical support for ensuring the stability and performance of distributed systems.

[0126] Those skilled in the art should understand that there are multiple technical paths for the specific implementation of the rule engine, the selection and training method of the time series prediction model, the construction method of the causal relationship diagram (manual definition, semi-automatic learning or fully automatic discovery), and the specific algorithm of causal reasoning. The selection and optimization of these specific implementations are all included in the concept of the present invention.

[0127] In summary, the distributed system intelligent stress testing method and related implementation provided by this disclosure achieves automated, personalized generation and precise matching of stress testing strategies by introducing business scenario tags and combining current system operating indicators with stress testing policy templates. During the stress testing process, not only can the stress testing load be dynamically adjusted based on real-time feedback, but fuzzy logic can also be used to intelligently assess the urgency level of system resources. Combined with robust mechanisms such as graded resource adjustment strategies (e.g., rapid, gradual, and pre-warming), impact estimation, smooth traffic switching (e.g., red-black deployment), and change cool-down periods, target system resources can be safely and effectively coordinated. Furthermore, by integrating indicator anomaly detection based on a rule engine, pattern anomaly detection based on time series prediction, and initiating deep root cause analysis based on causal reasoning models when anomalies occur, the method significantly improves the early identification, accuracy, and efficiency of problem resolution. These technical features work together to significantly enhance the automation, intelligence, adaptability, security, and reliability of distributed system stress testing, enabling more comprehensive and efficient identification of performance bottlenecks, assessment of the system's actual capacity and stability, and providing deep insights for rapid problem location and system optimization.

[0128] Corresponding to the above-mentioned stress testing method for a distributed system, an embodiment of the present disclosure further provides a stress testing device for a distributed system.

[0129] Figure 5 The structural block diagram of the stress testing device of the distributed system according to the embodiment of the present disclosure is schematically shown.

[0130] like Figure 5 As shown, the stress testing device 500 of the distributed system in this embodiment includes a data acquisition module 510 , a target stress testing strategy generation module 520 , and a dynamic adjustment module 530 .

[0131] The data acquisition module 510 can be used to obtain the business scenario label and current operating indicators of the target distributed system; the target stress testing strategy generation module. In one embodiment, the data acquisition module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0132] The target stress testing strategy generation module 520 can be used to obtain a stress testing strategy template based on the business scenario tag and generate a target stress testing strategy based on the stress testing strategy template and the current operating indicator. In one embodiment, the target stress testing strategy generation module 520 can be used to perform operation S220 described above, which will not be repeated here.

[0133] Dynamic adjustment module 530 can be configured to execute a stress testing task on the target distributed system according to the target stress testing strategy. During the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained. In response to the real-time operating indicators meeting preset adjustment conditions, the target stress testing strategy is dynamically adjusted. In one embodiment, dynamic adjustment module 530 can be configured to execute operation S230 described above, which will not be further described here.

[0134] According to an embodiment of the present disclosure, the target stress testing strategy generation module 520 can also be used to retrieve a stress testing strategy template that matches the business scenario label in a preset stress testing strategy template library based on the business scenario label; and fine-tune the initial stress testing strategy in the stress testing strategy template based on the current operating indicators to obtain the target stress testing strategy.

[0135] According to an embodiment of the present disclosure, the target stress testing strategy generation module 520 can also be used to obtain an initial stress testing strategy based on the initial preset parameters in the stress testing strategy template; adjust the target parameters in the initial stress testing strategy based on the current operating indicators to obtain a dynamic stress testing strategy; and merge the initial stress testing strategy and the dynamic stress testing strategy to form the target stress testing strategy.

[0136] According to an embodiment of the present disclosure, the dynamic adjustment module 530 can also be used to respond to the real-time operation indicator reaching a preset adjustment condition, use fuzzy logic to judge the real-time operation indicator, and obtain the resource urgency level; select a resource adjustment strategy based on the resource urgency level, and perform resource adjustment operations based on the resource adjustment strategy.

[0137] According to an embodiment of the present disclosure, the dynamic adjustment module 530 can also be used to select a resource adjustment strategy based on the resource urgency level, specifically including: in response to the resource urgency level being the first level, selecting a rapid capacity expansion operation; in response to the resource urgency level being the second level, selecting a gradual resource adjustment operation; and in response to the resource urgency level being the third level, selecting a preheating capacity expansion operation.

[0138] According to an embodiment of the present disclosure, the dynamic adjustment module 530 can also be used to switch stress testing traffic using a red-black deployment strategy during the process of performing resource adjustment operations based on the resource adjustment strategy; set a resource change cooling period so that the time interval between two consecutive dynamic adjustments is not less than a preset duration.

[0139] According to an embodiment of the present disclosure, the stress testing device 500 of the distributed system can also be used to predict the system impact of this resource adjustment and obtain an impact score before executing the resource adjustment operation based on the resource adjustment policy; and in response to the impact score exceeding the target threshold, prevent the execution of this resource change operation.

[0140] According to an embodiment of the present disclosure, the stress testing device 500 of the distributed system can also be used to perform indicator anomaly detection on the real-time operation indicators based on a rule engine; and perform trend prediction on the real-time operation indicators based on a time series prediction model, and perform pattern anomaly detection based on the results of the trend prediction.

[0141] According to an embodiment of the present disclosure, the stress testing device 500 of the distributed system can also be used to obtain root cause analysis results based on a causal reasoning model if indicator anomaly detection or pattern anomaly detection generates an abnormal result, wherein the causal reasoning model includes a causal correlation relationship diagram between multiple operating indicators.

[0142] According to embodiments of the present disclosure, any multiple modules among the data acquisition module 510, the target stress testing strategy generation module 520, and the dynamic adjustment module 530 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the data acquisition module 510, the target stress testing strategy generation module 520, and the dynamic adjustment module 530 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the data acquisition module 510, the target stress testing strategy generation module 520, and the dynamic adjustment module 530 can be at least partially implemented as a computer program module that, when executed, can perform the corresponding functionality.

[0143] Figure 6 A block diagram of an electronic device suitable for implementing a stress testing method for a distributed system according to an embodiment of the present disclosure is schematically shown.

[0144] like Figure 6As shown, an electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 606 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0145] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0146] According to an embodiment of the present disclosure, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0147] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0148] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above, and / or one or more memories other than ROM 602 and RAM 603.

[0149] Embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to enable the computer system to implement the distributed system stress testing method provided in the embodiments of the present disclosure.

[0150] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 601 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0151] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0152] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0153] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0155] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.

[0156] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A stress testing method for a distributed system, characterized in that: The method comprises: Obtain the business scenario tags and current operating indicators of the target distributed system; Obtaining a stress testing policy template based on the business scenario tag, and generating a target stress testing policy based on the stress testing policy template and the current operating indicator; and According to the target stress testing strategy, perform stress testing tasks on the target distributed system. During the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained; in response to the real-time operating indicators reaching preset adjustment conditions, the target stress testing strategy is dynamically adjusted.

2. The method according to claim 1, characterized in that The step of acquiring a stress testing policy template based on the business scenario tag and generating a target stress testing policy based on the stress testing policy template and the current operating indicator specifically includes: Based on the business scenario label, searching a preset stress testing policy template library for a stress testing policy template that matches the business scenario label; and The initial stress testing strategy in the stress testing strategy template is fine-tuned based on the current operating indicator to obtain the target stress testing strategy.

3. The method according to claim 2, characterized in that The fine-tuning of the initial stress testing strategy in the stress testing strategy template based on the current operating indicator specifically includes: Obtaining an initial stress testing strategy based on the initial preset parameters in the stress testing strategy template; Adjusting the target parameters in the initial stress testing strategy based on the current operating indicators to obtain a dynamic stress testing strategy; and The initial stress testing strategy and the dynamic stress testing strategy are integrated to form the target stress testing strategy.

4. The method according to any one of claims 1 to 3, characterized in that The dynamic adjustment of the target stress testing strategy specifically includes: In response to the real-time operation indicator reaching a preset adjustment condition, the real-time operation indicator is judged using fuzzy logic to obtain a resource emergency level; A resource adjustment policy is selected based on the resource urgency level, and a resource adjustment operation is performed based on the resource adjustment policy.

5. The method according to claim 4, characterized in that The selecting of a resource adjustment strategy based on the resource urgency level specifically includes: In response to the resource urgency level being the first level, selecting a rapid capacity expansion operation; In response to the resource urgency level being the second level, selecting a gradual resource adjustment operation; and In response to the resource urgency level being the third level, a preheating capacity expansion operation is selected.

6. The method according to claim 5, characterized in that In the process of performing resource adjustment operations based on the resource adjustment strategy, the red-black deployment strategy is used to switch the stress testing traffic; and / or a resource change cooling period is set so that the time interval between two dynamic adjustments is not less than a preset time length.

7. The method according to any one of claims 5 and 6, characterized in that The method further comprises: Before executing the resource adjustment operation based on the resource adjustment policy, predicting the system impact degree and obtaining an impact degree score; and In response to the impact score exceeding the target threshold, execution of the resource change operation is prevented.

8. The method according to any one of claims 1 to 3, 5 to 6, characterized in that: The method further comprises: Performing indicator anomaly detection on the real-time operating indicators based on a rule engine; and / or A trend forecast is performed on the real-time operating indicator based on a time series forecasting model, and pattern anomaly detection is performed based on the result of the trend forecast.

9. The method according to claim 8, characterized in that The method further comprises: If the indicator anomaly detection or the pattern anomaly detection generates an abnormal result, a root cause analysis result is obtained based on a causal reasoning model, wherein the causal reasoning model includes a causal correlation diagram between multiple operating indicators.

10. A stress testing device for a distributed system, characterized in that: The device comprises: The data acquisition module is used to obtain the business scenario labels and current operating indicators of the target distributed system; A target stress testing strategy generation module is configured to: obtain a stress testing strategy template based on the business scenario tag, and generate a target stress testing strategy based on the stress testing strategy template and the current operating indicator; A dynamic adjustment module is used to: perform a stress testing task on a target distributed system according to the target stress testing strategy, wherein, during the execution of the stress testing task, real-time operating indicators of the target distributed system are obtained; in response to the real-time operating indicators reaching preset adjustment conditions, the target stress testing strategy is dynamically adjusted.

11. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.