Fault isolation test method, device, equipment, medium and product

By configuring fault isolation parameters in Nginx and adjusting the test environment, the problem of Nginx open-source middleware being unable to quickly identify and isolate faults was solved, thus improving the system's stability and fault recovery capabilities.

CN121008950APending Publication Date: 2025-11-25CHINA CONSTRUCTION BANK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511347965.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, the Nginx open-source middleware is either not configured or configured improperly, which makes it impossible to quickly identify and isolate a server failure in the server cluster, thus affecting system stability.

Method used

By setting up a non-functional load testing environment, initialization and green light tests are performed, fault isolation parameters are configured, fault operation data of the downstream server cluster is obtained, fault isolation test results are determined based on the data, and parameters are adjusted when the expected standards are not met until the standards are met. Finally, the parameters are sent to the actual server cluster.

Benefits of technology

It enables rapid fault identification and timely isolation of Nginx, reducing the impact of a single failed server on overall production transactions and improving the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008950A_ABST
    Figure CN121008950A_ABST
Patent Text Reader

Abstract

The invention provides a fault isolation test method, device and equipment, a medium and a product, and relates to the technical field of computer software. The method comprises the following steps: initializing a pressure test non-functional test environment to obtain a configured upstream module; performing a green light test based on the module, and executing a parameter configuration operation to obtain a fault isolation parameter when passing the green light test; taking actual production pressure and production proportion, and taking fault operation data when any downstream server fails when stable transaction flow of preset duration is initiated according to the actual production pressure and production proportion through the pre-installed test tool; obtaining a fault isolation test result according to the operation data; when the result does not meet the expected test standard, adjusting the parameters, and returning to the step of taking the actual production pressure and production proportion until the test result meets the standard; and sending the parameters adjusted for the last time to an actual server cluster for configuration. The fault isolation parameters are reasonably configured, so that the fault can be isolated, and the influence of a single fault server on the whole production transaction is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer software technology, and in particular to a fault isolation testing method, apparatus, equipment, medium, and product. Background Technology

[0002] As business scales up and user numbers grow, a single server often struggles to handle high concurrency requests, and a single point of failure can paralyze the entire system. Therefore, it is necessary to use reasonable strategies to distribute traffic across multiple servers and ensure that the overall service is not affected when individual nodes malfunction. This requires attention to load balancing and fault isolation of the server cluster.

[0003] Currently, Nginx open-source middleware is used in the architecture and deployment of many systems. However, due to the lack of configuration or unreasonable configuration of Nginx parameters, when a server in the server cluster fails, it is impossible to quickly identify and isolate the fault in a timely manner, which affects the stability of the system. Summary of the Invention

[0004] This application provides a fault isolation testing method, apparatus, equipment, medium, and product to solve the technical problem that existing technologies using the Nginx open-source middleware cannot achieve rapid fault identification and timely isolation, thereby affecting system stability.

[0005] In a first aspect, embodiments of this application provide a fault isolation testing method applied to a load balancing server, the method comprising:

[0006] Set up a non-functional load testing environment and initialize it to obtain the configured upstream module;

[0007] A green light test is performed based on the configured upstream module, and when the green light test is detected to be passed, a parameter configuration operation is performed to obtain fault isolation parameters.

[0008] The actual production pressure and production ratio are obtained, and when a stable transaction flow of a preset duration is initiated by a pre-installed testing tool based on the actual production pressure and production ratio, the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails is obtained when the downstream server cluster runs according to the fault isolation parameters.

[0009] Based on the fault operation data of the downstream server cluster, determine the fault isolation test results;

[0010] When the fault isolation test result is found to be unsatisfactory, the fault isolation parameters are adjusted, and the process returns to the step of obtaining the actual production pressure and production ratio until the fault isolation test result meets the expected test standard.

[0011] Send the last adjusted fault isolation parameters to the actual server cluster for configuration.

[0012] In one possible design, the step of setting up a non-functional load testing environment and initializing the non-functional load testing environment to obtain a configured upstream module includes:

[0013] Receive user deployment configuration instructions and respond to the user deployment configuration instructions to complete the middleware service installation and generate middleware configuration files;

[0014] Start the middleware service and modify the middleware configuration file;

[0015] Round-robin was chosen as the load balancing strategy.

[0016] In one possible design, when the green light test is detected to have passed, a parameter configuration operation is performed to obtain fault isolation parameters, including:

[0017] When the green light test is detected as passed, a fault determination threshold and isolation duration are set.

[0018] Obtain the actual response time of the tested transaction, and set the timeout parameter based on the actual response time of the tested transaction;

[0019] The timeout parameter, the fault determination threshold, and the isolation duration constitute the fault isolation parameter.

[0020] In one possible design, the fault operation data of the downstream server cluster includes the number of transactions per second and resource changes of the downstream server cluster.

[0021] The step of determining the fault isolation test results based on the fault operation data of the downstream server cluster includes:

[0022] A transaction per second curve is generated based on the number of transactions per second of the downstream server cluster;

[0023] Determine the fault recovery time based on the aforementioned resource changes;

[0024] If the fault recovery time exceeds the expected target and / or the decrease in the number of transactions per second curve is greater than the preset decrease, then the failure to meet the expected test criteria is determined as a fault isolation test result.

[0025] Alternatively, if the fault recovery time is found to be within the expected range and the decrease in the number of transactions per second curve is less than or equal to the preset decrease range, then the fault isolation test result is determined to meet the expected test criteria.

[0026] In one possible design, the timeout parameters include proxy connection timeout, proxy read timeout, and proxy send timeout.

[0027] Accordingly, adjusting the fault isolation parameters includes:

[0028] The fault determination threshold is adjusted to obtain the adjusted fault determination threshold; or

[0029] The isolation duration is adjusted to obtain the adjusted isolation duration; or

[0030] The proxy connection timeout is adjusted to obtain the adjusted proxy connection timeout; or

[0031] The proxy read timeout is adjusted to obtain the adjusted proxy read timeout; or

[0032] The proxy sending timeout time is adjusted to obtain the adjusted proxy sending timeout time.

[0033] In one possible design, obtaining fault operation data of the downstream server cluster when any downstream server in the cluster fails, while the cluster operates according to the fault isolation parameters, includes:

[0034] The fault operation data of the downstream server cluster, when any downstream server in the downstream server cluster is out of service, is obtained through a preset monitoring system, and the cluster operates according to the fault isolation parameters.

[0035] The fault operation data of the downstream server cluster running according to the fault isolation parameters is obtained by acquiring fault operation data when any downstream server process in the downstream server cluster is suspended, through a preset monitoring system; or

[0036] The fault operation data of the downstream server cluster when any downstream server's network card fails is obtained through a preset monitoring system, and the downstream server cluster operates according to the fault isolation parameters.

[0037] In one possible design, the step of initiating a stable transaction flow of a preset duration based on the actual production pressure and production ratio using pre-installed testing tools includes:

[0038] By selecting a pre-installed pressure tool based on a preset transaction combination, a stable transaction flow with a preset duration of 5 minutes is initiated, simulating actual production pressure and 80% of the production rate.

[0039] Secondly, embodiments of this application provide a fault isolation testing device applied to a load balancing server, the device comprising:

[0040] The test environment setup module is used to set up a load testing non-functional test environment and initialize the load testing non-functional test environment to obtain the configured upstream module;

[0041] The test configuration module is used to perform a green light test based on the configured upstream module, and when the green light test is detected to be passed, to perform a parameter configuration operation to obtain fault isolation parameters.

[0042] The fault testing module is used to obtain the actual production pressure and production ratio, and when a stable transaction flow of a preset duration is initiated by the pre-installed testing tool based on the actual production pressure and production ratio, it obtains the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails, and the downstream server cluster runs according to the fault isolation parameters.

[0043] The fault testing module is also used to determine the fault isolation test results based on the fault operation data of the downstream server cluster.

[0044] The fault testing module is also used to adjust the fault isolation parameters and return to the step of obtaining the actual production pressure and production ratio when the fault isolation test result is detected to not meet the expected test standard, until the fault isolation test result meets the expected test standard.

[0045] The fault testing module is also used to send the last adjusted fault isolation parameters to the actual server cluster for configuration.

[0046] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0047] The memory stores computer-executed instructions;

[0048] The processor executes computer execution instructions stored in the memory to implement the fault isolation test method provided in the first aspect of this application.

[0049] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the fault isolation testing method provided in the first aspect of this application.

[0050] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, is used to implement the fault isolation testing method provided in the first aspect of this application.

[0051] This application provides a fault isolation testing method, apparatus, equipment, medium, and product. The method includes: setting up a non-functional stress testing environment and initializing it to obtain a configured upstream module; performing a green light test based on the configured upstream module, and when the green light test is detected as passed, performing parameter configuration operations to obtain fault isolation parameters; acquiring the actual production pressure and production ratio, and when a pre-installed testing tool initiates a stable transaction flow of a preset duration based on the actual production pressure and production ratio, acquiring fault operation data of the downstream server cluster running according to the fault isolation parameters when any downstream server in the downstream server cluster fails; determining the fault isolation test result based on the fault operation data of the downstream server cluster; when the fault isolation test result is detected as not meeting the expected test standard, adjusting the fault isolation parameters and returning to the step of acquiring the actual production pressure and production ratio until the fault isolation test result meets the expected test standard; and sending the last adjusted fault isolation parameters to the actual server cluster for configuration. The above method achieves the following technical effects: After passing the green light test, fault isolation parameters are configured to obtain fault operation data of the downstream server cluster running according to the fault isolation parameters. Fault isolation test results are obtained based on this data. If the fault isolation test results do not meet the expected test standards, the fault isolation parameters are continuously optimized and adjusted until the results do. By properly configuring the fault isolation parameters, Nginx can be endowed with proactive fault isolation capabilities. This method is simple and easy to use, and can effectively reduce the impact of a single failed server on overall production transactions. Attached Figure Description

[0052] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0053] Figure 1 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 1 ;

[0054] Figure 2 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 2 ;

[0055] Figure 3 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 3 ;

[0056] Figure 4 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 4 ;

[0057] Figure 5This is a schematic diagram of the fault isolation test device provided in the embodiments of this application;

[0058] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.

[0059] Explanation of reference numerals in the attached figures:

[0060] 901 - Processor; 902 - Memory; 903 - Communication components; 904 - Bus.

[0061] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0062] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0063] The collection, storage, use, processing, transmission, provision, and disclosure of financial data or user data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0064] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0065] First, let me explain the terms used in this application:

[0066] Nginx: A high-performance proxy server that can act as a load balancer to distribute client requests to multiple backend servers. It also boasts advantages such as fast response times, low memory usage, low resource consumption, open source, free availability, and simple configuration.

[0067] In order to clearly understand the technical solution of this application, the solutions of the prior art will be described in detail.

[0068] Currently, Nginx open-source middleware is used in the architecture and deployment of many systems. However, due to the lack of configuration or unreasonable configuration of Nginx parameters, when a server in the server cluster fails, it is impossible to quickly identify and isolate the fault in a timely manner, which affects the stability of the system.

[0069] In summary, the key issue this application aims to address is how to design a technology that can solve the problem of existing technologies using the Nginx open-source middleware failing to achieve rapid fault identification and timely isolation, thus affecting system stability.

[0070] Therefore, in view of the above-mentioned technical problems existing in the prior art, the embodiments of this application provide a fault isolation test method, device, equipment, medium and product, which aims to achieve rapid identification and timely isolation of faults through Nginx.

[0071] The following describes the application scenarios of the fault isolation testing methods, apparatus, equipment, media, and products provided in the embodiments of this application. These application scenarios are merely examples, intended to help those skilled in the art understand the technical content of this application, but do not imply that the embodiments of this application cannot be used in other devices, systems, environments, or scenarios.

[0072] 1) Daily performance testing of high-concurrency transaction systems: In scenarios with extremely high requirements for real-time performance and accuracy, such as payment and financial transactions, daily fault isolation testing is necessary to verify the system's stability under high-frequency requests. Simulating sudden failures of single or multiple servers, the fault isolation testing method provided in this application can quickly identify and isolate abnormal nodes, ensuring that the remaining servers can still process transactions normally, avoiding full-link blockage caused by localized failures, and guaranteeing the continuity of fund transactions and data consistency.

[0073] 2) Stability Verification Testing Before New Feature Launch: Before launching a new feature, it is necessary to verify its compatibility and fault tolerance with the existing system through fault isolation testing. For example, newly deployed recommendation algorithm service nodes may have memory leaks or logical vulnerabilities, causing slow response times for some requests. The fault isolation testing method provided in this application embodiment can quickly identify and isolate abnormal nodes, avoiding impact on all users.

[0074] Figure 1 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 1 The fault isolation testing method provided in this embodiment is applied to a load balancing server and includes the following steps:

[0075] S101. Set up a non-functional load testing environment and initialize it to obtain the configured upstream module.

[0076] In this embodiment, the load balancer is an Nginx load balancer. The load testing strategy is formulated using either JMeter or LoadRunner. JMeter is used to initiate the test, simulating the top 80% of transactions in production directly impacting the Nginx cluster. For example, the client forwards requests to three backend servers (IP1, IP2, and IP3) via Nginx. JMeter is an open-source performance testing tool developed by the Apache Software Foundation.

[0077] S102. Perform a green light test based on the configured upstream module, and when the green light test is detected to be passed, perform parameter configuration operation to obtain fault isolation parameters.

[0078] In this embodiment, the green light is tested first to ensure that all transactions during the stress test can be forwarded to the downstream server normally through Nginx.

[0079] S103. Obtain the actual production pressure and production ratio, and when a stable transaction flow of a preset duration is initiated by the pre-installed testing tool based on the actual production pressure and production ratio, obtain the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails, and the downstream server cluster runs according to the fault isolation parameters.

[0080] In this embodiment, the pre-installed testing tool is JMeter. A continuous and stable transaction scenario is initiated based on actual production pressure and production ratio. After running stably for approximately 5 minutes, a failure of a server in the downstream server cluster is simulated to obtain fault operation data of the downstream server cluster running according to fault isolation parameters.

[0081] S104. Determine the fault isolation test results based on the fault operation data of the downstream server cluster.

[0082] In this embodiment, after obtaining the fault isolation test results, it is determined whether the fault isolation test results meet the expected test standards.

[0083] S105. When the fault isolation test result is found to be unsatisfactory, the fault isolation parameters are adjusted and the process returns to the step of obtaining the actual production pressure and production ratio until the fault isolation test result meets the expected test standard.

[0084] In this embodiment, if the fault isolation test result does not meet the expected test standard, the fault isolation parameters are adjusted and the process returns to S103 until the fault isolation test result meets the expected test standard.

[0085] S106. Send the last adjusted fault isolation parameters to the actual server cluster for configuration.

[0086] In this embodiment, after the fault isolation test results meet the expected test standards, the last adjusted fault isolation parameters are sent to the actual server cluster for configuration.

[0087] By adjusting fault isolation parameters in a simulated testing environment and continuously adjusting these parameters when the fault isolation test results do not meet the expected test standards, a fault isolation parameter is found that minimizes the impact on overall production transactions when the load balancer and downstream service clusters fail. This parameter is then applied to the actual production server cluster to enable the actual server cluster to effectively achieve failover as expected, quickly restore transaction capabilities, and thus reduce the impact on actual production.

[0088] After passing the green light test, configure fault isolation parameters, obtain fault operation data of the downstream server cluster running according to the fault isolation parameters, obtain fault isolation test results based on the fault operation data, and continuously optimize and adjust the fault isolation parameters when the fault isolation test results do not meet the expected test standards until the fault isolation test results meet the expected test standards. By properly configuring the fault isolation parameters, Nginx can be endowed with proactive fault isolation capabilities. This method is simple to operate and easy to learn, and can effectively reduce the impact of a single failed server on overall production transactions.

[0089] This application provides a fault isolation testing method, comprising: setting up a load testing non-functional test environment and initializing the load testing non-functional test environment to obtain a configured upstream module; performing a green light test based on the configured upstream module, and when the green light test is detected to be passed, performing parameter configuration operations to obtain fault isolation parameters; obtaining the actual production pressure and production ratio, and when a pre-installed test tool initiates a stable transaction flow of a preset duration based on the actual production pressure and production ratio, obtaining fault operation data of the downstream server cluster running according to the fault isolation parameters when any downstream server in the downstream server cluster fails; determining the fault isolation test result based on the fault operation data of the downstream server cluster; when the fault isolation test result is detected to not meet the expected test standard, adjusting the fault isolation parameters and returning to the step of obtaining the actual production pressure and production ratio until the fault isolation test result meets the expected test standard; and sending the last adjusted fault isolation parameters to the actual server cluster for configuration. The above method achieves the following technical effects: After passing the green light test, fault isolation parameters are configured to obtain fault operation data of the downstream server cluster running according to the fault isolation parameters. Fault isolation test results are obtained based on this data. If the fault isolation test results do not meet the expected test standards, the fault isolation parameters are continuously optimized and adjusted until the results do. By properly configuring the fault isolation parameters, Nginx can be endowed with proactive fault isolation capabilities. This method is simple and easy to use, and can effectively reduce the impact of a single failed server on overall production transactions.

[0090] Figure 2 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 2 In the fault isolation test method provided in this embodiment, S101 includes the following steps:

[0091] S201. Receive user deployment configuration instructions and respond to user deployment configuration instructions to complete the middleware service installation and generate middleware configuration files.

[0092] In this embodiment, the middleware service includes the Nginx service and the downstream server service, and the middleware configuration file includes the Nginx configuration file nginx.conf.

[0093] S202. Start the middleware service and modify the middleware configuration file.

[0094] In this embodiment, the Nginx service and downstream server service are started, and the Nginx configuration file nginx.conf is modified. The main modification is to connect the IP address and port of the downstream server to the upstream module, enabling unified management of the downstream server cluster. This lays the foundation for subsequent load balancing strategy configuration and fault isolation mechanism deployment, ensuring that requests are distributed to the designated servers as expected. The upstream module is the upstream module.

[0095] S203. Select round-robin as the load balancing strategy.

[0096] In this embodiment, a load balancing strategy is configured, and round-robin is selected as the load balancing strategy. The Nginx load balancer maintains a list of servers to forward requests to. When the first request arrives, it forwards the request to the first server in the list; when the second request arrives, it forwards it to the second server in the list; and so on. After distributing the request to the last server in the list, it returns to the first server in the list and begins a new round of looping.

[0097] Figure 3 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 3 In the fault isolation test method provided in this embodiment, when the green light test is detected to be passed in S102, a parameter configuration operation is performed to obtain the fault isolation parameters, including the following steps:

[0098] S301. When the green light test is detected as passed, set the fault judgment threshold and isolation duration.

[0099] In this embodiment, the test green light can be a test to ensure that all transactions under stress testing can send requests from the client and be forwarded to the downstream server through the Nginx server.

[0100] The failure threshold is `max_fails`, and the isolation period is `fail_timeout`. These two parameters are used to determine whether a server in the load balancing upstream module is effective. The working principle is as follows: When a server fails, if the number of consecutive requests to that server reaches the number specified by `max_fails`, Nginx considers that server to be failed and removes it from the list of available servers. For the following time, specified by `fail_timeout`, that server remains considered unavailable. After the `fail_timeout` expires, requests are forwarded to the failed server again, and the number of requests is used to probe whether it has recovered. If it has not recovered, some requests will fail because they were assigned to the failed server; if it has recovered, that server is added back to the list of available servers and resumes processing requests.

[0101] S302. Obtain the actual response time of the tested transaction and set the timeout parameter based on the actual response time of the tested transaction.

[0102] As an optional implementation, historical transaction response time data is collected to determine the normal response time range and the maximum tolerable delay value; combined with business requirements and system fault tolerance, the timeout parameter is set to a value slightly higher than the actual response time of the tested transactions.

[0103] S303. The timeout parameter, fault determination threshold and isolation duration are combined to form the fault isolation parameter.

[0104] In this embodiment, the fault isolation parameters consist of timeout parameters, fault determination thresholds, and isolation duration. This enables standardized configuration of fault identification, determination, and handling, facilitating unified management and dynamic adjustment, ensuring clear and controllable fault isolation logic, and improving the system's response efficiency and processing accuracy for abnormal servers.

[0105] Figure 4 A flowchart illustrating the fault isolation testing method provided in this application embodiment. Figure 4 In the fault isolation testing method provided in this embodiment, the fault operation data of the downstream server cluster includes the number of transactions per second and resource changes of the downstream server cluster. S104 includes the following steps:

[0106] S401. Generate a transaction per second curve based on the number of transactions per second of the downstream server cluster.

[0107] In this embodiment, a Transactions Per Second (TPS) curve is generated based on the number of transactions per second of the downstream server cluster, i.e., the number of transactions that the servers in the downstream server cluster can successfully process per second. The TPS curve is automatically generated by a preset monitoring system.

[0108] S402. Determine the fault recovery time based on resource changes.

[0109] In this embodiment, normal thresholds are set for each resource. When a fault occurs, the starting time of abnormal fluctuations in each resource is recorded, and the process of each resource indicator recovering from the abnormal state to the normal threshold is continuously tracked. The fault recovery time is calculated by subtracting the fault occurrence time from the time when each resource returns to stability.

[0110] Based on the fault recovery time and TPS curve, determine whether the fault isolation test result is the situation in S403 or S404.

[0111] S403. When the fault recovery time exceeds the expected indicator and / or the decrease in the number of transactions per second curve is greater than the preset decrease, the failure to meet the expected test standard is determined as a fault isolation test result.

[0112] In this embodiment, three scenarios are defined as failure to meet the expected test criteria in the fault isolation test: 1) The fault recovery time exceeds the expected target, but the TPS curve decreases by less than or equal to the preset decrease; 2) The fault recovery time does not exceed the expected target, but the TPS curve decreases by more than the preset decrease; 3) The fault recovery time exceeds the expected target, and the TPS curve decreases by more than the preset decrease. If the fault recovery time does not exceed the expected target, but the TPS curve reaches its lowest point, this is generally considered an undesirable situation, meaning the entire system is unavailable during the fault, significantly impacting production.

[0113] If the fault isolation test results do not meet the expected test standards, it indicates that the fault isolation effect is poor and the fault isolation parameters need to be adjusted. For example, the maximum number of failures in the fault isolation parameters can be adjusted from 3 to 2.

[0114] S404. When the fault recovery time is found to be within the expected range and the decrease in the number of transactions per second curve is less than or equal to the preset decrease range, the fault isolation test result is determined to meet the expected test criteria.

[0115] In this embodiment, when the fault recovery time is detected to be within the expected range and the TPS curve decreases by less than or equal to the preset decrease range, the result is determined to meet the expected test criteria as the fault isolation test result.

[0116] Judging whether the fault isolation test results meet the expected test standards by analyzing the fault recovery time and transactions per second curves can intuitively quantify the fault isolation test results, quickly verify whether the system can restore normal transaction processing capabilities in a timely manner after a fault occurs, and ensure that the isolation mechanism is effective and meets the business requirements for availability and performance.

[0117] Based on the above embodiments, this application provides a fault isolation testing method. In the fault isolation testing method provided in this embodiment, the timeout parameters include proxy connection timeout, proxy read timeout, and proxy send timeout. Adjusting the fault isolation parameters in step S105 includes:

[0118] S501. Adjust the fault determination threshold to obtain the adjusted fault determination threshold.

[0119] S502. Adjust the isolation duration to obtain the adjusted isolation duration.

[0120] In this embodiment, the fault isolation parameters can be dynamically adjusted based on the fault isolation test results. For example, if the fault isolation delay is large, the isolation time can be shortened, such as from 10 seconds to 5 seconds; if false isolation is frequent, the fault judgment threshold can be increased, such as from 3 times to 5 times, or the timeout time can be extended.

[0121] S503. Adjust the proxy connection timeout to obtain the adjusted proxy connection timeout.

[0122] S504. Adjust the proxy read timeout to obtain the adjusted proxy read timeout.

[0123] S505. Adjust the proxy sending timeout to obtain the adjusted proxy sending timeout.

[0124] In this embodiment, the proxy connection timeout is `proxy_connect_timeout`, representing the connection timeout between Nginx and the backend server; the proxy read timeout is `proxy_read_timeout`, representing the timeout between two successful response operations between Nginx and the backend server; and the proxy send timeout is `proxy_send_timeout`, representing the timeout for Nginx to transfer files to the backend server. These three timeout parameters should be set based on the actual response times of tested transactions, and the values ​​should be greater than the response times of all transactions to avoid misidentification and transaction failure. These three timeout parameters can be obtained from historical data of previous server transactions or based on user experience.

[0125] When the fault isolation test results do not meet the expected test standards, the fault isolation parameters need to be adjusted. These parameters include the fault determination threshold, isolation duration, proxy connection timeout, proxy read timeout, and proxy send timeout. Therefore, adjusting the fault isolation parameters includes: adjusting the fault determination threshold via S501, adjusting the isolation duration via S502, adjusting the proxy connection timeout via S503, adjusting the proxy read timeout via S504, and adjusting the proxy send timeout via S505.

[0126] Adjusting fault isolation parameters is not a process that can be achieved in one or two attempts; it requires repeated adjustments, testing, and observation until an ideal result is found. For example, if the downstream server cluster has two servers, and one server fails while the processing capacity of the other remains unaffected, then the fault will only impact approximately 50% of the TPS curve.

[0127] When the fault isolation test results are found to be unsatisfactory, the fault identification sensitivity and isolation strategy can be optimized by adjusting the fault judgment threshold, isolation duration, proxy connection timeout, proxy read timeout, or proxy send timeout. This makes the fault isolation mechanism more aligned with actual business scenarios, thereby improving the system's accuracy in handling anomalies and its overall stability.

[0128] Based on the above embodiments, this application provides a fault isolation testing method. In the fault isolation testing method provided in this embodiment, step S103, obtaining fault operation data of the downstream server cluster running according to fault isolation parameters when any downstream server in the downstream server cluster fails, includes:

[0129] S601. Obtain fault operation data of the downstream server cluster when any downstream server in the downstream server cluster stops service, according to the fault isolation parameters.

[0130] S602. Obtain fault operation data of the downstream server cluster when any downstream server process in the downstream server cluster is suspended, according to the fault isolation parameters.

[0131] S603. Obtain fault operation data of the downstream server cluster when any downstream server's network card fails, based on the fault isolation parameters, through a preset monitoring system.

[0132] In this embodiment, the TPS curve, response time, and resource changes of the entire server cluster are observed in real time through a preset monitoring system.

[0133] Failures occurring on any downstream server in the downstream server cluster can be categorized into three scenarios: service outage, process suspension, and network card failure. Therefore, obtaining fault operation data of the downstream server cluster under fault isolation parameters when any downstream server fails includes the following three scenarios: S601 obtaining fault operation data of the downstream server cluster under fault isolation parameters when any downstream server fails (using a preset monitoring system); S602 obtaining fault operation data of the downstream server cluster under fault isolation parameters when any downstream server's process is suspended (using a preset monitoring system); and S603 obtaining fault operation data of the downstream server cluster under fault isolation parameters when any downstream server's network card fails (using a preset monitoring system). Optionally, failures occurring on any downstream server in the downstream server cluster can also include system crashes.

[0134] The system categorizes server failures into three types: service interruption, process suspension, and network card failure. It also collects operational data corresponding to each failure type, enabling precise identification of the root cause, rapid isolation of problematic components, and targeted design of automated repair strategies. This reduces the cost of manual intervention and ultimately enhances the overall stability and reliability of the system.

[0135] Based on the above embodiments, this application provides a fault isolation testing method. In the fault isolation testing method provided in this embodiment, step S103, which involves initiating a stable transaction flow of a preset duration based on actual production pressure and production ratio using pre-installed testing tools, includes:

[0136] S701: By selecting a pre-installed pressure tool according to a preset transaction combination, it simulates actual production pressure and 80% of the production ratio to initiate a stable transaction flow with a preset duration of 5 minutes.

[0137] In this embodiment, the production ratio is 80%, and the preset duration for stable transaction traffic is 5 minutes. Refining these two values—production ratio and preset duration for stable transaction traffic—allows for precise control of the traffic ratio between the test and production environments, avoiding resource waste or overload. Furthermore, setting a reasonable stabilization duration ensures system reliability under real-world load conditions.

[0138] Here is a specific example, taking a downstream server cluster consisting of one Nginx server and two downstream servers A and B, with an initial TPS of 1000.

[0139] Step 1: Set up a non-functional load testing environment in Nginx and configure the round-robin load balancing strategy.

[0140] Step 2: Perform a green light test in Nginx to ensure all traffic can be forwarded successfully to downstream servers A and B. After passing the green light test, configure fault isolation parameters. For example, set `max_fails` to a maximum of 3 failures and `fail_timeout` to a 10-second timeout.

[0141] Step 3: First, initiate continuous and stable transactions based on actual production pressure and production ratio. After the traffic stabilizes for 5 minutes, simulate the downtime of downstream server A. At this time, when Nginx sends a request to server A and fails to time out 3 times in a row, immediately isolate server A. In the following 10 seconds, all 1000 TPS are redirected to server B. Since the processing capacity of server B is limited to 500 TPS, the overall TPS of the downstream server cluster decreases and stabilizes at 500, rather than going to zero. After 10 seconds, attempt to restart server A.

[0142] Step 4: Based on the TPS curve and resource changes of the downstream server cluster collected during the isolation period, for example, if monitoring shows that it took 4 seconds from server A's crash to its isolation, with some requests failing during this period, it indicates that the isolation is too slow. In this case, the maximum number of failures is changed from 3 to 2, and step 3 is executed again. It was then found that the isolation was completed in just two seconds, thus restoring system stability more quickly.

[0143] Figure 5 This is a schematic diagram of the fault isolation test device provided in an embodiment of this application. Figure 5 As shown, in this embodiment, the fault isolation test device includes:

[0144] The test environment setup module 801 is used to set up a non-functional load testing environment and initialize the non-functional load testing environment to obtain the configured upstream module.

[0145] The test configuration module 802 is used to perform a green light test based on the configured upstream module, and when the green light test is detected to be passed, it performs parameter configuration operations to obtain fault isolation parameters.

[0146] The fault testing module 803 is used to obtain the actual production pressure and production ratio, and when a stable transaction flow of a preset duration is initiated by the pre-installed testing tool based on the actual production pressure and production ratio, it obtains the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails, and the downstream server cluster runs according to the fault isolation parameters.

[0147] The fault testing module 803 is also used to determine the fault isolation test results based on the fault operation data of the downstream server cluster.

[0148] The fault test module 803 is also used to adjust the fault isolation parameters and return to the step of obtaining the actual production pressure and production ratio when the fault isolation test result is detected to not meet the expected test standard, until the fault isolation test result meets the expected test standard.

[0149] The fault testing module 803 is also used to send the last adjusted fault isolation parameters to the actual server cluster for configuration.

[0150] The fault isolation test device provided in this embodiment can perform... Figure 1 The technical solution of the fault isolation test method embodiment shown herein, its implementation principle and technical effect are similar to Figure 1 The examples of the fault isolation test methods shown are similar and will not be described in detail here.

[0151] Meanwhile, the fault isolation test device provided by the present invention is a further refinement of the fault isolation test device provided in the previous embodiment.

[0152] Optionally, in this embodiment, the test environment setup module 801 is further used for:

[0153] Receive user deployment configuration instructions and respond to them by completing the middleware service installation and generating the middleware configuration file; start the middleware service and modify the middleware configuration file; select round-robin as the load balancing strategy.

[0154] Optionally, in this embodiment, the test configuration module 802 is further used for:

[0155] When the green light test is detected as passed, set the fault judgment threshold and isolation duration; obtain the actual response time of the tested transaction, and set the timeout parameter based on the actual response time of the tested transaction; combine the timeout parameter, fault judgment threshold and isolation duration to form the fault isolation parameter.

[0156] Optionally, in this embodiment, the fault operation data of the downstream server cluster includes the number of transactions per second and resource changes of the downstream server cluster. The fault testing module 803 is also used for:

[0157] A transaction-per-second (TPS) curve is generated based on the number of transactions per second (TPS) of the downstream server cluster. The fault recovery time is determined based on resource changes. If the fault recovery time exceeds the expected target and / or the TPS curve decreases by more than the preset decrease, the fault isolation test result is determined to be unsatisfactory. Alternatively, if the fault recovery time does not exceed the expected target and the TPS curve decreases by less than or equal to the preset decrease, the fault isolation test result is determined to be satisfactory.

[0158] Optionally, in this embodiment, the timeout parameters include the proxy connection timeout, the proxy read timeout, and the proxy send timeout. The fault test module 803 is also used for:

[0159] Adjust the fault determination threshold to obtain the adjusted fault determination threshold; or adjust the isolation duration to obtain the adjusted isolation duration; or adjust the proxy connection timeout to obtain the adjusted proxy connection timeout; or adjust the proxy read timeout to obtain the adjusted proxy read timeout; or adjust the proxy send timeout to obtain the adjusted proxy send timeout.

[0160] Optionally, in this embodiment, the fault test module 803 is further used for:

[0161] The system obtains fault operation data of the downstream server cluster when any downstream server in the downstream server cluster stops service and the downstream server cluster is running according to fault isolation parameters; or obtains fault operation data of the downstream server cluster when any downstream server process in the downstream server cluster is suspended and the downstream server cluster is running according to fault isolation parameters; or obtains fault operation data of the downstream server cluster when any downstream server's network card fails and the downstream server cluster is running according to fault isolation parameters.

[0162] Optionally, in this embodiment, the fault test module 803 is further used for:

[0163] By selecting a pre-installed pressure tool based on a preset transaction combination, a stable transaction flow with a preset duration of 5 minutes is initiated, simulating actual production pressure and 80% of the production rate.

[0164] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device provided in this embodiment includes at least one processor 901 and a memory 902. Optionally, the electronic device further includes a communication component 903. The processor 901, memory 902, and communication component 903 are connected via a bus 904.

[0165] In the specific implementation process, at least one processor 901 executes computer execution instructions stored in memory 902, causing at least one processor 901 to execute the above-mentioned fault isolation test method.

[0166] The specific implementation process of processor 901 can be found in the above-described fault isolation test method embodiment, which has a similar implementation principle and technical effect, and will not be repeated here.

[0167] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0168] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0169] Bus 904 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 904 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 904 in the accompanying drawings of this application is not limited to only one bus or one type of bus.

[0170] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described fault isolation test method.

[0171] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the above-described fault isolation test method.

[0172] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0173] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0174] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0175] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0176] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0177] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0178] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0179] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A fault isolation test method, characterized in that, Applied to a load balancing server, the method includes: Set up a non-functional load testing environment and initialize it to obtain the configured upstream module; A green light test is performed based on the configured upstream module, and when the green light test is detected to be passed, a parameter configuration operation is performed to obtain fault isolation parameters. The actual production pressure and production ratio are obtained, and when a stable transaction flow of a preset duration is initiated by a pre-installed testing tool based on the actual production pressure and production ratio, the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails is obtained when the downstream server cluster runs according to the fault isolation parameters. Based on the fault operation data of the downstream server cluster, determine the fault isolation test results; When the fault isolation test result is found to be unsatisfactory, the fault isolation parameters are adjusted, and the process returns to the step of obtaining the actual production pressure and production ratio until the fault isolation test result meets the expected test standard. Send the last adjusted fault isolation parameters to the actual server cluster for configuration.

2. The method according to claim 1, characterized in that, The process involves setting up a non-functional load testing environment and initializing it to obtain a configured upstream module, including: Receive user deployment configuration instructions and respond to the user deployment configuration instructions to complete the middleware service installation and generate middleware configuration files; Start the middleware service and modify the middleware configuration file; Round-robin was chosen as the load balancing strategy.

3. The method according to claim 1, characterized in that, When the green light test is detected as passed, a parameter configuration operation is performed to obtain fault isolation parameters, including: When the green light test is detected as passed, a fault determination threshold and isolation duration are set. Obtain the actual response time of the tested transaction, and set the timeout parameter based on the actual response time of the tested transaction; The timeout parameter, the fault determination threshold, and the isolation duration constitute the fault isolation parameter.

4. The method according to claim 1, characterized in that, The fault operation data of the downstream server cluster includes the number of transactions per second and resource changes of the downstream server cluster; The step of determining the fault isolation test results based on the fault operation data of the downstream server cluster includes: A transaction per second curve is generated based on the number of transactions per second of the downstream server cluster; Determine the fault recovery time based on the aforementioned resource changes; If the fault recovery time exceeds the expected target and / or the decrease in the number of transactions per second curve is greater than the preset decrease, then the failure to meet the expected test criteria is determined as a fault isolation test result. Alternatively, if the fault recovery time is found to be within the expected range and the decrease in the number of transactions per second curve is less than or equal to the preset decrease range, then the fault isolation test result is determined to meet the expected test criteria.

5. The method according to claim 3, characterized in that, The timeout parameters include proxy connection timeout, proxy read timeout, and proxy send timeout. Accordingly, adjusting the fault isolation parameters includes: The fault determination threshold is adjusted to obtain the adjusted fault determination threshold; or The isolation duration is adjusted to obtain the adjusted isolation duration; or The proxy connection timeout is adjusted to obtain the adjusted proxy connection timeout; or The proxy read timeout is adjusted to obtain the adjusted proxy read timeout; or The proxy sending timeout time is adjusted to obtain the adjusted proxy sending timeout time.

6. The method according to claim 1, characterized in that, The fault operation data obtained when any downstream server in the downstream server cluster fails, and the downstream server cluster operates according to the fault isolation parameters, includes: The fault operation data of the downstream server cluster, when any downstream server in the downstream server cluster is out of service, is obtained through a preset monitoring system, and the cluster operates according to the fault isolation parameters. The fault operation data of the downstream server cluster running according to the fault isolation parameters is obtained by acquiring fault operation data when any downstream server process in the downstream server cluster is suspended, through a preset monitoring system; or The fault operation data of the downstream server cluster when any downstream server's network card fails is obtained through a preset monitoring system, and the downstream server cluster operates according to the fault isolation parameters.

7. The method according to any one of claims 1 to 6, characterized in that, The step of initiating a stable transaction flow of a preset duration based on the actual production pressure and production ratio using pre-installed testing tools includes: By selecting a pre-installed pressure tool based on a preset transaction combination, a stable transaction flow with a preset duration of 5 minutes is initiated, simulating actual production pressure and 80% of the production rate.

8. A fault isolation test device, characterized in that, The device, used in a load balancing server, includes: The test environment setup module is used to set up a load testing non-functional test environment and initialize the load testing non-functional test environment to obtain the configured upstream module; The test configuration module is used to perform a green light test based on the configured upstream module, and when the green light test is detected to be passed, to perform a parameter configuration operation to obtain fault isolation parameters. The fault testing module is used to obtain the actual production pressure and production ratio, and when a stable transaction flow of a preset duration is initiated by the pre-installed testing tool based on the actual production pressure and production ratio, it obtains the fault operation data of the downstream server cluster when any downstream server in the downstream server cluster fails, and the downstream server cluster runs according to the fault isolation parameters. The fault testing module is also used to determine the fault isolation test results based on the fault operation data of the downstream server cluster. The fault testing module is also used to adjust the fault isolation parameters and return to the step of obtaining the actual production pressure and production ratio when the fault isolation test result is detected to not meet the expected test standard, until the fault isolation test result meets the expected test standard. The fault testing module is also used to send the last adjusted fault isolation parameters to the actual server cluster for configuration.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the fault isolation test method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the fault isolation test method as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the fault isolation test method as described in any one of claims 1 to 7.