High-availability testing method and system for task scheduling of computing power host management platform
The method and system for high-availability testing of task scheduling in resource management platforms address testing gaps in high-concurrency and cross-domain scenarios, improving system resilience and resource efficiency through mixed load simulation and intelligent scheduling, with real-time monitoring and data-driven optimization.
Patent Information
- Application Number
- CN202510503998.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-15
AI Technical Summary
The existing computing power host management platform has insufficient testing coverage rate in high concurrent scenarios, imperfect cross-domain scheduling verification, insufficient verification of fault recovery and self-healing capabilities, and weak data-driven optimization capabilities.
By simulating hybrid load testing, cross-regional intelligent scheduling, resource failure migration recovery, multi-dimensional monitoring and data acquisition, combined with analysis and optimization, a highly available testing method and system is built to fully cover complex scenarios, and the system's elastic scaling capacity, optimal resource allocation and fault recovery capabilities are verified.
It improves the test coverage rate in high concurrency scenarios, reduces the risk of system crashes, improves resource utilization and cross-domain task distribution efficiency, optimizes scheduling algorithms, and ensures the continuity and stability of the platform.
Smart Images

Figure CN120316014A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer high-availability testing, and specifically to a high-availability testing method and system for task scheduling of a computing power host management platform. Background Art
[0002] The "computing power host management platform" usually refers to a system used for centralized management and scheduling of high-performance computing resources (such as servers, GPU clusters, edge nodes, etc.). Its core goal is to improve resource utilization rate, simplify the complexity of operation and maintenance, and support the distribution and execution of large-scale computing tasks. The existing computing power host management platforms have the following pain points:
[0003] ①Insufficient testing in high-concurrency scenarios: Traditional testing schemes mostly use single-load simulation and cannot cover hybrid high-concurrency scenarios (such as sudden traffic, continuous high pressure, and superposition of multiple task types), resulting in risks such as resource competition and system crashes not being fully exposed.
[0004] ②Incomplete verification of cross-domain scheduling: The existing testing frameworks lack verification of the collaborative scheduling capabilities across regions and resource pools, and it is difficult to evaluate the dynamic resource allocation efficiency of intelligent scheduling algorithms in complex network environments.
[0005] ③Lack of verification of fault recovery and self-healing capabilities: The existing technologies have simple test scenario designs for failover, disaster recovery switching, and self-healing mechanisms, and do not simulate abnormal scenarios such as multi-node failures and resource isolation failures in real environments, resulting in insufficient verification of the system's disaster tolerance capabilities.
[0006] ④Weak data-driven optimization capabilities: Traditional test results fail to effectively quantify the performance indicators of scheduling algorithms and cannot provide multi-dimensional data support for algorithm optimization. Summary of the Invention
[0007] The technical task of the present invention is to provide a high-availability testing method and system for task scheduling of a computing power host management platform to solve the problems of insufficient test coverage in high-concurrency scenarios and incomplete verification of cross-domain scheduling, failover, and self-healing scenarios in the prior art.
[0008] The technical task of the present invention is realized in the following way. A high-availability testing method for task scheduling of a computing power host management platform is as follows:
[0009] Hybrid load testing: Verify the system's elastic scaling ability and performance stability by simulating long-term stable load and sudden traffic scenarios;
[0010] Cross-regional intelligent scheduling: Dynamically schedule tasks based on cost, latency, and regional policies to ensure optimal resource allocation and business continuity;
[0011] Resource failure migration and recovery: Simulate node-level or regional failures to verify the system's automatic detection, migration, and recovery capabilities;
[0012] Monitoring and data collection: Real-time collect transaction, resource, and stability metrics to provide multi-dimensional data support for verification;
[0013] Analysis and optimization: Identify bottlenecks based on test data to generate optimization suggestions and continuously improve the system's high availability.
[0014] Preferably, the mixed load test is as follows:
[0015] Identify key user roles and behavior patterns and design a mixing ratio that truly reflects the production environment;
[0016] Build a test environment similar to the production environment, prepare test data, consider data diversity, configure stress test tools and performance metric monitoring tools to capture comprehensive metrics;
[0017] Through script development and parameterization, achieve dynamic parameterization to simulate real changes, and set reasonable think times and paces;
[0018] Gradually increase the load to observe the system's behavior, perform stability tests (long-term mixed load) and peak tests (burst mixed tests);
[0019] Identify performance bottlenecks and optimize them.
[0020] More preferably, identifying key user roles and behavior patterns and designing a mixing ratio that truly reflects the production environment is as follows:
[0021] Long-term load: Simulate scenarios of long-term high-load operation to verify whether the system will experience performance degradation or crashes due to resource leakage or cache overflow issues;
[0022] Burst load: Simulate a large number of concurrent user requests to test whether the system's throughput, response time, and elastic scaling ability meet the declared values.
[0023] More preferably, identifying performance bottlenecks and optimizing them is as follows:
[0024] Judge whether the system can run continuously and stably, and the elastic scaling mechanism automatically adds resources when the system load is too high and releases resources when the load decreases;
[0025] Judge whether the system's throughput increases linearly with the increase in the number of concurrent users, and the average response time of users is less than or equal to 2 seconds;
[0026] Judge whether the key performance metrics of the system's CPU usage rate and memory occupancy exceed 80% utilization of the hardware resources and there are no abnormal fluctuations.
[0027] Preferably, the cross-region intelligent scheduling is as follows:
[0028] The cost-optimal strategy is specifically as follows: ① Configure dynamic scheduling tests: When the CPU utilization rate reaches a fixed threshold, automatically scale up or down; ② Simulate task calls to increase the CPU utilization rate of nodes; ③ When the system detects changes in node resources and reaches the expansion trigger condition, the algorithm task selects a suitable computing power node for expansion according to the optimal cost plan of the computing network, and the process does not affect business operations;
[0029] The minimum latency strategy: Specifically: ① Configure dynamic scheduling tests: When the application response latency reaches a fixed threshold, automatically scale up or down; ② Simulate task scheduling and perform operations to simulate increasing latency on the deployed computing power nodes; ③ When the strategy set threshold is reached, trigger the latency-sensitive scheduling mechanism, and the algorithm task selects a suitable computing power node for expansion according to the lowest latency plan of the computing network, and the process does not affect business operations;
[0030] The local priority strategy; Specifically: ① Configure dynamic scheduling tests: The local priority strategy, automatically scale up or down; ② Simulate task scheduling, and the initiator is a user in an undeployed area; ③ The algorithm task performs expansion according to the optimal cost plan of the computing network, and the process does not affect business operations.
[0031] Preferably, the resource failure migration and recovery is as follows:
[0032] Simulate continuous scheduling tasks, simulate node-level failures, and continuously observe the resources and system conditions of all nodes;
[0033] The load balancer can detect the failure status within a short time after a node failure and immediately start redistributing traffic to other normal nodes;
[0034] After the failed node automatically recovers, the load balancer can timely adjust the traffic distribution so that the system traffic reaches a balanced state again among all nodes, and the resource usage and performance indicators of each node smoothly transition to the normal level.
[0035] Preferably, during monitoring and data collection, multi-dimensional test indicators are monitored. During the execution of the test task, the execution status and resource utilization of the task are monitored in real time, or monitoring tools and log analysis are used to collect and analyze the data of task execution; among them, the data of task execution includes transaction indicators, resource indicators, execution logs of scheduling policies, and fault recovery records; transaction indicators include success rate, response time, and throughput; resource indicators include CPU, memory, disk I / O, and network; execution logs of scheduling policies include scheduling decisions, trigger conditions, and execution time consumption; fault recovery records include RTO, RPO, and recovery success rate.
[0036] Preferably, the analysis and optimization are as follows:
[0037] Data analysis: After aggregating and preprocessing the data during the testing process, data analysis and evaluation are carried out.
[0038] Resource bottleneck detection: Identify nodes with long-term high load (>80%) on CPU / memory / disk / network;
[0039] Evaluation of scheduling policy effectiveness: Verify whether the cost-optimal policy causes the latency of critical services to exceed the standard, and detect whether the local priority policy causes resource fragmentation;
[0040] Disaster recovery ability evaluation: Statistically calculate the success rate of failover (such as the RTO compliance rate);
[0041] Elastic scaling adjustment: Adjust the scaling threshold (such as the CPU trigger threshold from 70% to 65%) and recommend a more optimal instance specification (such as using GPU instances for compute-intensive tasks);
[0042] Dynamic weight adjustment: Increase the weight of the latency policy during peak business hours and switch to the cost-optimal policy during idle hours;
[0043] Cross-region scheduling improvement: Add edge nodes to reduce cross-region latency;
[0044] Fault recovery strategy optimization: Shorten the heartbeat detection interval (such as from 5 seconds to 2 seconds).
[0045] A high-availability test system for task scheduling of a computing power host management platform, which is used to implement the high-availability test method for task scheduling of the computing power host management platform as described above; the system includes:
[0046] Hybrid load test module, used to verify the elastic scaling and performance stability of the system by simulating long-term stable load and burst traffic scenarios;
[0047] Cross-region intelligent scheduling module, used to dynamically schedule tasks based on cost, latency, and regional policies to ensure optimal resource allocation and business continuity;
[0048] Resource fault migration and recovery module, used to simulate node-level or regional-level faults and verify the system's automatic detection, migration, and recovery capabilities;
[0049] Monitoring and data collection module, used to collect transaction, resource, and stability metrics in real time to provide multi-dimensional data support for verification;
[0050] Analysis and optimization module, used to identify bottlenecks based on test data, generate optimization suggestions, and continuously improve the high availability of the system.
[0051] The high-availability test method and system for task scheduling of the computing power host management platform of the present invention have the following advantages:
[0052] (1) Comprehensive coverage of complex scenarios: Through hybrid load and cross-domain testing, the present invention addresses the verification blind spots in high-concurrency and multi-fault scenarios in traditional solutions, with the test coverage rate increased by more than 40%.
[0053] (2) Enhancement of system stability: 1. The present invention improves the test coverage rate in high-concurrency scenarios, accurately exposes resource competition and scheduling failure problems, reduces the system crash risk by 60%, and decreases the task interruption rate by 35%.
[0054] (3) Optimization of resource utilization: After being verified through testing by the dynamic scheduling algorithm of the present invention, the resource fragmentation index is reduced by 25%, and the cross-domain task distribution efficiency is increased by 30%.
[0055] (4) Data-driven algorithm iteration: The present invention provides multi-dimensional performance indicators and deviation analysis, supporting a 50% shortening of the optimization cycle of the scheduling algorithm.
[0056] (5) The present invention simulates long-term stable load and burst traffic scenarios, monitors the threshold of system resource utilization, and combines multi-dimensional test indicators to achieve high-availability verification of the platform.
[0057] (6) The present invention dynamically executes the strategies of optimal cost, minimum delay, and local priority, injects node-level faults, and verifies the self-healing mechanism, verifying the dynamic resource allocation ability and fault recovery efficiency of the cross-domain scheduling strategy, and ensuring the continuity and stability of the computing power platform.
[0058] (7) The present invention constructs an intelligent strategy verification system, simulates complex fault scenarios to evaluate the system's self-healing ability, and provides a quantifiable decision-making basis for the optimization of the scheduling algorithm through data collection and analysis.
[0059] (8) Through links such as hybrid load simulation, cross-domain intelligent strategy verification, fault recovery and self-healing testing, the present invention conducts multi-dimensional tests (pressure, disaster tolerance, resource optimization), verifies the effectiveness of the intelligent scheduling algorithm (dynamic resource allocation ability, fault recovery and disaster tolerance ability), ensures the continuity of platform services and the stability of the system in high-concurrency scenarios, prevents system crash risks, avoids resource competition problems, and provides data support for the continuous optimization of the scheduling algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The present invention will be further described below in conjunction with the drawings.
[0061] Appendix Figure 1 It is a flowchart of the high-availability test method for task scheduling of the computing power host management platform. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The high-availability test method and system for task scheduling of the computing power host management platform of the present invention will be described in detail below with reference to the drawings of the specification and specific embodiments.
[0063] Example 1:
[0064] As shown in the Figure 1 appendix, this embodiment provides a high - availability test method for task scheduling of a computing power host management platform, and the method is as follows:
[0065] S1. Hybrid load test: By simulating long - term stable load and burst traffic scenarios, verify the system's elastic scaling ability and performance stability;
[0066] S2. Cross - region intelligent scheduling: Dynamically schedule tasks based on cost, latency, and regional policies to ensure optimal resource allocation and business continuity;
[0067] S3. Resource failure migration and recovery: Simulate node - level or regional - level failures to verify the system's automatic detection, migration, and recovery capabilities;
[0068] S4. Monitoring and data collection: Real - time collect transaction, resource, and stability metrics to provide multi - dimensional data support for verification;
[0069] S5. Analysis and optimization: Identify bottlenecks based on test data and generate optimization suggestions to continuously improve the system's high availability.
[0070] The hybrid load test in step S1 of this embodiment is specifically as follows:
[0071] S101. Identify key user roles and behavior patterns, and design a mixing ratio that truly reflects the production environment;
[0072] S102. Build a test environment similar to the production environment, prepare test data, consider data diversity, configure stress - testing tools and performance - metric monitoring tools to capture comprehensive metrics;
[0073] S103. Through script development and parameterization, achieve dynamic parameterization to simulate real - world changes, and set reasonable think times and paces;
[0074] S104. Gradually increase the load and observe the system behavior, perform stability tests (long - term hybrid load) and peak tests (burst hybrid tests);
[0075] S105. Identify performance bottlenecks and optimize them.
[0076] The identification of key user roles and behavior patterns in step S101 of this embodiment, and the design of a mixing ratio that truly reflects the production environment are specifically as follows:
[0077] S10101. Long - term load: Simulate a scenario of running at a high load for a long time to verify whether the system will experience performance degradation or crashes due to resource leakage or cache overflow problems;
[0078] S10101, Sudden Load: Simulate a large number of concurrent user requests to test whether the throughput, response time, and elastic scaling ability of the system meet the declared values.
[0079] The identification of performance bottlenecks and optimization in step S105 of this embodiment are specifically as follows:
[0080] S10501, Determine whether the system can operate continuously and stably. The elastic scaling mechanism automatically adds resources when the system load is too high and releases resources when the load decreases;
[0081] S10502, Determine whether the throughput of the system increases linearly with the increase in the number of concurrent users, and the average response time of users is less than or equal to 2 seconds;
[0082] S10503, Determine whether the key performance indicators of CPU usage and memory occupancy of the system exceed 80% utilization of hardware resources and there are no abnormal fluctuations.
[0083] The cross-regional intelligent scheduling in step S2 of this embodiment is specifically as follows:
[0084] S201, Cost-optimal strategy, specifically: ① Configure dynamic scheduling test: Automatically scale when the CPU utilization reaches a fixed threshold; ② Simulate task invocation to increase the CPU utilization of nodes; ③ When the system detects changes in node resources and reaches the expansion trigger condition, the algorithm task selects a suitable computing power node for expansion according to the cost-optimal solution of the computing network, and the process does not affect business operations;
[0085] S202, Lowest latency strategy: Specifically: ① Configure dynamic scheduling test: Automatically scale when the application response latency reaches a fixed threshold; ② Simulate task scheduling and perform an operation to simulate increasing the latency on the deployed computing power nodes; ③ When the policy set threshold is reached, trigger the latency-sensitive scheduling mechanism, and the algorithm task selects a suitable computing power node for expansion according to the lowest latency solution of the computing network, and the process does not affect business operations;
[0086] S203, Local priority strategy; specifically: ① Configure dynamic scheduling test: Local priority strategy, automatically scale; ② Simulate task scheduling, and the initiator is a user in an undeployed area; ③ The algorithm task expands according to the cost-optimal solution of the computing network, and the process does not affect business operations.
[0087] The resource fault migration and recovery in step S3 of this embodiment are specifically as follows:
[0088] S301, Simulate continuous scheduling tasks, simulate node-level faults, and continuously observe the resources and system conditions of all nodes;
[0089] S302. The load balancer can detect the fault status within a short time after a node fails and immediately start redistributing traffic to other normal nodes;
[0090] S303. After the faulty node automatically recovers, the load balancer can timely adjust the traffic distribution so that the system traffic reaches a balanced state again among all nodes, and the resource usage and performance metrics of each node smoothly transition to the normal level.
[0091] In the monitoring and data collection in step S4 of this embodiment, for multi-dimensional test metric monitoring, during the execution of the test task, the execution status and resource utilization of the task are monitored in real time, or monitoring tools and log analysis are used to collect and analyze the data of task execution; among them, the data of task execution includes transaction metrics, resource metrics, scheduling policy execution logs, and fault recovery records; transaction metrics include success rate, response time, and throughput; resource metrics include CPU, memory, disk I / O, and network; scheduling policy execution logs include scheduling decisions, triggering conditions, and execution time consumption; fault recovery records include RTO, RPO, and recovery success rate.
[0092] The analysis and optimization in step S5 of this embodiment are specifically as follows:
[0093] S501. Data analysis: After aggregating and preprocessing the data during the test process, data analysis and evaluation are performed.
[0094] S502. Resource bottleneck detection: Identify nodes with long-term high load (>80%) on CPU / memory / disk / network;
[0095] S503. Scheduling policy effectiveness evaluation: Verify whether the cost-optimal policy causes the key business delay to exceed the standard, and detect whether the local priority policy causes resource fragmentation;
[0096] S504. Disaster recovery ability evaluation: Statistically calculate the failure switch success rate (such as the RTO compliance rate);
[0097] S505. Elastic scaling adjustment: Adjust the expansion threshold (such as the CPU trigger threshold from 70% → 65%) and recommend a more optimal instance specification (such as changing to a GPU instance for compute-intensive tasks);
[0098] S506. Dynamic weight adjustment: Increase the weight of the delay policy during the business peak period and switch to the cost-optimal policy during the idle time;
[0099] S507. Cross-region scheduling improvement: Add edge nodes to reduce cross-region delay;
[0100] S508. Fault recovery policy optimization: Shorten the heartbeat detection interval (such as from 5 seconds → 2 seconds).
[0101] Embodiment 2:
[0102] This embodiment provides a high-availability test system for task scheduling of a computing power host management platform, which is used to implement the high-availability test method for task scheduling of the computing power host management platform as in Embodiment 1; the system includes:
[0103] A hybrid load test module, which is used to verify the system's elastic scaling and performance stability by simulating long-term stable load and burst traffic scenarios;
[0104] A cross-regional intelligent scheduling module, which is used to dynamically schedule tasks based on cost, latency, and regional policies to ensure optimal resource allocation and business continuity;
[0105] A resource failure migration and recovery module, which is used to simulate node-level or regional-level failures to verify the system's automatic detection, migration, and recovery capabilities;
[0106] A monitoring and data collection module, which is used to collect transaction, resource, and stability metrics in real time to provide multi-dimensional data support for verification;
[0107] An analysis and optimization module, which is used to identify bottlenecks based on test data, generate optimization suggestions, and continuously improve the high availability of the system.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A high-availability test method for task scheduling of a computing power host management platform, characterized in that, The method is as follows: Mixed load testing: By simulating long-term stable load and burst traffic scenarios, verify the system's elastic scaling ability and performance stability; Cross-region intelligent scheduling: Dynamically schedule tasks based on cost, latency, and regional policies to ensure optimal resource allocation and business continuity; Resource failure migration and recovery: Simulate node-level or regional-level failures to verify the system's automatic detection, migration, and recovery capabilities; Monitoring and data collection: Real-time collect transaction, resource, and stability metrics to provide multi-dimensional data support for verification; Analysis and optimization: Identify bottlenecks based on test data and generate optimization suggestions to continuously improve the system's high availability.
2. The high-availability test method for task scheduling of the computing power host management platform according to claim 1, wherein The specific steps of mixed load testing are as follows: Identify key user roles and behavior patterns, and design a mix ratio that truly reflects the production environment; Build a test environment similar to the production environment, prepare test data, consider data diversity, configure stress testing tools and performance metric monitoring tools to capture comprehensive metrics; Through script development and parameterization, achieve dynamic parameterization to simulate real changes, and set reasonable think times and paces; Gradually increase the load and observe the system's behavior, perform stability testing and peak testing; Identify performance bottlenecks and optimize them.
3. The high-availability test method for task scheduling of the computing power host management platform according to claim 2, wherein The specific steps of identifying key user roles and behavior patterns, and designing a mix ratio that truly reflects the production environment are as follows: Long-term load: Simulate a scenario of long-term high-load operation to verify whether the system will experience performance degradation or crashes due to resource leaks and cache overflows; Burst load: Simulate a large number of concurrent user requests to test whether the system's throughput, response time, and elastic scaling ability meet the claimed values.
4. The high-availability test method for task scheduling of the computing power host management platform according to claim 2, characterized in that, The specific steps of identifying performance bottlenecks and optimizing them are as follows: Judge whether the system can run continuously and stably. The elastic scaling mechanism automatically adds resources when the system load is too high and releases resources when the load decreases; Judge whether the system's throughput increases linearly with the increase in the number of concurrent users, and the average response time of users is less than or equal to 2 seconds; Judge whether the key performance metrics of the system's CPU utilization rate and memory occupancy exceed 80% utilization of the hardware resources, and there are no abnormal fluctuations.
5. The high-availability test method for task scheduling of the computing power host management platform according to claim 1, wherein The specific steps of cross-region intelligent scheduling are as follows: Cost-optimal strategy: Specifically, ① Configure dynamic scheduling tests: When the CPU utilization rate reaches a fixed threshold, automatically scale; ② Simulate task calls to increase the CPU utilization rate of nodes; ③ When the system detects changes in node resources and reaches the expansion trigger condition, the algorithm task selects a suitable computing power node for expansion according to the optimal cost plan of the computing network, and the process does not affect business operations; Lowest latency strategy: Specifically, ① Configure dynamic scheduling tests: When the application response latency reaches a fixed threshold, automatically scale; ② Simulate task scheduling and perform operations to increase the latency of the deployed computing power nodes; ③ When the strategy set threshold is reached, trigger the latency-sensitive scheduling mechanism, and the algorithm task selects a suitable computing power node for expansion according to the lowest latency plan of the computing network, and the process does not affect business operations; Local priority strategy; specifically: ① Configure dynamic scheduling test: local priority strategy, automatic scaling; ② Simulate task scheduling, and the initiator is a user in an undeployed area; ③ The algorithm task performs the expansion process according to the optimal solution of the computing network cost without affecting business operation.
6. The high-availability test method for task scheduling of the computing power host management platform according to claim 1, wherein The resource fault migration and recovery are as follows: Simulate continuous scheduling tasks, simulate node-level faults, and continuously observe the resources and system conditions of all nodes; The load balancer can detect the fault status within a short time after a node fails and immediately starts to redistribute the traffic to other normal nodes; After the faulty node automatically recovers, the load balancer can timely adjust the traffic distribution so that the system traffic reaches a balanced state among all nodes again, and the resource usage and performance indicators of each node smoothly transition to the normal level.
7. The high-availability test method for task scheduling of the computing power host management platform according to claim 1, characterized in that During monitoring and data collection, multi-dimensional test indicators are monitored. During the execution of the test task, the execution status and resource utilization of the task are monitored in real time, or monitoring tools and log analysis are used to collect and analyze the data of task execution; among them, the data of task execution includes transaction indicators, resource indicators, execution logs of scheduling policies, and fault recovery records; transaction indicators include success rate, response time, and throughput; resource indicators include CPU, memory, disk I / O, and network; execution logs of scheduling policies include scheduling decisions, triggering conditions, and execution time consumption; fault recovery records include RTO, RPO, and recovery success rate.
8. The high-availability test method for task scheduling of the computing power host management platform according to claim 1, wherein The analysis and optimization are as follows: Data analysis: After aggregating and preprocessing the data during the test process, perform data analysis and evaluation. Resource bottleneck detection: Identify nodes with long-term high loads on CPU / memory / disk / network; Evaluation of the effectiveness of the scheduling policy: Verify whether the cost-optimal policy causes the key business delay to exceed the standard, and detect whether the local priority policy causes resource fragmentation; Evaluation of the disaster recovery and recovery ability: Statistically calculate the success rate of fault switching; Elastic scaling adjustment: Adjust the expansion threshold and recommend a better instance specification; Dynamic weight adjustment: Increase the weight of the delay policy during peak business hours and switch to the cost-optimal policy during idle hours; Improvement of cross-regional scheduling: Add edge nodes to reduce cross-regional delay; Optimization of the fault recovery policy: Shorten the heartbeat detection interval.
9. A high-availability test system for task scheduling of a computing power host management platform, characterized in that, This system is used to implement the high-availability test method for task scheduling of the computing power host management platform described in any one of claims 1-8; this system includes: A hybrid load test module, which is used to verify the elastic scaling ability and performance stability of the system by simulating long-term stable load and burst traffic scenarios; A cross-regional intelligent scheduling module, which is used to dynamically schedule tasks based on cost, delay, and regional policies to ensure optimal resource allocation and business continuity; A resource fault migration and recovery module, which is used to simulate node-level or regional-level faults and verify the system's automatic detection, migration, and recovery capabilities; A monitoring and data collection module, which is used to collect transaction, resource, and stability indicators in real time to provide multi-dimensional data support for verification; An analysis and optimization module, which is used to identify bottlenecks based on test data, generate optimization suggestions, and continuously improve the high availability of the system.
Citation Information
Cited By
Server test method, electronic equipment, storage medium and program product
CN120639659A
Resource allocation method and device of server
CN120723484A
Software performance load test method and system for large-scale video conference
CN120872493A
Software performance load testing method and system for large-scale video conference
CN120872493B
GPU virtualization system based on CUDA cross-level translation and multi-pooling scheduling
CN121143952A