Distributed test resource dynamic allocation and management system and method
By designing a distributed testing resource dynamic allocation and management system, the problem of unbalanced resource allocation in software testing is solved, efficient resource utilization and automated management are achieved, and the flexibility and reliability of the system are improved.
Patent Information
- Application Number
- CN202510110362.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology has uneven resource allocation in software testing, resulting in some resources being idle and other resources being overloaded, resulting in resource waste.
Design a distributed testing resource dynamic allocation and management system, including resource monitoring and collection modules, resource scheduling modules, resource prediction and optimization modules, and automated decision-making and execution modules. By monitoring resource status in real time, generating resource scheduling strategies, predicting resource requirements, and automatically allocating test tasks to achieve dynamic resource adjustment.
It improves resource utilization, avoids resource waste, enhances system flexibility and adaptability, realizes automated resource management, improves fault handling efficiency, and realizes real-time monitoring and prediction.
Smart Images

Figure CN119988025A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer software testing, and in particular relates to a distributed test resource dynamic allocation and management system and method. Background Art
[0002] Software testing is a very important part of the software system and one of the most effective ways to evaluate software quality. With the rapid development of the Internet, the iteration cycle of many software has been shortened, but testing is very time-consuming. Many software systems are put online without repeated testing, resulting in slow response and even server paralysis.
[0003] To ensure the correctness of the test indicators, the efficiency of the test is particularly important. During the test, we must first ensure that the hardware configuration of the test machine can support the test, so that the software system can be tested normally and the correct results can be obtained. In actual work, the test cost must be considered, so the available hardware resources are limited, so it is very important to reasonably allocate the use of these resources.
[0004] Existing technologies improve resource utilization through technologies such as automated scheduling, containerization, and virtualization. However, these operations generally result in unbalanced or static resource allocation, which causes some resources to be idle while other resources may be overloaded, resulting in resource waste. Summary of the invention
[0005] In order to solve the problems of complex structure, single point failure and the like in the existing resource allocation scheme, the present invention provides a distributed test resource dynamic allocation and management system.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: A distributed test resource dynamic allocation and management system, comprising: Resource monitoring and collection module, used to monitor and collect indicator data of each node of the test machine; Resource scheduling module, used to generate resource scheduling strategies according to task attributes and resource status of each node of the test machine; The resource prediction and optimization module is used to establish a linear regression model of historical data of resource demand, and use the machine learning library to train the linear regression model of historical data, and predict new test data based on the trained linear regression model of historical data; the new test data includes test cases and executable scripts; The automated decision-making and execution module is used to receive the indicator data of each node of the test machine fed back by the resource monitoring and collection module, the resource scheduling strategy generated by the resource scheduling module, and the new data predicted by the resource prediction and optimization module, generate test tasks based on the received data, and assign the test tasks to the corresponding test nodes.
[0007] Preferably, the resource monitoring and collection module includes a monitoring agent and an indicator collector. The indicator collector is installed on the central monitoring server, and the monitoring agent is installed and configured on each test node. The indicator collector is used to regularly extract indicator data from the monitoring agent or the agent node operating system and send it to the central monitoring server; the indicator data includes CPU usage, memory usage, disk usage, disk IO and network bandwidth usage; The monitoring agent is used to regularly pull the service endpoints that expose the indicator data through the HTTP protocol, and store the indicator data extracted by the indicator collector in the local time series database. The indicator collector is also responsible for extracting the indicator data monitored by Prometheus and drawing charts based on the extracted data.
[0008] Preferably, the indicator collector is a Grafana indicator collector, and the monitoring agent is a Prometheus monitoring agent.
[0009] Preferably, the service endpoints include applications, databases and network devices.
[0010] Preferably, the resource scheduling module includes a resource allocator Kubernetes and a task scheduling algorithm; The task scheduling algorithm is the shortest job first algorithm, which is used in conjunction with the Redis database to implement priority queue strategy comprehensive task scheduling; according to the task attributes and node resource conditions, the resource allocator Kubernetes allocates tasks to nodes and monitors the execution status of tasks; in Redis, the ordered set Sorted Set is used to implement the priority queue, in which each member in the set has an associated score score, and the members are sorted according to the size of the score; The resource allocator Kubernetes includes multiple nodes, one of which is called the Master node. The Master manages and controls the state and behavior of the cluster through the API Server, controller and scheduler, thereby managing the state of the entire cluster and controlling its behavior.
[0011] Preferably, the resource prediction and optimization module is used to establish a linear regression model of historical data of resource demand using Python and Scikit-learn library, and train the linear regression model of historical data using a machine learning library, and predict new test data based on the trained linear regression model of historical data.
[0012] Preferably, the training process of the linear regression model of the historical data uses the least squares method to solve the parameters; The new test data is predicted based on the linear regression model of the trained historical data, specifically by substituting the new independent variable value into the linear regression equation to calculate the predicted value of the dependent variable.
[0013] Preferably, the automated decision and execution module includes a decision engine, an automated executor and an event driver; the decision engine configures a Drools rule engine, defines resource allocation and scheduling rules, and integrates with the system, configures an Airflow task flow, and executes resource allocation and scheduling operations based on the decision of the rule engine; the automated executor integrates a Kafka message queue service, sends system events to a message queue, and triggers the automated executor to perform corresponding operations; the event driver is used to detect system failures and send corresponding messages based on the detection results.
[0014] The present invention also provides a distributed test resource dynamic allocation and management method, comprising the following steps: Monitor and collect indicator data of each node of the test machine; Generate resource scheduling strategies based on task attributes and resource status of each node of the test machine; Establish a linear regression model for the historical data of resource requirements, and use the machine learning library to train the linear regression model of the historical data. Use the trained linear regression model of the historical data to predict new test data; the new test data includes test cases and executable scripts. Generate test tasks based on indicator data, resource scheduling strategies and predicted new data, and assign test tasks to corresponding test nodes.
[0015] The distributed test resource dynamic allocation and management system provided by the present invention has the following beneficial effects: The present invention provides indicator data of test nodes by a resource monitoring and collection module, and generates a resource scheduling strategy according to task attributes and resource status of each node of the test machine through a resource scheduling module. At the same time, the new data predicted by the resource prediction and optimization module is combined with historical data and new data, and resource allocation is dynamically adjusted according to actual needs, which can maximize resource utilization and avoid waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 A schematic diagram of a distributed test resource dynamic allocation and management system provided in Example 1 of the present invention; Figure 2 This is a flow chart of the distributed test resource dynamic allocation and management method provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.
[0019] Example 1 The present invention studies how to effectively and dynamically allocate and manage resources in a distributed test system, including computing resources, storage resources and network bandwidth, so as to optimize the execution efficiency and resource utilization of test tasks.
[0020] Specifically, the present invention provides a distributed test resource dynamic allocation and management system, such as Figure 1 As shown, it includes a resource monitoring and collection module, a resource scheduling module, a resource prediction and optimization module, and an automated decision-making and execution module.
[0021] The resource monitoring and collection module is used to monitor and collect the index data of each node of the test machine, and draw charts for users to view based on the index data; the resource scheduling module is used to generate resource scheduling strategies based on task attributes and resource status of each node of the test machine; the resource prediction and optimization module is used to establish a linear regression model of historical data of resource demand, and use the machine learning library to train the linear regression model of historical data, and predict new test data based on the trained linear regression model of historical data; resource prediction is to predict the required resources for new test cases and executable scripts. The automated decision-making and execution module is used to receive the index data of each node of the test machine fed back by the resource monitoring and collection module, the resource scheduling strategy generated by the resource scheduling module, and the new data predicted by the resource prediction and optimization module, generate test tasks based on the received data, and assign the test tasks to the corresponding test nodes.
[0022] Specifically, the resource monitoring and collection module consists of a monitoring agent and an indicator collector. Install the Grafana indicator collector on the central monitoring server, install and configure the Prometheus monitoring agent on each test node, and use the Grafana indicator collector to regularly extract indicator data such as CPU usage, memory usage, disk usage, disk IO, and network bandwidth usage from the monitoring agent or agent node operating system and send it to the central monitoring server.
[0023] Prometheus regularly pulls service endpoints that expose indicator data, such as applications, databases, network devices, etc., through the HTTP protocol, and then stores this data in a local time series database. Grafana is responsible for extracting monitoring data (i.e. indicator data) from Prometheus and charting the extracted data to facilitate viewing the status of each node.
[0024] The resource scheduling module consists of the resource allocator Kubernetes, task scheduling algorithm, and priority queue management. The shortest job first algorithm is used in conjunction with the Redis database to implement priority queue strategy comprehensive task scheduling. According to the task attributes and node resource conditions, the resource allocator Kubernetes allocates tasks to nodes and monitors the execution status of tasks to ensure that high-priority tasks are executed first. The task scheduling algorithm is a set of strategies and rules that determine which tasks should be executed when, where, and by whom in a computing system. These algorithms usually take into account factors such as task priority, resource availability, dependencies between tasks, and system load conditions. The shortest job first algorithm is a strategy in the task scheduling algorithm, and the two are inclusive. The shortest job algorithm sorts the tasks in the task queue according to the estimated execution time, and then prioritizes the tasks with the shortest execution time. This can minimize the average waiting time.
[0025] A Kubernetes cluster usually consists of multiple nodes, one of which is called the Master node. The Master manages and controls the state and behavior of the cluster through components such as the API Server, controller, and scheduler, thereby managing the state of the entire cluster and controlling its behavior.
[0026] In Redis, you can use a sorted set to implement a priority queue, where each member in the set has an associated score, and the members are sorted according to the size of the score. You can use the command ZADD to add a task to an ordered set, where the score represents the priority of the task and the member represents the content of the task. You can use the command ZPOPMAX to obtain and remove the member with the highest score in the ordered set (that is, the task with the highest priority). You can use the command ZRANGEBYSCORE to obtain the members within a specified score range in the ordered set (that is, a list of tasks within a specified priority range). You can use the command ZREM to remove a specified member from the ordered set (that is, remove a specified task). The purpose is to take into account the monitoring data of the test machine collected by the resource monitoring and collection modules in the priority queue strategy for comprehensive queue sorting.
[0027] The resource prediction and optimization module is an additional function that can make a prediction for new test data (test cases, executable scripts) to achieve the purpose of quickly allocating resources. This module consists of machine learning models. Use Python and Scikit-learn and other libraries to build a linear regression model of historical data of resource requirements, and use machine learning libraries to train the model. When training the model, optimize the parameters of the model through differential verification technology, integrate the trained model into the system, and optimize resource allocation based on the prediction results.
[0028] The training process of the linear regression model usually uses the least squares method to solve the parameters. The goal of the least squares method is to minimize the sum of squared residuals between the observed data points and the linear model's predicted values. By minimizing the sum of squared residuals, the parameter values that best fit the model to the data can be obtained.
[0029] Several major advantages of the linear regression model are simplicity and easy to understand: Linear regression is a simple model that is easy to understand and explain. Its basic principle is to find a linear relationship between the independent variable and the dependent variable, which makes the model highly interpretable; high computational efficiency: the amount of computation required to train a linear regression model is relatively small, and the method for solving the parameters is usually an analytical solution or an iterative method, which has a fast calculation speed and is suitable for large-scale data sets; strong generalization ability: under appropriate circumstances, the linear regression model can achieve good generalization performance, that is, good predictive ability for new data; strong interpretability: since the coefficients of the linear regression model represent the degree of influence of the independent variable on the dependent variable, the results of the model are very easy to explain. This allows decision makers to make meaningful decisions based on the model results; robust to outliers: The linear regression model has a relatively small impact on outliers because it fits the data by minimizing the residual sum of squares, rather than absolutely fitting each data point.
[0030] Once training is complete, the trained model can be used to predict new data. The prediction process is to substitute the new independent variable value into the linear regression equation and calculate the predicted value of the dependent variable. The new data here refers to new test cases, test scripts, etc. The test data predicted by the model will no longer be involved in the scheduling. Automated decision-making and execution will directly allocate resources based on the predicted value.
[0031] The automated decision and execution module is the leading program for resource allocation. This module distributes tasks (i.e., distributes them to the corresponding test nodes). This module consists of a decision engine, an automated executor, and an event driver. Configure the Drools rule engine, define resource allocation and scheduling rules, and integrate with the system. Configure the Airflow task flow to execute resource allocation and scheduling operations based on the rule engine decision. Integrate message queue services such as Kafka, send system events to the message queue, and trigger the automated executor to perform corresponding operations. The event driver is a fault tolerance of the system itself. If any module of the system itself has a problem or is offline, a corresponding message will be issued.
[0032] Drools is a rule-based engine that uses the RETE algorithm for rule matching. The RETE algorithm is an efficient pattern matching algorithm that can quickly match rules to determine which rules apply to a given fact. The RETE algorithm decomposes the rule condition part into multiple nodes, each of which represents a condition test. When a new fact is entered into the system, starting from the top of the network, the fact passes through the nodes layer by layer, and each node is checked according to the attributes or relationships of the fact. The connection between the nodes ensures that only facts that meet all conditions can activate the relevant rules. In layman's terms, it is to avoid a piece of test data being sent to different test nodes at the same time, resulting in repeated testing.
[0033] Use the Airflow scheduler to manage the execution of tasks. The scheduler is responsible for determining the execution order of tasks based on the dependencies between tasks and assigning tasks to distributed task queues.
[0034] Kafka is a powerful and highly scalable distributed stream processing platform for processing real-time data streams and provides features such as high throughput, persistence, and reliability.
[0035] Each module in the above-mentioned distributed test resource dynamic allocation and management system can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0036] Based on the same inventive concept, the present invention also provides a distributed test resource dynamic allocation and management method, specifically as follows Figure 2 As shown, the following steps are included: Step 1: Monitor and collect the indicator data of each node of the test machine, and draw a chart for the user to view based on the indicator data.
[0037] Step 2: Generate a resource scheduling strategy based on the task attributes and the resource status of each node of the test machine.
[0038] Step 3: Establish a linear regression model of the historical data of resource demand, and use the machine learning library to train the linear regression model of the historical data, and predict the new test data based on the trained linear regression model of the historical data; wherein the new test data includes test cases and executable scripts.
[0039] Step 4: Generate test tasks based on indicator data, resource scheduling strategies and predicted new data, and assign the test tasks to corresponding test nodes.
[0040] The distributed test resource dynamic allocation and management system provided by the present invention has the following advantages: Solve the problem of low resource utilization in the prior art: the prior art has the situation of unbalanced or static resource allocation, which causes some resources to be idle while other resources may be overloaded. The present invention can maximize the resource utilization and avoid resource waste by monitoring the resource utilization of the system in real time and dynamically adjusting resource allocation according to actual needs.
[0041] Solve the problem of lack of flexibility and adaptability in the prior art: the prior art lacks a flexible resource management mechanism and cannot dynamically adjust resource allocation according to actual needs. The present invention introduces event-driven technology to make the system more flexible and adaptable. The system can adopt different response strategies according to different event types and dynamically adjust resource allocation to meet actual needs.
[0042] Solve the problem of cumbersome manual operation in the prior art: the prior art requires manual intervention for resource management, which is cumbersome and error-prone. The present invention can reduce the need for manual intervention and improve the efficiency and reliability of resource management by realizing automated resource management. The system can automatically monitor and adjust resource allocation according to pre-set rules and policies, reducing the workload of administrators.
[0043] Solve the problem of untimely fault handling in the prior art: The prior art has the problem of untimely fault handling, which leads to system downtime or performance degradation. The present invention can timely discover and handle resource faults by real-time monitoring of the system resource status and introducing fault detection and automated fault handling mechanisms, thereby reducing system downtime and improving system availability and reliability.
[0044] Solve the problem of lack of real-time monitoring and prediction in the prior art: The prior art lacks the function of real-time monitoring of system resource utilization and resource demand prediction based on historical data. The present invention introduces event-driven technology to monitor the resource status of the system in real time and predict resource demand based on historical data, so as to discover potential resource bottlenecks or performance problems in advance and take preventive measures to avoid system failures or performance degradation.
[0045] In summary, the present invention can improve resource utilization, enhance system flexibility and adaptability, realize automated resource management, improve fault handling efficiency, and achieve beneficial effects such as real-time monitoring and prediction.
[0046] It should be understood by those skilled in the art that embodiments of the present invention may provide methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0048] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0050] It should be pointed out that the specific implementation methods described above can enable those skilled in the art to understand the invention more comprehensively, but do not limit the invention in any way. Therefore, although the invention has been described in detail in this specification and embodiments, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved. Any simple change or equivalent replacement of the technical solution that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belongs to the protection scope of the present invention.
Claims
1. A distributed test resource dynamic allocation and management system, characterized in that: include: Resource monitoring and collection module, used to monitor and collect indicator data of each node of the test machine; Resource scheduling module, used to generate resource scheduling strategies according to task attributes and resource status of each node of the test machine; The resource prediction and optimization module is used to establish a linear regression model of historical data of resource demand, and use the machine learning library to train the linear regression model of historical data, and predict new test data based on the trained linear regression model of historical data; the new test data includes test cases and executable scripts; The automated decision-making and execution module is used to receive the indicator data of each node of the test machine fed back by the resource monitoring and collection module, the resource scheduling strategy generated by the resource scheduling module, and the new data predicted by the resource prediction and optimization module, generate test tasks based on the received data, and assign the test tasks to the corresponding test nodes.
2. The distributed test resource dynamic allocation and management system according to claim 1, characterized in that: The resource monitoring and collection module includes a monitoring agent and an indicator collector. The indicator collector is installed on the central monitoring server, and the monitoring agent is installed and configured on each test node. The indicator collector is used to regularly extract indicator data from the monitoring agent or the agent node operating system and send it to the central monitoring server; the indicator data includes CPU usage, memory usage, disk usage, disk IO and network bandwidth usage; The monitoring agent is used to regularly pull the service endpoints that expose the indicator data through the HTTP protocol, and store the indicator data extracted by the indicator collector in the local time series database. The indicator collector is also responsible for extracting the indicator data monitored by Prometheus and drawing charts based on the extracted data.
3. The distributed test resource dynamic allocation and management system according to claim 2, characterized in that: The indicator collector is a Grafana indicator collector, and the monitoring agent is a Prometheus monitoring agent.
4. The distributed test resource dynamic allocation and management system according to claim 2, characterized in that: The service endpoints include applications, databases, and network devices.
5. The distributed test resource dynamic allocation and management system according to claim 2, characterized in that: The resource scheduling module includes a resource allocator Kubernetes and a task scheduling algorithm; The task scheduling algorithm is the shortest job first algorithm, which is used in conjunction with the Redis database to implement priority queue strategy comprehensive task scheduling; according to the task attributes and node resource conditions, the resource allocator Kubernetes allocates tasks to nodes and monitors the execution status of tasks; in Redis, the ordered set Sorted Set is used to implement the priority queue, in which each member in the set has an associated score score, and the members are sorted according to the size of the score; The resource allocator Kubernetes includes multiple nodes, one of which is called the Master node. The Master manages and controls the state and behavior of the cluster through the API Server, controller and scheduler, thereby managing the state of the entire cluster and controlling its behavior.
6. The distributed test resource dynamic allocation and management system according to claim 5, characterized in that: The resource prediction and optimization module is used to establish a linear regression model of historical data of resource demand using Python and Scikit-learn library, and to train the linear regression model of historical data using machine learning library, and to predict new test data based on the trained linear regression model of historical data.
7. The distributed test resource dynamic allocation and management system according to claim 6, characterized in that: The training process of the linear regression model of the historical data uses the least squares method to solve the parameters; The new test data is predicted based on the linear regression model of the trained historical data, specifically by substituting the new independent variable value into the linear regression equation to calculate the predicted value of the dependent variable.
8. The distributed test resource dynamic allocation and management system according to claim 6, characterized in that: The automated decision and execution module includes a decision engine, an automated executor and an event driver; the decision engine configures the Drools rule engine, defines resource allocation and scheduling rules, and integrates with the system, configures the Airflow task flow, and executes resource allocation and scheduling operations based on the rule engine decision; the automated executor integrates the Kafka message queue service, sends system events to the message queue, and triggers the automated executor to perform corresponding operations; the event driver is used to detect system failures and send corresponding messages based on the detection results.
9. A distributed test resource dynamic allocation and management method, characterized in that: The following steps are involved: Monitor and collect indicator data of each node of the test machine; Generate resource scheduling strategies based on task attributes and resource status of each node of the test machine; Establish a linear regression model for the historical data of resource requirements, and use the machine learning library to train the linear regression model of the historical data. Use the trained linear regression model of the historical data to predict new test data; the new test data includes test cases and executable scripts. Generate test tasks based on indicator data, resource scheduling strategies and predicted new data, and assign the test tasks to corresponding test nodes.