Performance tuning method of cluster operating system, electronic equipment and storage medium
By acquiring and simulating the allocation and acquisition task collection set in a hyper-large-scale cluster operating system, predicting execution time and completion time, and optimizing task allocation strategy, the problem of real-time processing delay of monitoring data is solved, and system performance and equipment operation efficiency are improved.
Patent Information
- Application Number
- CN202510734657.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the super-large-scale cluster operating system, the real-time processing of monitoring data and the delay in decision-making responses, resulting in untimely performance monitoring, affecting operation and maintenance efficiency and effect.
By obtaining the collection task collection and simulating the allocation to the management terminal, the execution time and completion time of the performance acquisition task are predicted, the performance acquisition task is allocated based on the iterative optimization task allocation strategy, and the performance status of the equipment is analyzed and the performance status is tuned.
It improves the allocation efficiency and accuracy of performance acquisition tasks, reduces task waiting time, and improves the overall performance of the management terminal cluster and the operation efficiency and stability of the collected equipment.
Smart Images

Figure CN120276865A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method for optimizing the performance of a cluster operating system, an electronic device, and a storage medium. Background Art
[0002] In recent years, with the rapid development of cloud computing and distributed system architectures, large-scale server clusters have been widely used in key fields such as data centers, financial trading platforms, and the Internet of Things. To ensure the stable operation of the system, the performance tuning system of modern cluster operating systems needs to perform full-stack performance monitoring on underlying hardware resources (including CPUs, memory, disk arrays, network devices), the operating state of the operating system, and upper-layer service components (such as database systems, distributed storage systems like Ceph, etc.). Traditional architectures usually adopt a centralized management mode, where a single management cluster is uniformly responsible for multi-dimensional metric collection (covering real-time metrics such as CPU utilization, memory page error rate, disk IOPS, network packet loss rate, etc.), synchronization of software and hardware configuration information, and handling of exception alarms.
[0003] However, with the exponential growth of the cluster scale and the improvement of the refinement degree of monitoring metrics (for example, a Ceph storage cluster needs to synchronously monitor dozens of sub-metrics such as OSD status, PG distribution, and RADOS layer latency), the existing technology has exposed serious architectural defects: while the management cluster undertakes core responsibilities such as regular task scheduling and resource allocation, it also needs to process real-time aggregation and analysis of massive monitoring data, resulting in double occupation of its computing resources and network bandwidth. Especially in scenarios of sudden load surges (such as cascading monitoring alarms caused by a sharp increase in distributed transactions), the processing delay of the management node for monitoring data will exceed the metric collection interval period, forming a continuous data lag effect, resulting in untimely performance monitoring of the cluster and affecting the operation and maintenance efficiency and effect of the cluster operating system.
[0004] Therefore, there is an urgent need to develop a method that can timely achieve real-time processing and decision response for monitoring data streams in ultra-large-scale cluster operating systems to timely improve the overall performance and resource utilization rate of the system. Summary of the Invention
[0005] The purpose of this application is to provide a method for optimizing the performance of a cluster operating system, an electronic device, and a storage medium to solve the above problems.
[0006] To achieve the above purpose, in the first aspect, this application proposes a method for optimizing the performance of a cluster operating system, and the method includes: Obtain a collection of acquisition tasks, where the collection of acquisition tasks is a collection formed by periodic task subsets for N devices to be acquired on which a cluster operating system is deployed. The The performance collection task is the j-th performance collection task for the i-th device to be collected; Simulate and allocate the collection task set to M management terminals, and predict the simulated execution time and simulated completion time of each performance collection task according to the task generation time and simulated allocation result of each performance collection task; Iteratively optimize the simulated allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and perform performance collection task allocation according to the optimized task simulated allocation strategy; Analyze the device performance status of each device to be collected based on the performance index data obtained by each management terminal executing the allocated performance collection task, and perform performance tuning on the device to be collected according to the device performance status.
[0007] In some embodiments, the iteratively optimizing the simulated allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and performing performance collection task allocation according to the optimized task simulated allocation strategy includes: Calculate the task completion duration and task execution duration corresponding to the performance collection task according to the task generation time, simulated execution time, and simulated completion time of each performance collection task respectively; Regard the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution cycle as tasks to be transferred out, and allocate the tasks to be transferred out to one of the terminals to be transferred in. The terminal to be transferred in is a management terminal among the M management terminals, excluding the scheduling terminal that simulates the execution of the task to be transferred out under the current simulated allocation strategy.
[0008] In some embodiments, the allocating the task to be transferred out to one of the terminals to be transferred in includes: Calculate the comprehensive turnover degree corresponding to the entire collection task set when each terminal to be transferred in simulates the execution of the task to be transferred out. The comprehensive turnover degree is the weighted sum of the individual turnover degrees of each performance collection task, and the individual turnover degree is the quotient of the task completion duration and task execution duration of the corresponding performance collection task; Regard the terminal to be transferred in corresponding to the simulated allocation strategy with the minimum comprehensive turnover degree as the target transfer terminal, and allocate the task to be transferred out to the target transfer terminal.
[0009] In some embodiments, the allocating the task to be transferred out to one of the terminals to be transferred in includes: Calculate the comprehensive turnover of the entire collection task set when each terminal to be transferred simulates the execution of the task to be retrieved and the associated tasks. The comprehensive turnover is the weighted sum of the individual turnovers of each performance collection task. The individual turnover is the quotient of the task completion duration and the task execution duration of the corresponding performance collection task. The associated tasks are performance collection tasks that belong to the same task subset as the task to be retrieved. Use the terminal to be transferred corresponding to the simulation allocation strategy with the smallest comprehensive turnover as the target transfer terminal, and allocate the task to be retrieved and the associated tasks to the target transfer terminal.
[0010] In some embodiments, the step of allocating the task to be retrieved to one of the terminals to be transferred includes: Determine a target transfer terminal from multiple terminals to be transferred, and calculate the first cumulative increase in the simulated execution duration and the second cumulative increase in the simulated execution duration of the target transfer terminal respectively. When the second cumulative increase in the simulated execution duration is greater than the first cumulative increase in the simulated execution duration, only allocate the task to be retrieved to the target transfer terminal. When the second cumulative increase in the simulated execution duration is less than or equal to the first cumulative increase in the simulated execution duration, allocate both the task to be retrieved and the associated tasks to the target transfer terminal. The associated tasks are performance collection tasks that belong to the same task subset as the task to be retrieved. The first cumulative increase in the simulated execution duration is the weighted cumulative increase in the simulated execution duration of the performance collection tasks originally allocated by the target transfer terminal after only adding the task to be retrieved. The second cumulative increase in the simulated execution duration is the cumulative increase in the simulated execution duration of the performance collection tasks originally allocated by the target transfer terminal after adding both the task to be retrieved and the associated tasks.
[0011] In some embodiments, the step of iteratively optimizing the simulation allocation strategy based on the execution cycle, task generation time, simulation execution time, and simulation completion time of each performance collection task, and performing performance collection task allocation according to the optimized task simulation allocation strategy includes: Calculate the task completion duration and the task execution duration of each performance collection task according to the task generation time, simulation execution time, and simulation completion time of each performance collection task respectively. Use the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution cycle as the tasks to be retrieved. Calculate the influence degree of each performance collection task simulated and executed by the terminal to be scheduled on the task to be retrieved, and use the preset number of performance collection tasks with the highest influence degree as the tasks to be transferred out. Assign the task to be transferred out to one of the terminals to be transferred in, where the terminal to be transferred in is a management terminal among the M management terminals, excluding the terminal to be scheduled that simulates the execution of the task to be transferred out under the current simulated allocation policy.
[0012] In some embodiments, the performance tuning of the device to be collected according to the device performance status includes: Determine the device to be collected with abnormal device performance status as the device to be processed; Obtain the task execution information of the device to be processed, and screen out the tasks corresponding to the abnormal type of the device performance status from the task execution information as the tasks to be processed; Perform performance tuning on the device to be processed according to several resource requirement combinations of the tasks to be processed.
[0013] In some embodiments, the performance tuning of the device to be processed according to several resource requirement combinations of the tasks to be processed includes: Judge whether the performance index data associated with the device to be processed meets at least one of the several resource requirement combinations; If so, based on the satisfied resource requirement combination, perform resource reallocation of the device to be processed for the task to be processed; If not, determine the device to be collected that meets at least one resource requirement combination according to the performance index data of other devices to be collected as the standby device, so as to transfer the task to be processed to the standby device.
[0014] In a second aspect, the present application proposes an electronic device, including: One or more processors; A memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the performance tuning method of the cluster operating system as described above.
[0015] In a third aspect, the present application proposes a storage medium, the storage medium stores executable instructions, and when the instructions are executed by a processor, the processor executes the performance tuning method of the cluster operating system as described above.
[0016] Compared with the prior art, the beneficial effects of the present application include: In the first aspect, by obtaining a set of acquisition tasks formed from a subset of periodic tasks for the device to be acquired and simulating their allocation to the management terminal, it is possible to predict in advance the simulated execution time and simulated completion time of each performance acquisition task. This helps to reasonably plan the allocation of performance acquisition tasks, avoid problems such as task backlog, data lag, and untimely performance monitoring, and improve the allocation efficiency and accuracy of performance acquisition tasks. In the second aspect, by iteratively optimizing the simulated allocation strategy, it is possible to continuously adjust the task allocation strategy to make it more adaptable to the actual operating environment, reduce the waiting time and execution time of tasks, and improve the overall performance of the management terminal cluster. In the third aspect, by analyzing the performance metric data obtained from each management terminal executing the allocated performance acquisition tasks, it is possible to accurately evaluate the device performance status of each device to be acquired. Based on this detailed performance status data, it is possible to specifically optimize the performance of the device to be acquired and improve the operating efficiency and stability of the device to be acquired. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and should not be regarded as limiting the scope of the present application.
[0018] Figure 1 FIG. is an application scenario diagram of a method for optimizing the performance of a cluster operating system in an embodiment; Figure 2 FIG. is a system architecture diagram of a method for optimizing the performance of a cluster operating system in an embodiment; Figure 3 FIG. is a flowchart showing the process of a method for optimizing the performance of a cluster operating system in an embodiment; Figure 4 FIG. is a flowchart showing the process of iteratively optimizing a simulated allocation strategy based on the execution period, task generation time, simulated execution time, and simulated completion time of each performance acquisition task, and allocating performance acquisition tasks according to the optimized task simulated allocation strategy in an embodiment; Figure 5 FIG. is a flowchart showing the process of iteratively optimizing a simulated allocation strategy based on the execution period, task generation time, simulated execution time, and simulated completion time of each performance acquisition task, and allocating performance acquisition tasks according to the optimized task simulated allocation strategy in another embodiment; Figure 6 FIG. is a schematic structural diagram of an electronic device involved in the method for optimizing the performance of a cluster operating system in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0021] For example, terms such as "first" and "second" used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of this application, the first execution comprehensive duration may be referred to as the second execution comprehensive duration, and similarly, the second execution comprehensive duration may be referred to as the first execution comprehensive duration. Both the first execution comprehensive duration and the second execution comprehensive duration are execution comprehensive durations, but they are not the same execution comprehensive duration.
[0022] Another example is that terms such as "including" and "comprising" used in this application indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] As mentioned above, while the management cluster (such as M management terminals) undertakes core responsibilities such as scheduling and resource allocation for routine tasks (such as performance collection tasks for N devices to be collected), it also needs to process real-time aggregation and analysis of massive monitoring data, resulting in double occupation of its computing resources and network bandwidth. Especially in scenarios of sudden load surges (such as cascading monitoring alarms caused by a sharp increase in distributed transactions), the processing delay of a certain management terminal in the management cluster for performance collection data will exceed the index collection interval period, forming a continuous data lag effect, resulting in untimely performance monitoring of the devices to be collected in the cluster, and affecting the operation and maintenance efficiency and effect of the cluster operating system. Therefore, there is an urgent need to develop a method that can timely realize real-time processing and decision response for the monitoring data stream in a super-large-scale cluster operating system to timely improve the overall performance and resource utilization rate of the system. For this purpose, the present application proposes a performance tuning method, an electronic device, and a storage medium for a cluster operating system, which can timely realize real-time processing and decision response for the monitoring data stream in a super-large-scale cluster operating system.
[0024] The performance tuning method for the cluster operating system in this application can be applied to Figure 1 the application scenarios as shown in Figure 1 and Figure 2As shown in the figure, the cluster operating system environment includes M cluster devices (i.e., the cluster of devices to be collected) and N management terminals (i.e., the cluster of management terminals), namely cluster device 1 (i.e., device to be collected 1), cluster device 2 (i.e., device to be collected 2) …… cluster device N (i.e., device to be collected N), management terminal 1, management terminal 2 …… management terminal M. Each cluster device / device to be collected serves as a device node, and each management terminal serves as a management node. The cluster operating system is deployed on each device to be collected and each management terminal. The M management terminals monitor the metric performance of the N devices to be collected through the deployed cluster operating system, and collect the performance metrics of each device to be collected in real time according to a preset frequency. These performance metrics may include one or more of CPU utilization rate, I / O interface read and write speed, memory occupancy rate, network bandwidth and latency, disk usage rate, concurrent user number, throughput, response time (RT), etc. It can be understood that usually N is much larger than M.
[0025] Each service probe (agent, such as service probe 1 to service probe N) in the service probe layer is correspondingly deployed on each device node of the cluster operating system, responsible for collecting the performance data of the device node in real time, and transmitting the data through a distributed message queue. The middleware and storage layer is used to store the received performance data. The middleware and storage layer may include one or more components such as MySQL, Redis, OSS cloud storage, etc. For example, RabbitMQ message queue is used for asynchronous processing, and Redis is used to cache data. Relevant operation and maintenance personnel can access each management terminal through the front-end display layer (such as the APP side and / or the WEB side) and perform configuration management, data collection and analysis, and issue tuning instructions for the visual web application. The NGINX service in the gateway layer is mainly used to achieve load balancing of the management terminals and distribute the request tasks occurring at the front end to the management nodes. The cluster operating system deployed on each node can be a domestic cluster operating system, such as the Galaxy Kylin operating system. The device to be collected can be related devices such as the server to be collected or the database to be collected.
[0026] In one embodiment, as Figure 1 and Figure 2 、 Figure 3 shown, the embodiment of the present application proposes a performance tuning method for a cluster operating system. This method can be applied to the scenario as Figure 1 shown, and includes the following steps: Step S10, obtain a collection task set.
[0027] In this embodiment, the collection task set is a set formed by periodic task subsets for the N devices to be collected on which the cluster operating system is deployed. The task subset in the collection task set The performance collection task is for the th performance collection task of the th device to be collected.
[0028] Specifically, relevant operation and maintenance personnel can log in to the web management terminal and configure the device information of the devices to be collected that need to be monitored (such as server information and / or database information, etc.) on the configuration management page of the management terminal. Record the IP address and port information of the server to be collected in the relevant device information table to be collected, and the information such as the IP address, port, username, password, and database name of the database to be collected that needs to be monitored. Configure the performance indicators to be collected, such as one or more of CPU performance indicators, memory performance indicators, disk performance indicators, network performance indicators, software and hardware information, databases, distributed storage ceph, etc., and configure the collection frequency of each performance indicator. Among them, the collection frequency (collection period) of each performance indicator can be the same or different, and the collection frequency (collection period) of the same performance indicator for different devices to be collected can also be the same or different. Provide a task record form, which records the performance indicators to be collected and the collection frequency and other information. The task generator of the system will generate a batch of task lists to be executed according to the configuration information of the device information table to be collected and the task record form, and this task list is the above-mentioned collection task set.
[0029] It can be understood that each task in the collection task set is a task that needs to be repeatedly executed in real time periodically. The performance collection tasks for the same device to be collected constitute a task subset of this device to be collected. Denote the i-th device to be collected among N devices to be collected as device to be collected S i , and the j-th performance collection task for device to be collected S i is performance collection task d ij , i = 1, 2,... N. Denote the total number of performance collection tasks for device to be collected S i (that is, the total number of performance collection tasks in the performance collection task subset D i corresponding to device to be collected S i ) as a i , and the total number of performance collection tasks in the collection task set is . It can be understood that the task subsets corresponding to different devices to be collected can be the same or different, and their corresponding totals a i can also be the same or different.
[0030] In some embodiments, the monitoring requirements of each device to be collected for performance metrics are obtained, and periodic performance collection tasks are formulated according to the monitoring requirements. The performance collection tasks include, but are not limited to, the collection and processing of key performance metrics such as CPU usage rate, memory occupancy rate, disk I / O, and network traffic. The periodic task subsets of all devices to be collected are integrated to form a collection task set. Each task subset contains multiple performance collection tasks for a specific device to be collected, and each performance collection task has a clear execution period and task generation time. The task generation time is determined according to the collection frequency (collection period) of the configured performance collection task. For example, for a performance collection task of a certain CPU occupancy rate, the set collection frequency is once every 10 seconds. Then the task generation times of this performance collection task can be 14:59:00, 14:59:10, 14:59:20, 14:59:30, 14:59:40, etc.
[0031] It can be understood that the performance collection tasks in the collection task set change dynamically with the changes in the operating state of the device cluster to be collected and the monitoring requirements of performance metrics. By constructing the collection task set, the integrity and periodicity of performance data can be ensured, thereby providing accurate data support for performance tuning.
[0032] Step S20: The collection task set is simulated and assigned to M management terminals. According to the task generation time of each performance collection task and the simulation assignment result, the simulated execution time and simulated completion time of each performance collection task are predicted.
[0033] In this embodiment, the management terminal is an entity responsible for executing the performance collection task, which can be a physical server, a virtual machine, or other computing units. The simulated execution time refers to the time point when the management terminal is predicted to start executing a certain performance collection task during the simulation assignment phase. The simulated completion time refers to the time point when the management terminal is predicted to complete executing a certain performance collection task during the simulation assignment phase. The management terminal can execute the assigned performance collection tasks periodically until it receives a stop execution instruction for the performance collection task.
[0034] Since the tasks in the collection task set are fixed and the resource pools of each management terminal are also known, it is possible to predict the execution status of each performance collection task by the management terminal in different performance states (the resource pool of the management terminal, the allocated performance collection tasks) based on the historical execution results of the management terminal for each performance collection task (including the execution time, completion time, execution duration, completion duration, etc. described below). The cluster monitor of the system defines an environment to simulate the task scheduling scenario for a subset of tasks, which includes the above-mentioned D tasks and the status information of M management terminals. Based on this, a simulated allocation strategy for the task subset is formed. The simulated allocation strategy defines the performance collection tasks allocated to each management terminal and the execution order of each performance collection task.
[0035] In some embodiments, a task allocation model is established based on the task attributes of each performance collection task (generation time, execution period, CPU / memory / I / O resource requirements) and the resource pool of the management terminal (computing power benchmark, concurrent capacity, real-time load status). For example, the task allocation model is a deep reinforcement learning model, and the mapping relationship between each performance collection task and the resource pool of the management terminal is determined according to the deep reinforcement learning model. Through a discrete time axis simulation engine, an initial allocation decision is dynamically triggered based on the task generation time of each performance collection task, and the task subset in the collection task set is allocated to M management terminals according to the initial allocation decision to form a simulated allocation result. Among them, the initial allocation decision can be based on principles such as polling, load balancing, and optimal performance.
[0036] In some embodiments, the simulated execution time is determined according to the position of each performance collection task in the task queue of the allocated management terminal (the execution order of the performance collection task). The simulated completion time is determined according to the simulated execution time and the benchmark execution duration of each performance collection task. Among them, the positions of each performance collection task in the corresponding task queue are in a dynamic change state. For example, at a certain moment, the performance collection task d ij is at the head of the task queue. After completing the execution of the performance collection task d ij , the performance collection task d ij is in a frozen state. Based on the collection period of the performance collection task d ij , at the next task generation time of the performance collection task d ij , the performance collection task d ij is activated, and the performance collection task d ij is re-generated. At this time, the performance collection task d ij is at the end of the task queue.
[0037] In some embodiments, based on the simulated execution time, the benchmark execution duration, and the dynamic correction factors (such as CPU preemption latency, I / O bandwidth competition coefficient) of each performance collection task, the simulated completion time is determined, where the dynamic correction factors can be predicted according to the historical execution data of the management terminal and the performance model of the management terminal.
[0038] Step S30: Iteratively optimize the simulation allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and allocate the performance collection tasks according to the optimized task simulation allocation strategy.
[0039] In this embodiment, the system can predict the simulated execution time and simulated completion time of each management terminal for the allocated performance collection tasks under various simulation allocation strategies. Based on the execution cycle, task generation time, simulated execution time, and simulated completion time, the execution efficiency for the task set can be calculated, and the loss function of the corresponding task allocation model can be calculated according to this execution efficiency, and the parameters in the task allocation model can be adjusted to optimize the task allocation model and the simulation allocation strategy.
[0040] By iteratively optimizing the simulation allocation strategy, the task allocation can be continuously adjusted to timely achieve the real-time processing and decision response of the performance collection data stream for the ultra-large-scale device cluster to be collected.
[0041] After completing the iterative optimization of the simulation allocation strategy, each performance collection task in the collection task set can be allocated to M management terminals according to the finally optimized simulation allocation strategy, so that each management terminal executes the allocated performance collection tasks.
[0042] Step S40: Analyze the device performance status of each device to be collected based on the performance index data obtained by each management terminal executing the allocated performance collection tasks, and perform performance tuning on the devices to be collected according to the device performance status.
[0043] In this embodiment, the performance index data refers to the data obtained after the management terminal executes the performance collection task, including CPU usage rate, memory occupancy rate, disk I / O, network traffic, etc. The device performance status refers to the performance condition of the device to be collected comprehensively analyzed according to the performance index data, which may include the computing power, storage resources, network bandwidth, etc. of the device. By executing the allocated performance collection task, the management terminal can obtain the performance index data of the corresponding device to be collected, for example, the CPU occupancy rate of a certain device to be collected is 88%.
[0044] By analyzing the device performance status of each collected device, it is possible to identify whether the corresponding collected device has reached a performance bottleneck. After reaching the performance bottleneck, the tasks being executed in the collected device can be migrated. For example, a certain computing task in the collected device S i is migrated to the collected device S j and processed by the collected device S j to achieve performance tuning for the collected device and improve the overall performance of the cluster operating system.
[0045] In one embodiment, as Figure 4 shown, step S30 includes: Step A10, calculate the task completion duration and task execution duration corresponding to each performance collection task based on the task generation time, simulated execution time, and simulated completion time of each performance collection task respectively.
[0046] In this embodiment, the task completion duration T1 refers to the time interval from the task generation time T0 to the simulated completion time Tb, reflecting the total time required for the performance collection task to be created and completed. The task execution duration T2 refers to the time interval from the simulated execution time Ta to the simulated completion time Tb, reflecting the time required for the performance collection task to be actually executed on the management terminal. That is, T1 = Tb - T0; T2 = Tb - Ta.
[0047] By calculating the task completion duration and task execution duration, the performance of the performance collection task in the time dimension can be quantified, providing data support for subsequent identification of tasks to be transferred out.
[0048] Step A20, regard the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution period as tasks to be transferred out, and assign the tasks to be transferred out to one of the terminals to be transferred in.
[0049] In this embodiment, the terminals to be transferred in are the management terminals among the M management terminals excluding the scheduling terminal that simulates the execution of the tasks to be transferred out under the current simulation allocation strategy.
[0050] The execution period Tf refers to the periodic interval of the performance collection task, that is, the time interval between two consecutive task generation times. For example, a performance collection task is generated every 10 seconds, and its execution period is 10 seconds. The tasks to be transferred out refer to the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution period, indicating that these tasks may be delayed or have low execution efficiency under the current allocation strategy, that is, the performance collection tasks with T1 > Tf and / or T2 > Tf are tasks to be transferred out. The management terminal that (simulates) executes the task to be transferred out is the scheduling terminal, and the other management terminals among the M management terminals excluding the scheduling terminal are the terminals to be transferred in.
[0051] Optionally, one of the multiple terminals to be transferred can be selected as the target transfer terminal, and the task to be transferred out is assigned to the target transfer terminal. Specifically, the target transfer terminal can be selected according to one or more of the load degree, resource abundance, and comprehensive turnover degree of simulated transfer of each terminal to be transferred. For example, the terminal to be transferred with the lightest load, the most abundant resources, or the smallest comprehensive turnover degree of simulated transfer can be selected as the target transfer terminal.
[0052] As an alternative implementation of assigning the task to be transferred out to one of the terminals to be transferred, calculate the comprehensive turnover degree corresponding to the entire collection task set when each terminal to be transferred simulates the execution of the task to be transferred out. The comprehensive turnover degree will use the terminal to be transferred corresponding to the simulation assignment strategy with the smallest comprehensive turnover degree as the target transfer terminal, and assign the task to be transferred out to the target transfer terminal.
[0053] By calculating the comprehensive turnover degree, the efficiency of different management terminals in simulating the execution of the task to be transferred out can be quantified, providing a basis for selecting the optimal target transfer terminal.
[0054] Among them, the turnover degree is used to measure the completion efficiency of the task, and this completion efficiency is related to the task completion duration and the task execution duration. The single turnover degree y is used to measure the completion efficiency of a single task, and the comprehensive turnover degree Y is used to measure the task completion efficiency of the entire collection task set. It can be the weighted sum of the single turnover degrees of all performance collection tasks, that is, for each performance collection task in the collection task set, each execution is completed, and the weighted sum of the corresponding single turnover degrees of each performance collection task is obtained. Among them, the single turnover degree = task completion duration / task execution duration, that is, y = T1 / T2. Denote the single turnover degree of the performance collection task d ij as y ij and the comprehensive turnover degree Y is: .
[0055] is the weight of the th performance collection task. Optionally, the weights of each performance collection task can all be set to 1, or different weights can be set according to the different priorities of different performance collection tasks.
[0056] Optionally, the comprehensive turnover degree can be used as the loss function of the corresponding task allocation model, or based on this comprehensive turnover degree as a part of the loss function of the task allocation model, the task allocation model is iteratively optimized, so as to iteratively optimize the simulation allocation strategy.
[0057] The task allocation model can be a deep reinforcement learning model, which includes two neural networks, namely the Actor network and the Critic network. The Actor network is responsible for selecting the simulation task allocation strategy of the performance collection task to the management terminal, and the Critic network is responsible for evaluating the state value of the current simulation task allocation strategy. Before performing model iteration and simulation task allocation strategy iteration optimization, first initialize the network parameters, including the learning rate, discount factor, etc. By interacting with the environment, collect the experience of state, action, reward, and next state, use the collected experience to update the parameters of the Actor and Critic networks, and perform iterative optimization on the deep reinforcement learning model according to the calculated comprehensive turnover Y, so as to finally obtain the optimized task allocation strategy.
[0058] As another alternative implementation for allocating the task to be transferred out to one of the terminals to be transferred in, calculate the comprehensive turnover corresponding to the entire collection task set when each terminal to be transferred in simulates the execution of the task to be transferred out and the associated tasks, and use the terminal to be transferred in corresponding to the simulation allocation strategy with the smallest comprehensive turnover as the target transfer-in terminal, and allocate the task to be transferred out and the associated tasks to the target transfer-in terminal.
[0059] Among them, the associated task is a performance collection task belonging to the same task subset as the task to be transferred out, that is, the performance collection task d ij and the performance collection task d ik belong to the same task subset. When the performance collection task d ij is the task to be transferred out, the performance collection task d ik is the associated task of this performance collection task d ij Since the performance collection task and its associated tasks are all periodic collection tasks for the same device to be collected, if the performance collection task and its associated tasks are distributed to multiple management terminals for execution, it is easy to cause task allocation chaos. Therefore, here, the performance collection tasks belonging to the same task subset can be transferred out at the same time to avoid multiple management terminals collecting the same device to be collected and consuming too much scheduling resources.
[0060] Exemplarily, assume there is a device to be collected, and its task subset contains three performance collection tasks: Task A, Task B, and Task C, where Task A is the task to be transferred out. Task B and Task C are associated tasks of Task A. Among them, the task completion duration of Task A is 15 seconds, the task execution duration is 10 seconds, and the single turnover degree is 15 / 10 = 1.5. The task completion duration of Task B is 20 seconds, the task execution duration is 10 seconds, and the single turnover degree is 20 / 10 = 2.0. The task completion duration of Task C is 25 seconds, the task execution duration is 15 seconds, and the single turnover degree is 25 / 15 ≈ 1.67. Assume there are two terminals to be transferred in, Terminal X and Terminal Y. After simulating the execution of these three tasks, the comprehensive turnover degree of the entire collection task set of Terminal X is 150, and the comprehensive turnover degree of Terminal Y is 180. Then the execution efficiency of Terminal X is higher, and it receives Task A, Task B, and Task C as the target terminal to be transferred in.
[0061] As another alternative implementation manner of allocating the task to be transferred out to one of the terminals to be transferred in, determine a target terminal to be transferred in from multiple terminals to be transferred in, and calculate the first change in the comprehensive execution duration and the second change in the comprehensive execution duration of the target terminal to be transferred in respectively; when the second change in the comprehensive execution duration is greater than the first change in the comprehensive execution duration, only allocate the task to be transferred out to the target terminal to be transferred in; when the second change in the comprehensive execution duration is less than or equal to the first change in the comprehensive execution duration, allocate both the task to be transferred out and the associated tasks to the target terminal to be transferred in, and the associated tasks are performance collection tasks belonging to the same task subset as the task to be transferred out.
[0062] Among them, the change in duration is used to reflect the change in the execution duration of the management terminal for each performance collection task after changing the task allocation strategy. Specifically, the first change in the comprehensive execution duration is the weighted cumulative increase in the simulated execution duration of the performance collection tasks originally allocated by the target terminal to be transferred in after only adding the task to be transferred out; the second change in the comprehensive execution duration is the cumulative increase in the simulated execution duration of the performance collection tasks originally allocated by the target terminal to be transferred in after adding both the task to be transferred out and the associated tasks at the same time, and the associated tasks are performance collection tasks belonging to the same task subset as the task to be transferred out.
[0063] Specifically, the first change in the comprehensive execution duration can be set as , and the second change in the comprehensive execution duration . Among them, n is the number of performance collection tasks originally allocated by the target terminal to be transferred in, is the original execution duration of the k-th original task among the originally allocated performance collection tasks, is the simulated execution duration of the k-th task among the originally allocated performance collection tasks after only adding the task to be transferred out, To simultaneously increase the simulated execution duration of the k-th task in the original allocated performance collection tasks after adding both the tasks to be transferred out and the associated tasks, is the weight coefficient when only transferring out the tasks to be transferred out. The weight value can be determined according to the dispersion degree of the performance collection tasks in the task subset to which the tasks to be transferred out belong. The dispersion degree is used to reflect the degree to which the performance collection tasks belonging to the same task subset are distributed to multiple management terminals for processing. The higher the dispersion degree, the greater the weight value (penalty factor) of the corresponding performance collection task.
[0064] In one embodiment, , where, L O is the task to be transferred out, L C is the associated task of the task to be transferred out, represents the correlation degree between the task to be transferred out and its associated task, is a customizable collaborative gain coefficient, which can be a fixed value or a value adaptively determined according to the number of the original allocated performance collection tasks in the target transfer-in terminal. Among them, the larger the number of the original allocated performance collection tasks, the smaller the collaborative gain coefficient.
[0065] Optionally, the more dispersed the task to be transferred out L O is from its associated task L C (i.e., distributed to multiple different management terminals), the higher the correlation degree and the greater the corresponding weight value; the closer the task to be transferred out L O is to the associated task L C , the higher the correlation degree. The correlation degree can be calculated based on the combination of any suitable one or more correlation coefficient calculation models such as the Pearson correlation coefficient and the Spearman rank correlation coefficient. The specific values of the coefficients / parameters involved in the correlation coefficient calculation model can be determined according to the relationship between the task to be transferred out and its associated task, so as to obtain a suitable correlation degree.
[0066] For example , represents the Pearson correlation coefficient calculated according to the Pearson model, which is used to capture the linear correlation between the task to be transferred out and its associated task, represents the Spearman rank correlation coefficient calculated according to the Spearman model, which is used to capture the monotonic non-linear correlation between the task to be transferred out and its associated task, , represents the mixing ratio parameter, which can be set according to the specific relationship between the task to be transferred out and its associated task.
[0067] Further, it can also be simply set . T O1 is the first quantity of the tasks to be transferred out, T C1 is the second quantity of the associated tasks of the tasks to be transferred out.
[0068] When , only assign the tasks to be transferred out to the target incoming terminal; when , assign both the tasks to be transferred out and the associated tasks to the target incoming terminal.
[0069] Exemplarily, assume that there are 4 periodic performance collection tasks for a certain device to be collected. When one of the performance collection tasks is used as the task to be transferred out (i.e., T O1 = 1), the remaining 3 performance collection tasks are the associated tasks of this task to be transferred out (i.e., T C1 = 3). For the target incoming terminal determined for this task to be transferred out, the simulated execution duration of all its original performance collection tasks is 100 seconds. When only adding the task to be transferred out, the simulated execution duration of only the original all performance collection tasks increases to 115 seconds. Although the execution duration only increases by 15 seconds, based on the above weighted summation method, the change in the first execution comprehensive duration calculated will be greater than 15 seconds. When adding both the task to be transferred out and the associated tasks at the same time, the simulated execution duration of only the original all performance collection tasks increases to 120 seconds, and the change in the second execution comprehensive duration is 20 seconds. At this time, if the change in the first execution comprehensive duration calculated according to the above embodiment is still less than 20 seconds (i.e., ), then only assign the task to be transferred out to the target incoming terminal; if the calculated change in the first execution comprehensive duration is greater than or equal to 20 seconds (i.e., ), then assign both the task to be transferred out and the associated tasks to the target incoming terminal.
[0070] In this embodiment, by introducing a penalty factor to adjust the calculation of the weighted cumulative increase duration, the purpose is to tend to maintain the integrity of the task subset in the task scheduling decision, that is, under feasible circumstances, try to transfer the entire task subset (including the tasks to be transferred out and the associated tasks) together, rather than transferring only the tasks to be transferred out alone. Thereby reducing the scheduling complexity and resource coordination overhead caused by task dispersion, and thus improving the efficiency of task scheduling and resource utilization.
[0071] In another embodiment, as Figure 5 shown, step S30 includes: Step B10: Calculate the task completion duration and task execution duration corresponding to each performance collection task based on the task generation time, simulated execution time, and simulated completion time of each performance collection task.
[0072] Step B20: Take the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution period as the tasks to be transferred out.
[0073] In this embodiment, the execution period refers to the periodic interval of the performance collection task, that is, the time interval between two consecutive generations of tasks. The tasks to be transferred out refer to the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution period, indicating that these tasks may be delayed or have low execution efficiency under the current allocation strategy.
[0074] Step B30: Calculate the influence degree of each performance collection task simulated and executed by the terminal to be scheduled on the tasks to be transferred out, and take the preset number of performance collection tasks with the highest influence degree as the tasks to be transferred.
[0075] In this embodiment, the terminal to be scheduled refers to the management terminal that simulates and executes the tasks to be transferred out under the current simulated allocation strategy. The influence degree refers to the degree of influence on the execution efficiency of the tasks to be transferred out by other performance collection tasks on the terminal to be scheduled. The tasks to be transferred refer to the preset number of performance collection tasks with the highest influence degree, and these tasks will be selected to be transferred out from the terminal to be scheduled.
[0076] In some embodiments, calculate the influence degree of each performance collection task on the terminal to be scheduled on the tasks to be transferred out respectively. The influence degree can be comprehensively evaluated by factors such as resource occupancy, execution priority, and dependency relationship for the tasks to be transferred out.
[0077] In some embodiments, identify the resource types required by the tasks to be transferred out, such as CPU, memory, disk I / O, etc. Calculate the occupancy rate of each performance collection task on the terminal to be scheduled on each required resource type of the tasks to be transferred out. For example, for each performance collection task i , its occupancy rate on the resource type r can be expressed as: Ui , r = the occupancy amount of the performance collection task Ui , r = the total amount of the resource type i on the resource type r / the total amount of the resource type r . Based on the occupancy rate of each performance collection task on the terminal to be scheduled on each required resource type of the tasks to be transferred out, calculate the influence degree of each performance collection task on the terminal to be scheduled on the tasks to be transferred out. For example, for each performance collection task i , its influence degree IiIt can be expressed as: . Among them, R is the set of resource types required for the task to be retrieved, Wr is the resource type r 's weight, which is set according to its impact on task execution. Among them, the weight of the resource type r can be dynamically adjusted according to the real-time state of the task to be retrieved and the load condition of the to-be-scheduled terminal.
[0078] Step B40: Allocate the task to be transferred out to one of the to-be-transferred-in terminals, where the to-be-transferred-in terminals are the management terminals among the M management terminals except the to-be-scheduled terminal that is simulating the execution of the task to be retrieved under the current simulated allocation strategy.
[0079] In this embodiment, the target transferred-in terminal can be determined from multiple to-be-transferred-in terminals according to the above method of calculating the turnover degree, and the task to be transferred out is allocated to the target transferred-in terminal, or the task to be transferred out and its associated tasks are both allocated to the target transferred-in terminal.
[0080] By analyzing the performance execution task that has the greatest impact on the task to be retrieved to determine the task to be transferred out, and allocating the task to be transferred out to the target transferred-in terminal instead of directly allocating the task to be retrieved for transfer, the comprehensive collection efficiency of the performance collection tasks in the collection task set can be further improved.
[0081] As an implementation method for performance optimization of the device to be collected according to the device performance state, the threshold of the performance state can be preset. When the device performance state exceeds the set threshold, it is determined that the device has an abnormality. Further, the device to be collected with an abnormal device performance state is determined as the device to be processed.
[0082] Obtain the task execution information of the device to be processed. The task execution information includes information such as the resource requirements of each task being executed in the task list of the device to be processed. Screen out the tasks corresponding to the abnormal type of the device performance state from the task execution information as the tasks to be processed. For example, if the device performance state of the device to be processed is in a high-risk state of GPU overload, then screen out the tasks with high CPU resource requirements as the tasks to be processed.
[0083] Optimize the performance of the device to be processed according to several resource requirement combinations of the task to be processed. The resource requirement combination refers to different resource combinations that the task to be processed may need during execution, including CPU, memory, disk I / O, etc. Moreover, due to the dynamic changes in the execution environment and requirements of the task, the resource requirement combination is not fixed and may have multiple combination methods due to resource compensation. Resource compensation means that by adjusting the resource usage method of the task, part or all of the demand for one resource is transferred to another resource to optimize the task execution efficiency and device performance. For example, the CPU load of some tasks can be reduced by GPU acceleration, thereby reducing the demand for CPU resources.
[0084] Specifically, determine whether the performance index data associated with the device to be processed meets at least one of several resource requirement combinations; if so, based on the satisfied resource requirement combination, perform resource reallocation of the device to be processed for the task to be processed; if not, according to the performance index data of other collected devices, determine the collected device that meets at least one resource requirement combination as the standby device to transfer the task to be processed to the standby device.
[0085] For example, assume that the resource requirement combinations of the task to be processed include: Combination 1: CPU demand 10%, GPU demand 20%; Combination 2: CPU demand 5%, GPU demand 30%. The performance index data of the device to be processed shows that the CPU usage rate that can be allocated to the task to be processed is 5%, and the GPU usage rate is 40%. Since the device to be processed currently only allocates 5% of the CPU usage rate and 20% of the GPU usage rate to the task to be processed according to the resource requirements of Combination 1, the task to be processed is abnormal. However, at this time, it is judged that the performance index of the device to be processed meets the conditions of Combination 2 (CPU usage rate 5% ≤ 5%, GPU usage rate 40% ≥ 30%). Therefore, Combination 2 can be selected as the resource requirement combination for optimization, and an additional 10% of the GPU usage rate can be allocated to the task to be processed.
[0086] For another example, assume that the resource requirement combinations of the tasks to be processed include: Combination 1: CPU requirement 10%, GPU requirement 20%; Combination 2: CPU requirement 5%, GPU requirement 30%. The performance indicator data of the device to be processed shows that the CPU usage that can be allocated to the tasks to be processed is 5%, and the GPU usage is 20%. At this time, it is determined that the performance indicator of the device to be processed does not meet the conditions of combination 1 or combination 2. Therefore, based on the performance indicator data of other collected devices, the collected devices that meet at least one resource requirement combination are determined as backup devices. For example, there are also devices B and C. Among them, the CPU usage that can be allocated to device B is 10%, and the GPU usage is 20%. The CPU usage that can be allocated to device C is 5%, and the GPU usage is 30%. It is detected that device C meets combination 2 of the tasks to be processed, so the tasks to be processed are transferred to device C.
[0087] In the performance tuning method of the cluster operating system proposed in the embodiment of the present application, on the one hand, by obtaining the collection task set formed by the periodic task subset for the collected device and simulating the allocation to the management terminal, the simulated execution time and the simulated completion time of each performance collection task can be predicted in advance. This helps to reasonably plan the allocation of performance collection tasks, avoid problems such as task accumulation, data lag and untimely performance monitoring, and improve the allocation efficiency and accuracy of performance collection tasks. On the other hand, by iteratively optimizing the simulation allocation strategy, the task allocation strategy can be continuously adjusted to make it more adaptable to the actual operating environment, reduce the waiting time and execution time of the task, and improve the overall performance of the management terminal cluster. On the third hand, by analyzing the performance indicator data obtained by each management terminal executing the assigned performance collection task, the device performance status of each collected device can be accurately evaluated. Based on these detailed performance status data, the performance of the collected device can be tuned in a targeted manner to improve the operating efficiency and stability of the collected device.
[0088] In one embodiment, a computer-readable storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the steps in the above-mentioned method embodiments.
[0089] In one embodiment, an electronic device is also provided, including one or more processors; a memory, in which one or more programs are stored, wherein when the one or more programs are executed by one or more processors, the one or more processors execute the steps in the above-mentioned method embodiments.
[0090] In one embodiment, Figure 6As shown, it shows a schematic structural diagram of an electronic device for implementing an embodiment of the present application. The electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0091] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read from it can be installed into the storage section 608 as needed.
[0092] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer-readable medium carrying instructions. In such an embodiment, the instructions can be downloaded and installed from the network through the communication section 609, and / or installed from the removable medium 611. When the instructions are executed by the central processing unit (CPU) 601, the various method steps described in the present application are executed.
[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0094] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, any one of the above embodiments can be used in any combination. The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present application, and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those skilled in the art.
Claims
1. A method for performance tuning of a cluster operating system, characterized in that, The method includes: Obtain a collection of collection tasks, where the collection task collection is a collection formed by periodic task subsets for N devices to be collected on which a cluster operating system is deployed. The jth performance collection task in the ith task subset in the collection task collection is the jth performance collection task for the ith device to be collected; The performance collection task is the jth performance collection task for the ith device to be collected; Simulating and allocating the collection task set to M management terminals, and predicting the simulated execution time and simulated completion time of each performance collection task according to the task generation time and simulated allocation result of each performance collection task; Iteratively optimizing the simulated allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and allocating the performance collection tasks according to the optimized task simulated allocation strategy; Analyzing the device performance status of each device to be collected based on the performance index data obtained by each management terminal executing the allocated performance collection tasks, and performing performance tuning on the device to be collected according to the device performance status.
2. The performance tuning method of the cluster operating system according to claim 1, wherein The iteratively optimizing the simulated allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and allocating the performance collection tasks according to the optimized task simulated allocation strategy includes: Calculating the task completion duration and task execution duration of the performance collection task corresponding to the task respectively according to the task generation time, simulated execution time, and simulated completion time of each performance collection task; Regarding the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution cycle as tasks to be transferred out, and allocating the tasks to be transferred out to one of the terminals to be transferred in, where the terminal to be transferred in is a management terminal among the M management terminals excluding the scheduling terminal that simulates the execution of the tasks to be transferred out under the current simulated allocation strategy.
3. The method according to claim 2, characterized in that, The allocating the tasks to be transferred out to one of the terminals to be transferred in includes: Calculating the comprehensive turnover degree corresponding to the entire collection task set in the state where each terminal to be transferred in simulates the execution of the tasks to be transferred out, where the comprehensive turnover degree is the weighted sum of the individual turnover degrees of each performance collection task, and the individual turnover degree is the quotient of the task completion duration and task execution duration of the corresponding performance collection task; Regarding the terminal to be transferred in corresponding to the simulated allocation strategy with the smallest comprehensive turnover degree as the target terminal to be transferred in, and allocating the tasks to be transferred out to the target terminal to be transferred in.
4. The method according to claim 2, wherein The allocating the tasks to be transferred out to one of the terminals to be transferred in includes: Calculating the comprehensive turnover degree corresponding to the entire collection task set in the state where each terminal to be transferred in simulates the execution of the tasks to be transferred out and associated tasks, where the comprehensive turnover degree is the weighted sum of the individual turnover degrees of each performance collection task, the individual turnover degree is the quotient of the task completion duration and task execution duration of the corresponding performance collection task, and the associated tasks are performance collection tasks belonging to the same task subset as the tasks to be transferred out; Regarding the terminal to be transferred in corresponding to the simulated allocation strategy with the smallest comprehensive turnover degree as the target terminal to be transferred in, and allocating the tasks to be transferred out and the associated tasks to the target terminal to be transferred in.
5. The method according to claim 2, characterized in that, The allocating the tasks to be transferred out to one of the terminals to be transferred in includes: Determining a target terminal to be transferred in from multiple terminals to be transferred in, and calculating the first change in the comprehensive execution duration and the second change in the comprehensive execution duration of the target terminal to be transferred in respectively; When the change in the second comprehensive execution duration is greater than the change in the first comprehensive execution duration, only allocate the task to be transferred out to the target incoming terminal; When the change in the second comprehensive execution duration is less than or equal to the change in the first comprehensive execution duration, allocate both the task to be transferred out and the associated task to the target incoming terminal, where the associated task is a performance collection task belonging to the same task subset as the task to be transferred out; The change in the first comprehensive execution duration is the weighted cumulative increase duration of the simulated execution duration of the originally allocated performance collection tasks after only adding the task to be transferred out to the target incoming terminal; The change in the second comprehensive execution duration is the cumulative increase duration of the simulated execution duration of the originally allocated performance collection tasks after adding both the task to be transferred out and the associated task to the target incoming terminal.
6. The performance tuning method of the cluster operating system according to claim 1, wherein Iteratively optimizing the simulated allocation strategy based on the execution cycle, task generation time, simulated execution time, and simulated completion time of each performance collection task, and performing performance collection task allocation according to the optimized task simulation allocation strategy, including: Calculating the task completion duration and task execution duration of the corresponding performance collection task according to the task generation time, simulated execution time, and simulated completion time of each performance collection task; Regarding the performance collection tasks whose task completion duration and / or task execution duration exceed the corresponding execution cycle as tasks to be transferred out; Calculating the influence degree of each performance collection task simulated and executed by the terminal to be scheduled on the task to be transferred out, and regarding the preset number of performance collection tasks with the highest influence degree as tasks to be transferred out; Allocating the tasks to be transferred out to one of the incoming terminals, where the incoming terminal is a management terminal among the M management terminals, excluding the terminal to be scheduled that simulates the execution of the task to be transferred out under the current simulated allocation strategy.
7. The performance tuning method of the cluster operating system according to claim 1, wherein The performance tuning of the device to be collected according to the device performance state includes: Determining the device to be collected with abnormal device performance state as the device to be processed; Obtaining the task execution information of the device to be processed, and screening out the tasks corresponding to the abnormal type of the device performance state from the task execution information as the tasks to be processed; Performing performance tuning on the device to be processed according to several resource requirement combinations of the tasks to be processed.
8. The performance tuning method of the cluster operating system according to claim 7, characterized in that Performing performance tuning on the device to be processed according to several resource requirement combinations of the tasks to be processed, including: Judging whether the performance index data associated with the device to be processed meets at least one of the several resource requirement combinations; If so, based on the satisfied resource requirement combination, perform resource reallocation of the device to be processed for the task to be processed; If not, according to the performance index data of other devices to be collected, determine the device to be collected that meets at least one resource requirement combination as the standby device, so as to transfer the task to be processed to the standby device.
9. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method for performance tuning of the cluster operating system according to any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores executable instructions, and when the instructions are executed by a processor, the processor is caused to execute the method for performance tuning of the cluster operating system according to any one of claims 1 to 8.
Citation Information
Patent Citations
Task distributing method and scanner
CN103699443A
Cluster scaling and expansion method and system, scaling and expansion control terminal and medium
CN113835824A
Task scheduling method and device, computer equipment and storage medium
CN115437770A
Manage and Worker-based task distribution method and device
CN118606016A
Large-model heterogeneous cluster scheduling system and method based on adaptive parallel co-optimization
CN118916156A