Resource optimization scheduling method for high-performance parallel computing data center
By collecting and processing multi-dimensional resource scheduling parameters in high-performance parallel computing data centers, calculating the comprehensive resource scheduling index and building a multi-objective optimization model, the problem that traditional scheduling methods are difficult to comprehensively consider multi-dimensional factors, and the optimal balance of resource utilization efficiency, system security and task execution coordination is achieved.
Patent Information
- Application Number
- CN202510130259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional resource scheduling methods are difficult to effectively consider multidimensional factors in high-performance parallel computing data centers, such as storage redundancy, security policies, virtualized resource management and task execution coordination, which makes it difficult to improve resource utilization efficiency, system security and task execution coordination.
By collecting storage redundant parameters, virtualized resource coordination parameters and task execution coordination parameters, comprehensive scheduling target indicators such as data shield coordination index, virtual coordination complex index and task collaborative execution index, and integrating them into comprehensive resource scheduling index, a multi-objective optimization model is built to generate resource allocation plans.
It realizes comprehensive and comprehensive evaluation and optimization scheduling of high-performance parallel computing data center resources, improves the accuracy and comprehensiveness of resource management, and ensures the best balance of resource utilization efficiency, system security and task execution coordination.
Smart Images

Figure CN120066780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource optimization scheduling, and specifically provides a resource optimization scheduling method for a high-performance parallel computing data center. Background Art
[0002] With the rapid development of high-performance parallel computing technology and the expansion of large-scale data center infrastructure, the management and scheduling problems of resources (including computing, storage, and network) in computing clusters have become increasingly complex. Traditional resource scheduling often focuses on the optimization of a single dimension (such as CPU utilization or task waiting time), ignoring the comprehensive consideration of multi-dimensional factors such as storage redundancy, security policies, virtualized resource management, and the coordination degree during task execution. Especially in data centers adopting virtualization and containerization technologies, parameters such as the number of virtual machine migrations, container health status, resource demand change rate, and allocation fluctuation amplitude will dynamically affect the effectiveness and feasibility of resource scheduling decisions. In addition, in scenarios dealing with data redundancy and security detection requirements, only focusing on the scheduling efficiency at the task level while ignoring parameters such as storage redundancy, security vulnerability detection frequency, and data access error rate may lead to an increased operation risk of the storage system and reduce data availability and system reliability.
[0003] Most of the existing methods only perform one-dimensional optimization for some elements. For example, they rely on heuristic algorithms to minimize task waiting time as much as possible, or linearly process resource constraints through simple weighted combinations. However, there are often multiple conflicting goals in a data center: it is required to ensure task execution efficiency, guarantee storage system redundancy and security, and maintain the coordination and stability of virtualized resources. In this context, there is a lack of a method to extract comprehensive scheduling target metrics from multi-dimensional parameters and fuse them into a comprehensive resource scheduling index. Traditional optimization methods have limited capabilities in dealing with the fusion of complex multi-dimensional parameters and non-linear interactions, and are difficult to accurately solve multi-objective optimization models. Summary of the Invention
[0004] Based on the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a resource optimization scheduling method for a high-performance parallel computing data center to solve the above technical problems.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A resource optimization scheduling method for a high-performance parallel computing data center, comprising:
[0006] Collect the storage redundancy parameters, virtualized resource coordination parameters, and task execution coordination parameters of the high-performance parallel computing data center, and preprocess the parameters;
[0007] Calculate the comprehensive scheduling target metrics based on the preprocessed storage redundancy parameters, virtualized resource coordination parameters, and task execution coordination parameters. The comprehensive scheduling target metrics include the data shield collaboration index, virtual coordination complexity index, and task collaboration execution index;
[0008] Calculate the comprehensive resource scheduling index based on the comprehensive scheduling target metrics, set the comprehensive resource scheduling index as the objective function, define the constraint conditions, and construct a multi-objective optimization model;
[0009] Solve the multi-objective optimization model according to the multi-objective optimization algorithm to generate a resource allocation plan. According to the resource allocation plan, allocate tasks to the corresponding resource nodes to complete resource optimization scheduling.
[0010] The present invention is further configured that the storage redundancy parameters include data redundancy, security vulnerability detection frequency, data access error rate, number of vulnerability detection events, and vulnerability detection security factor; the virtualized resource coordination parameters include the number of virtual machine migrations, container health status score, virtual machine startup time, resource allocation fluctuation range, and resource demand change rate; the task execution coordination parameters include task waiting time, task execution time, scheduling delay, scheduling success rate, and resource allocation success rate; the preprocessing includes data cleaning, data transformation, and data smoothing and filtering.
[0011] The present invention is further configured to calculate the data shield collaboration index according to the data redundancy, security vulnerability detection frequency, data access error rate, number of vulnerability detection events, and vulnerability detection security factor;
[0012] Calculate the virtual coordination complexity index according to the number of virtual machine migrations, container health status score, virtual machine startup time, resource allocation fluctuation range, and resource demand change rate;
[0013] Calculate the task collaboration execution index according to the task waiting time, task execution time, scheduling delay, scheduling success rate, and resource allocation success rate.
[0014] The present invention is further configured that the calculation logic of the data shield collaboration index is: where DSSI is the data shield collaboration index, P is the total number of storage devices in the high-performance parallel computing data center, k is the index of the storage device, is the data redundancy of the kth storage device, is the security vulnerability detection frequency of the kth storage device, is the data access error rate of the kth storage device, Q is the total number of vulnerability detection systems, l is the index of the vulnerability detection system, is the number of vulnerability detection events of the lth vulnerability detection system, is the vulnerability detection security coefficient of the l-th vulnerability detection system, and α is the weight factor of the vulnerability detection security coefficient.
[0015] The present invention is further configured such that the calculation logic of the virtual coordination complexity index is: Wherein, VCCI is the virtual
[0016] coordination complexity index, R is the total number of virtual machines, m is the index of the virtual machine, is the number of virtual machine migrations of the m-th virtual machine, is the container health status score of the m-th virtual machine, is the startup time of the m-th virtual machine, S is the total number of resource nodes, n is the index of the resource node,
[0017] is the resource allocation fluctuation range of the n-th resource node, is the resource demand change rate of the n-th resource node, and β is the weight factor of the resource demand change rate.
[0018] The present invention is further configured such that the calculation logic of the task collaborative execution index is:
[0019]
[0020] collaborative execution index, V is the total number of tasks, t is the index of the task, is the task waiting time of the t-th task, is the task execution time of the t-th task, is the scheduling delay of the t-th task, D is the total number of scheduling decisions, s is the index of the scheduling decision, is the scheduling success rate of the s-th scheduling decision, is the resource allocation success rate of the s-th scheduling decision.
[0021] The present invention is further configured such that the calculation logic of the comprehensive resource scheduling index is: CRSI = ω 1 · DSSI + ω 2 · VCCI + ω 3 · TCEI, where CRSI is the comprehensive resource scheduling index, DSSI is the data shield collaboration index, VCCI is the virtual coordination complexity index, TCEI is the task collaborative execution index, ω 1 、ω 2 and ω 3 are weight coefficients, and ω 1 + ω 2 + ω 3 = 1.
[0022] The present invention is further configured such that, in the multi-objective optimization model, when the constraint conditions are satisfied, the optimization objective comprehensive resource scheduling index CRSI is maximized, where the constraint conditions include: resource capacity constraint and the allocation availability constraint of any task in the node. The resource capacity constraint is: where SD is the task that needs to perform resource scheduling, and D ij is the demand of the i-th task that needs to perform resource scheduling for the j-th resource node, and x ij is the decision indication for allocating the i-th task that needs to perform resource scheduling to the resource node j, and C j is the capacity of the j-th resource node; the allocation availability constraint of any task in the node is: and S is the total number of resource nodes.
[0023] The present invention is further configured to solve the multi-objective optimization model according to a multi-objective optimization algorithm to generate a resource allocation plan, including:
[0024] For the SD tasks that need to perform resource scheduling and the S resource nodes, samples are taken from a uniform distribution to generate an initial solution set {x ij};
[0025] Calculate the comprehensive resource scheduling index CRSI of the initial solution set {x ij}, verify whether the constraint conditions are satisfied, and record the verification result as: Calculate the comprehensive optimization index of the solution set: SSCOI = CRSI({x ij})·exp(H({x ij}) - 1);
[0026] Through iterative optimization by the multi-objective optimization algorithm, obtain the solution set with the maximum comprehensive optimization index SSCOI of the solution set, generate an actual mapping table of tasks and resource nodes according to the solution set, and allocate the tasks to the corresponding resource nodes to complete resource optimization scheduling.
[0027] The present invention is further configured such that the multi-objective optimization algorithm includes: multi-objective particle swarm optimization algorithm, Pareto-based optimization algorithm, and multi-objective simulated annealing optimization algorithm.
[0028] The present invention provides a resource optimization scheduling method for a high-performance parallel computing data center. By collecting storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters of the high-performance parallel computing data center, preprocessing the parameters; calculating a comprehensive scheduling target index according to the preprocessed storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters, where the comprehensive scheduling target index includes a data shield collaboration index, a virtual coordination complexity index, and a task collaboration execution index; calculating a comprehensive resource scheduling index according to the comprehensive scheduling target index, setting the comprehensive resource scheduling index as the objective function, defining constraint conditions, and constructing a multi-objective optimization model; solving the multi-objective optimization model according to a multi-objective optimization algorithm, generating a resource allocation plan, and according to the resource allocation plan, allocating tasks to corresponding resource nodes to complete resource optimization scheduling. The beneficial effects generated include:
[0029] 1. Comprehensive resource evaluation ability: By collecting storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters, and combining the data shield collaboration index, virtual coordination complexity index, and task collaboration execution index, it is possible to comprehensively and meticulously evaluate the usage status and coordination efficiency of the data center resources. Compared with traditional scheduling methods that only focus on a single dimension, it provides a multi-level and multi-angle resource evaluation, significantly improving the accuracy and comprehensiveness of resource management;
[0030] 2. Efficient construction of the comprehensive resource scheduling index: By non-linearly combining and performing complex mathematical operations, the data shield collaboration index, virtual coordination complexity index, and task collaboration execution index are integrated into a comprehensive resource scheduling index, effectively reflecting the overall resource utilization efficiency, system security, and comprehensive state of task execution coordination in the data center. This index avoids the linear deviation caused by simple weighted summation, ensures the non-linear coupling relationship between different indicators, and more accurately reflects the real needs and dynamic changes of resource scheduling;
[0031] 3. Multi-objective optimization model and solution method: Using a multi-objective optimization algorithm to solve the constructed optimization model can efficiently search for the optimal solution in a complex and high-dimensional solution space. Combining a non-linear scoring function and a rigorous definition of constraint conditions, the optimization process has good global search ability and convergence performance, significantly improving the optimization quality and practicality of the resource allocation plan.
[0032] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. In the accompanying drawings:
[0034] Figure 1 It is a flowchart of a resource optimization scheduling method for a high-performance parallel computing data center shown in an exemplary embodiment of the present invention. Detailed implementation manners
[0035] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention and not for limiting the protection scope of the present invention.
[0036] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0037] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0038] A resource optimization scheduling method for a high-performance parallel computing data center, as Figure 1 shown, includes:
[0039] Collect the storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters of the high-performance parallel computing data center, and preprocess the parameters;
[0040] Calculate the comprehensive scheduling target metrics according to the preprocessed storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters. The comprehensive scheduling target metrics include a data shield cooperation index, a virtual coordination complexity index, and a task cooperation execution index;
[0041] Calculate the comprehensive resource scheduling index according to the comprehensive scheduling target index, set the comprehensive resource scheduling index as the objective function, define the constraint conditions, and construct a multi-objective optimization model;
[0042] Solve the multi-objective optimization model according to the multi-objective optimization algorithm, generate a resource allocation plan, and allocate tasks to the corresponding resource nodes according to the resource allocation plan to complete resource optimization scheduling.
[0043] The present invention is further configured such that the storage redundancy parameters include data redundancy, security vulnerability detection frequency, data access error rate, number of vulnerability detection events, and vulnerability detection security factor; the virtualized resource coordination parameters include the number of virtual machine migrations, container health status score, virtual machine startup time, resource allocation fluctuation range, and resource demand change rate; the task execution coordination parameters include task waiting time, task execution time, scheduling delay, scheduling success rate, and resource allocation success rate; the preprocessing includes data cleaning, data transformation, and data smoothing and filtering. Specifically, among the storage redundancy parameters, the data redundancy is used to measure the data redundancy level of the storage device, reflecting the distribution of data replication or check information among different storage units. A high redundancy means stronger fault tolerance and data availability; the security vulnerability detection frequency represents the number of security vulnerability detections performed by the storage device per unit time, reflecting the monitoring intensity of the system for potential security threats; the data access error rate records the proportion of data access operation errors that occur in the storage device within a specific time, reflecting the reliability of data access; the number of vulnerability detection events represents the total number of vulnerability events detected by the vulnerability detection system within a certain time, reflecting the degree of security threats faced by the system; the vulnerability detection security factor is used to quantify the security of the vulnerability detection system, and the higher the value, the stronger the effectiveness and reliability of the vulnerability detection system. Among the virtualized resource coordination parameters, the number of virtual machine migrations records the number of times a virtual machine migrates from one physical node to another within a certain time, reflecting the dynamics of resource scheduling and the flexibility of virtualization management; the container health status score is used to evaluate the running health status of the containers in the virtual machine, and the higher the score, the better the running condition of the containers and the higher the system stability; the virtual machine startup time is used to measure the time required for the virtual machine to transition from startup to an available state, affecting the response speed of tasks and the overall system performance; the resource allocation fluctuation range represents the change range of the resource allocation amount on the resource node within a certain time, reflecting the stability and consistency of resource scheduling; the resource demand change rate is used to measure the change rate or proportion of the virtual machine resource demand on the resource node over time, affecting the adaptability and flexibility of the resource scheduling strategy. Among the task execution coordination parameters, the task waiting time records the time elapsed from when the task is submitted to when it starts to execute, reflecting the response efficiency of the scheduling system; the task execution time is used to measure the time required for the task to complete from the start of execution, affecting the overall execution efficiency of the task; the scheduling delay represents the time required for the scheduler to allocate resources to the task, affecting the startup speed of the task and the overall throughput of the system; the scheduling success rate records the proportion of scheduling decisions that successfully allocate tasks to resource nodes, reflecting the effectiveness of the scheduling strategy; the resource allocation success rate measures the proportion of scheduling decisions that successfully meet the resource requirements of the tasks, reflecting the accuracy of resource matching and the adequacy of resource utilization. The preprocessing including data cleaning, data transformation, and data smoothing and filtering is prior art and will not be elaborated herein.
[0044] The present invention is further configured to calculate a data shield collaboration index based on data redundancy, security vulnerability detection frequency, data access error rate, the number of vulnerability detection events, and vulnerability detection security factor; the present invention is further configured that the calculation logic of the data shield collaboration index is as follows: where DSSI is the data shield collaboration index, P is the total number of storage devices in the high-performance parallel computing data center, k is the index of the storage device, is the data redundancy of the k-th storage device, is the security vulnerability detection frequency of the k-th storage device, is the data access error rate of the k-th storage device, Q is the total number of vulnerability detection systems, l is the index of the vulnerability detection system, is the number of vulnerability detection events of the l-th vulnerability detection system, is the vulnerability detection security factor of the l-th vulnerability detection system, and α is the weight factor of the vulnerability detection security factor. Specifically, the calculation of the data shield collaboration index DSSI aims to comprehensively evaluate the redundancy and security of the storage system in the high-performance parallel computing data center. By combining multiple key parameters, DSSI can reflect the comprehensive performance of the storage system in data redundancy and security protection, thereby providing a quantitative basis for resource optimization scheduling. Its calculation logic can be divided into two major parts. Redundancy and security assessment at the storage device level: For each storage device, calculate the product of its redundancy and security vulnerability detection frequency, and adjust the influence of the data access error rate. Take the geometric mean of the evaluation results of all storage devices to obtain the overall storage redundancy and security level; for each storage device, multiply its data redundancy by the security vulnerability detection frequency to reflect the comprehensive performance of the device in data redundancy and security protection. Divide the above product by (1 + data access error rate) to reduce the contribution of high error rates to DSSI and reflect the importance of data access reliability. Take the P-th root of the adjusted product of all storage devices to balance the weights of different storage devices in DSSI and ensure the comparability and stability of the overall evaluation; Security event impact assessment at the vulnerability detection system level: Summarize the detection event numbers of all vulnerability detection systems and adjust the influence of the security factor. Quantify the negative impact of vulnerability detection events on the overall DSSI through an exponential function. For each vulnerability detection system, calculate the number of vulnerability detection events divided by (1 + security factor multiplied by the weight factor) to reflect the negative impact of the security events of each system on DSSI. A high security factor means that the same number of vulnerability events has a smaller impact on DSSI. Add up the influence degrees of all vulnerability detection systems to obtain the total negative impact of security events. Take the negative value of the cumulative impact and map it through an exponential function to ensure that the negative impact of security events decays in a non-linear manner, enhance the sensitivity to high-impact events, and further reduce the DSSI value. Finally, DSSI comprehensively reflects the redundancy and security collaboration ability of the data center storage system through the combination of these two parts.
[0045] Calculate the virtual coordination complexity index based on the number of virtual machine migrations, the container health status score, the virtual machine startup time, the resource allocation fluctuation range, and the resource demand change rate; the present invention is further set such that the calculation logic of the virtual coordination complexity index is as follows: where VCCI is the virtual coordination complexity index, R is the total number of virtual machines, m is the index of the virtual machine, is the number of virtual machine migrations of the m-th virtual machine, is the container health status score of the m-th virtual machine, is the startup time of the m-th virtual machine, S is the total number of resource nodes, n is the index of the resource node, is the resource allocation fluctuation range of the n-th resource node, Let \(\Delta r_n\) be the rate of change of resource demand for the \(n\)th resource node, and \(\beta\) be the weight factor of the rate of change of resource demand. Specifically, the calculation of the Virtual Coordination Complexity Index (VCCI) aims to comprehensively evaluate the coordination and management complexity of virtualized resources in a high-performance parallel computing data center. By combining multiple key parameters, VCCI can reflect the comprehensive impact of factors such as virtual machine migration, container health status, resource allocation fluctuations, and demand changes on resource coordination in a virtualized environment, thereby providing a quantitative basis for resource optimization scheduling. Its calculation logic can be divided into two main parts. Coordination complexity assessment at the virtual machine and container management level: By multiplying the number of virtual machine migrations by the container health status score and combining the logarithmic adjustment of the virtual machine startup time, the dynamics and stability of virtualized resource management are quantified. Sum up the evaluation results of all virtual machines and take the root of the total number of virtual machines to the power of the exponent to balance the weights of different virtual machines in VCCI and ensure the comparability and stability of the overall evaluation. For each virtual machine, multiply its number of migrations by the container health status score to reflect the dynamics of virtualized resource management and the stability of container operation. Divide the above product by \((1 + \log(\text{virtual machine startup time}))\) to reduce the contribution of high startup time to VCCI and reflect the importance of virtual machine startup efficiency. Sum up the adjusted products of all virtual machines and take the reciprocal power of the total number of virtual machines to balance the weights of different virtual machines in VCCI and ensure the comparability and stability of the overall evaluation. Coordination complexity assessment at the level of resource allocation fluctuations and demand changes: By taking the ratio of the resource allocation fluctuation amplitude to the rate of change of resource demand and combining the weight factor of the rate of change of resource demand, the stability and adaptability of the resource scheduling strategy are quantified. Take the product of the evaluation results of all resource nodes and take the fifth root of the total number of resource nodes to comprehensively reflect the impact of resource allocation and demand changes on coordination complexity. For each resource node, calculate the ratio of the resource allocation fluctuation amplitude to \((1+\text{rate of change of resource demand}\times\text{weight factor})\) to reflect the stability and adaptability of the resource allocation strategy. Take the product of the adjusted ratios of all resource nodes and take the fifth root to comprehensively reflect the impact of resource allocation and demand changes on coordination complexity. Finally, through the combination of these two parts, VCCI comprehensively reflects the complexity and management efficiency of resource coordination in a virtualized environment.
[0046] Calculate the task collaborative execution index according to the task waiting time, task execution time, scheduling delay, scheduling success rate, and resource allocation success rate. The present invention is further set such that the calculation logic of the task collaborative execution index is: where \(TCEI\) is the task collaborative execution index, \(V\) is the total number of tasks, \(t\) is the index of the task, is the task waiting time of the \(t\)th task, is the task execution time of the \(t\)th task, is the scheduling delay of the \(t\)th task, \(D\) is the total number of scheduling decisions, \(s\) is the index of the scheduling decision, is the scheduling success rate of the sth scheduling decision, and is the resource allocation success rate of the sth scheduling decision. Specifically, the calculation of the Task Cooperative Execution Index (TCEI) aims to comprehensively evaluate the coordination and efficiency of task execution and scheduling processes in a high-performance parallel computing data center. By combining multiple key parameters, TCEI can reflect the comprehensive performance of tasks during waiting, execution, and scheduling, thus providing a quantitative basis for resource optimization scheduling. Its calculation logic can be divided into two main parts. Coordination evaluation of task execution and waiting time: Considering the waiting time and execution time of tasks comprehensively, adjusting the impact of scheduling delay in a non-linear manner, quantifying the evaluation results of task execution efficiency and scheduling response speed, taking the reciprocal power of the total number of tasks, balancing the weights of different tasks in TCEI, and ensuring the comparability and stability of the overall evaluation. By multiplying the waiting time and execution time of tasks, it reflects the total time consumption of tasks in the system. Dividing the above product by (1 + scheduling delay) reduces the impact of scheduling delay on task time consumption and emphasizes the importance of scheduling efficiency. Summing up the adjusted products of all tasks and taking the reciprocal power of the total number of tasks, balancing the weights of different tasks in TCEI, and ensuring the comparability and stability of the overall evaluation; Coordination evaluation of scheduling decision success rate and resource allocation success rate: By the ratio of the scheduling success rate to the resource allocation success rate, quantifying the effectiveness of the scheduling strategy and the accuracy of resource matching. Taking the product of the evaluation results of all scheduling decisions and taking the reciprocal power of the total number of scheduling decisions, comprehensively reflecting the impact of scheduling decisions on task cooperative execution. Dividing the scheduling success rate by the resource allocation success rate reflects the accuracy of scheduling decisions in resource allocation. A higher scheduling success rate and a lower resource allocation failure rate indicate the efficiency and precision of the scheduling strategy. Adding 1 to the denominator avoids division by zero errors and reduces the impact of a high resource allocation success rate on the ratio. Taking the product of the adjusted ratios of all scheduling decisions and taking the reciprocal power of the total number of scheduling decisions, comprehensively reflecting the overall impact of the scheduling strategy on task cooperative execution. Finally, through the combination of these two parts, TCEI comprehensively reflects the cooperative efficiency and coordination complexity in task execution and scheduling processes.
[0047] The present invention is further configured such that the calculation logic of the comprehensive resource scheduling index is: CRSI = ω 1 ·DSSI + ω 2 ·WCCI + ω 3 ·TCEI, where CRSI is the comprehensive resource scheduling index, DSSI is the data shield cooperation index, VCCI is the virtual coordination complexity index, TCEI is the task cooperative execution index, ω 1 , ω 2 and ω 3 are weight coefficients, and ω 1 + ω 2 + ω3 = 1. Specifically, the calculation of the Comprehensive Resource Scheduling Index (CRSI) aims to comprehensively evaluate the overall efficiency and effectiveness of resource scheduling in a high-performance parallel computing data center by integrating multiple key performance indicators. CRSI reflects the comprehensive performance of the data center in multiple dimensions such as data security, virtualized resource coordination, and task execution coordination through a weighted combination of the Data Shield Synergy Index (DSSI), the Virtual Coordination Complexity Index (VCCI), and the Task Cooperative Execution Index (TCEI). Weight coefficients ω 1 , ω 2 and ω 3 are assigned to each indicator, and these weights reflect the relative importance of each indicator in the overall resource scheduling. The indicators are combined into CRSI through weighted summation to ensure that the influence of different indicators is reasonably reflected according to the weight coefficients.
[0048] The present invention is further configured such that in the multi-objective optimization model, when the constraint conditions are satisfied, the optimization objective Comprehensive Resource Scheduling Index (CRSI) is maximized, where the constraint conditions include: resource capacity constraint and the availability constraint of any task allocated in a node. The resource capacity constraint is: where SD is the task for which resource scheduling is required, D ij is the demand of the i-th task for which resource scheduling is required for the j-th resource node, x ij is the decision indication for the i-th task for which resource scheduling is required to be allocated to the resource node j, and C j is the capacity of the j-th resource node; the availability constraint of any task allocated in a node is: and S is the total number of resource nodes. Specifically, in the resource optimization scheduling method of a high-performance parallel computing data center, the key to constructing a multi-objective optimization model lies in defining appropriate objective functions and constraint conditions. The core objective of the multi-objective optimization model is to maximize the Comprehensive Resource Scheduling Index (CRSI) to achieve the best balance of resource utilization efficiency, system security, and task execution coordination. To ensure the feasibility and rationality of the optimization process, the following two types of constraint conditions are introduced: resource capacity constraint: ensuring that the resource allocation of each resource node does not exceed its capacity to prevent resource overload. For each resource node j, the sum of the resource requirements of all tasks i allocated to it shall not exceed its capacity C j . Through this constraint, it is ensured that resource nodes are not over-allocated, thus avoiding resource overload and system performance degradation; task allocation availability constraint: ensuring that each task is allocated to and only allocated to one resource node to ensure the uniqueness of the task and the effective utilization of resources. Each x ijIt can only take 0 or 1 to ensure that a task is either not assigned to a certain resource node or is completely assigned. Each task i must be assigned to and only assigned to one resource node j. That is, for each task i, the sum of all x ij must be equal to 1. This ensures the uniqueness of tasks and the effectiveness of resource allocation. Through these constraints, the optimization model can seek the optimal resource allocation scheme on the premise of meeting the physical resource limitations and task scheduling requirements, thereby maximizing the CRSI.
[0049] The present invention is further configured to solve the multi-objective optimization model according to a multi-objective optimization algorithm to generate a resource allocation scheme, including:
[0050] For SD tasks that need to perform resource scheduling and S resource nodes, sample from a uniform distribution to generate an initial solution set {x ij}; specifically, the initial solution set {x ij} represents the initial allocation relationship between tasks and resource nodes. Among them, x ij is a decision variable, indicating whether task i is assigned to resource node j, and is generated by using a uniform distribution sampling method to ensure that the initial solution set has good diversity and covers a wide range of possible spaces;
[0051] Calculate the comprehensive resource scheduling index CRSI of the initial solution set {x ij}, verify whether the constraint conditions are satisfied, and record the verification result as: Calculate the comprehensive optimization index of the solution set: SSCOI = CRSI({x ij})·exp(H({x ij}) - 1); specifically, H({x ij}) is a constraint condition verification function, used to determine whether the solution set {x ij} satisfies the predetermined constraint conditions to ensure the feasibility of the resource allocation scheme under the physical resource limitations and task scheduling requirements; the comprehensive optimization index of the solution set combines CRSI and the constraint condition verification result to strengthen the preferential selection of feasible solutions;
[0052] Through iterative optimization with a multi-objective optimization algorithm, the solution set with the maximum comprehensive optimization index SSCOI is obtained. According to the solution set, an actual mapping table of tasks and resource nodes is generated, and tasks are assigned to the corresponding resource nodes to complete resource optimization scheduling. Specifically, the multi-objective optimization algorithm is used to continuously adjust and optimize the solution set with the goal of maximizing SSCOI. During the optimization process, the solution set with the maximum comprehensive optimization index SSCOI is selected as the final resource allocation plan. An actual mapping table of tasks and resource nodes is generated based on the optimal solution set, and tasks are assigned to the corresponding resource nodes to achieve resource optimization scheduling. The present invention is further configured such that the multi-objective optimization algorithm includes: a multi-objective particle swarm optimization algorithm, a Pareto-based optimization algorithm, and a multi-objective simulated annealing optimization algorithm. Specifically, the multi-objective particle swarm optimization algorithm (MOPSO, Multi-Objective Particle Swarm Optimization) is an extension based on the particle swarm optimization (PSO) algorithm, specifically designed to solve optimization problems involving multiple objectives. The particle swarm optimization algorithm simulates the group behavior of bird flocks or fish schools, and the search process is jointly guided by the individual and group experiences of particles. In multi-objective optimization, MOPSO ensures finding multiple non-dominated solutions by maintaining a Pareto Front archive, reflecting the trade-off relationships between different objectives. The Pareto-based optimization algorithms aim to find multiple non-dominated solutions that lie on the Pareto Front and cannot be improved in terms of other objectives without deteriorating at least one objective. Typical representatives include the Non-dominated Sorting Genetic Algorithm II (NSGA-II) and the Non-dominated Sorting Genetic Algorithm III (NSGA-III). These algorithms can effectively maintain the diversity and uniformity of the solution set and approach the Pareto Front through fast non-dominated sorting, crowding distance calculation, and selection mechanisms. The multi-objective simulated annealing optimization algorithm (MOSA, Multi-Objective Simulated Annealing) is a multi-objective extension based on the simulated annealing (SA) algorithm, simulating the physical annealing process, escaping from local optima through random perturbations and temperature control, and searching for the global optimal solution. MOSA balances between multiple objectives and explores the solution space using a probabilistic acceptance mechanism, being suitable for dealing with complex multi-objective optimization problems, especially performing excellently when there are multiple local optima in the solution space. By introducing the multi-objective optimization algorithms (MOPSO, Pareto-based optimization algorithm, MOSA), efficient, comprehensive, and intelligent resource scheduling optimization is achieved in complex and high-dimensional multi-objective optimization models.This not only significantly improves the resource utilization efficiency, system security, and task execution coordination, but also enhances the adaptability, robustness, and automation level of data center resource management, with broad application prospects and significant economic and technical value.
[0053] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0054] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0055] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0056] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not indicate the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0057] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0058] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0059] In several embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0060] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0061] In addition, the functional units in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0062] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0063] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A resource optimization scheduling method for a high-performance parallel computing data center, characterized in that: include: Collecting storage redundancy parameters, virtualization resource coordination parameters, and task execution coordination parameters of a high-performance parallel computing data center, and preprocessing the parameters; Calculate the comprehensive scheduling target index according to the preprocessed storage redundancy parameter, virtualization resource coordination parameter and task execution coordination parameter, wherein the comprehensive scheduling target index includes a data shield coordination index, a virtual coordination complexity index and a task coordination execution index; A comprehensive resource scheduling index is calculated according to the comprehensive scheduling target index, the comprehensive resource scheduling index is set as the objective function, constraints are defined, and a multi-objective optimization model is constructed; The multi-objective optimization model is solved according to the multi-objective optimization algorithm to generate a resource allocation plan. According to the resource allocation plan, tasks are allocated to corresponding resource nodes to complete resource optimization scheduling.
2. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 1, characterized in that: The storage redundancy parameters include data redundancy, security vulnerability detection frequency, data access error rate, number of vulnerability detection events and vulnerability detection safety factor; the virtualization resource coordination parameters include number of virtual machine migrations, container health status score, virtual machine startup time, resource allocation fluctuation range and resource demand change rate; the task execution coordination parameters include task waiting time, task execution time, scheduling delay, scheduling success rate and resource allocation success rate; the preprocessing includes data cleaning, data transformation and data smoothing and filtering.
3. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 2, characterized in that: The Data Shield Synergy Index is calculated based on data redundancy, security vulnerability detection frequency, data access error rate, number of vulnerability detection events, and vulnerability detection safety factor; The virtual coordination complexity index is calculated based on the number of virtual machine migrations, container health status score, virtual machine startup time, resource allocation fluctuation range, and resource demand change rate; The task collaborative execution index is calculated based on task waiting time, task execution time, scheduling delay, scheduling success rate and resource allocation success rate.
4. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 3 is characterized in that: The calculation logic of the data shield synergy index is: Where DSSI is the data shield synergy index, P is the total number of storage devices in the high-performance parallel computing data center, k is the index of the storage device, is the data redundancy of the kth storage device, is the security vulnerability detection frequency of the kth storage device, is the data access error rate of the kth storage device, Q is the total number of vulnerability detection systems, l is the index of the vulnerability detection system, is the number of vulnerability detection events of the first vulnerability detection system, is the vulnerability detection safety factor of the lth vulnerability detection system, and α is the weight factor of the vulnerability detection safety factor.
5. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 1, characterized in that: The calculation logic of the virtual coordination complexity index is: Where VCCI is the virtual coordination complexity index, R is the total number of virtual machines, m is the index of the virtual machine, is the number of virtual machine migrations of the mth virtual machine, Score the health status of the container of the mth virtual machine, is the startup time of the mth virtual machine, S is the total number of resource nodes, n is the index of the resource node, is the resource allocation fluctuation range of the nth resource node, is the resource demand change rate of the nth resource node, and β is the weight factor of the resource demand change rate.
6. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 1, characterized in that: The calculation logic of the task collaborative execution index is: Among them, TCEI is the task collaborative execution index, V is the total number of tasks, t is the index of the task, is the task waiting time of the tth task, is the task execution time of the tth task, is the scheduling delay of the tth task, D is the total number of scheduling decisions, s is the index of the scheduling decision, is the scheduling success rate of the s-th scheduling decision, is the resource allocation success rate of the sth scheduling decision.
7. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 3, characterized in that: The calculation logic of the comprehensive resource scheduling index is: CRSI=ω1·DSSI+ω2·VCCI+ω3·TCEI, wherein CRSI is the comprehensive resource scheduling index, DSSI is the data shield coordination index, VCCI is the virtual coordination complexity index, TCEI is the task collaborative execution index, ω1, ω2 and ω3 are weight coefficients, and ω1+ω2+ω3=1.
8. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 7, characterized in that: In the multi-objective optimization model, when the constraints are met, the optimization target comprehensive resource scheduling index CRSI is maximized, wherein the constraints include: resource capacity constraints and any task allocation availability constraints in the node, and the resource capacity constraints are: Among them, SD is the task that needs resource scheduling, D ij is the demand of the i-th task that needs resource scheduling for the j-th resource node, x ij is the decision instruction for assigning the i-th task that needs resource scheduling to resource node j, C j is the capacity of the jth resource node; the availability constraint for any task to be allocated in the node is: and S is the total number of resource nodes.
9. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 8, characterized in that: The multi-objective optimization model is solved according to a multi-objective optimization algorithm to generate a resource allocation plan, including: For SD tasks and S resource nodes that need resource scheduling, sample from uniform distribution to generate the initial solution set {x ij }; Calculate the initial solution set {x ij }’s comprehensive resource scheduling index CRSI is used to verify whether the constraints are met, and the verification result is recorded as: Calculate the solution set comprehensive optimization index: SSCOI = CRSI ({x ij })·exp(H({x ij })-1); Through iterative optimization of the multi-objective optimization algorithm, the solution set with the largest solution set comprehensive optimization index SSCOI is obtained. According to the solution set, the actual mapping table of tasks and resource nodes is generated, and the tasks are assigned to the corresponding resource nodes to complete the resource optimization scheduling.
10. The resource optimization scheduling method for a high-performance parallel computing data center according to claim 9, characterized in that: Multi-objective optimization algorithms include: multi-objective particle swarm optimization algorithm, Pareto-based optimization algorithm and multi-objective simulated annealing optimization algorithm.
Citation Information
Cited By
Offshore multi-target supply decision optimization method, system, equipment and product
CN120634145A
Cost optimization method for resource scheduling management of cloud data center
CN120762920A
Heterogeneous quantum computing resource scheduling method and device, equipment and medium
CN121900915A
A heterogeneous quantum computing resource scheduling method, device, equipment and medium
CN121900915B
Data center load elastic regulation potential assessment method and device
CN122152543A