Intelligent Computing Center Computing Power Asymmetric Cooperative Scheduling Method for Inclusive Computing Power

By obtaining the operating parameters of computing power operation tasks in the intelligent computing center, dividing them into different groups and scheduling them to the acceleration card for execution, the problems of low resource utilization and high leasing costs are solved, efficient resource utilization and cost reduction are achieved, and the application of universal computing power is promoted.

CN119938284BActive Publication Date: 2025-07-01DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510439992.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-01
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The low resource utilization rate of intelligent computing centers and high computing power leasing costs have led to limited application of universal computing power.

Method used

By obtaining the operating parameters of the computing power operation tasks in the intelligent computing center, dividing them into different groups, and scheduling the tasks to different accelerator cards for execution, the preset parameter conditions are met to improve resource utilization and reduce rental costs.

Benefits of technology

It improves the resource utilization rate of the intelligent computing center, reduces the cost of computing power leasing, and realizes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938284B_ABST
    Figure CN119938284B_ABST
Patent Text Reader

Abstract

The present invention provides an asymmetric cooperative scheduling method for computing power in an intelligent computing center for inclusive computing power, which relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure. The method includes: Step S1, obtaining the operation parameters corresponding to each computing power operation task among multiple computing power operation tasks in the intelligent computing center. The operation parameters include the utilization rate of the streaming multiprocessor (SM) and other parameters, and the other parameters include at least one of the video memory utilization rate, the activity of the Compute Unified Device Architecture (CUDA) cores, the memory occupancy rate, and the bandwidth occupancy rate; Step S2, dividing the multiple computing power operation tasks based on the operation parameters to obtain at least one group; Step S3, scheduling the computing power operation tasks included in different groups to different acceleration cards in the intelligent computing center, and each acceleration card in the intelligent computing center is used to execute the computing power operation tasks within one group. The present invention can greatly improve the resource utilization rate and greatly reduce the computing power leasing cost, realizing the wide application of inclusive computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to an asymmetric collaborative scheduling method for the computing power of an intelligent computing center for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios for the development, training, and inference of artificial intelligence deep learning models) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", is the ability of computer devices or computing / data centers to process information, is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, is the computing ability to achieve the output of the target result by processing information data, is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] Currently, intelligent computing centers provide acceleration cards to provide acceleration computing power services for users. In the prior art, multiple acceleration cards are arranged in an intelligent computing center, and each acceleration card respectively executes different computing power operation tasks. However, the resources required by different users are different. For example, some computing power operation tasks require more video memory resources, and some other computing power operation tasks require more computing power resources. In the process of the acceleration cards in the intelligent computing center providing computing power services for users, there is a situation where the video memory resources or computing power resources of some acceleration cards are wasted, resulting in a very low resource utilization rate of the intelligent computing center. At the same time, since the computing power services provided by the intelligent computing center are limited by the number of acceleration cards, when users lease computing power services, they usually need to lease the entire acceleration card, resulting in a very high lease cost and making it difficult to achieve the widespread application of inclusive computing power.

[0008] It can be seen that there are problems in the prior art such as a very low resource utilization rate of intelligent computing centers and a very high computing power lease cost. Summary of the Invention

[0009] The present invention provides a method for asymmetric cooperative scheduling of computing power in an intelligent computing center for inclusive computing power, so as to solve the problems of low resource utilization rate of the intelligent computing center and high computing power leasing cost in the prior art.

[0010] To solve the above problems, the present invention is implemented as follows:

[0011] In a first aspect, the present invention provides a method for asymmetric cooperative scheduling of computing power in an intelligent computing center for inclusive computing power, including:

[0012] Step S1, obtaining the operation parameters corresponding to each computing power operation task in multiple computing power operation tasks in the intelligent computing center, where the operation parameters include the utilization rate of the streaming multiprocessor (SM) and other parameters, and the other parameters include at least one of the video memory utilization rate, the activity of the Compute Unified Device Architecture (CUDA) cores, the memory occupancy rate, and the bandwidth occupancy rate;

[0013] Step S2, dividing the multiple computing power operation tasks based on the operation parameters to obtain at least one group, and the operation parameters corresponding to each group meet the preset parameter conditions;

[0014] Step S3, scheduling the computing power operation tasks included in different groups to different acceleration cards of the intelligent computing center, and each acceleration card of the intelligent computing center is used to execute the computing power operation tasks within one group.

[0015] In one embodiment, the preset parameter conditions include:

[0016] The sum of the SM utilization rates corresponding to each computing power operation task included in each group is greater than or equal to a second set threshold;

[0017] The preset parameter conditions further include at least one of the following:

[0018] The sum of the video memory utilization rates corresponding to each computing power operation task included in each group is less than or equal to a first set threshold;

[0019] The sum of the CUDA core activities corresponding to each computing power operation task included in each group is less than or equal to a third set threshold;

[0020] The sum of the memory occupancy rates corresponding to each computing power operation task included in each group is less than or equal to a fourth set threshold;

[0021] The sum of the bandwidth occupancy rates corresponding to each computing power operation task included in each group is less than or equal to a fifth set threshold.

[0022] In one embodiment, the step S2 includes:

[0023] Step S21: Traverse multiple groups in the intelligent computing center that allow adding computing power operation tasks;

[0024] Step S22: When the sum of the SM utilization rate of the first group and the SM utilization rate of the first computing power operation task is less than the first set threshold among the multiple groups, add the first computing power operation task to the first group, where the first computing power operation task is the computing power operation task to be added to a group among the multiple computing power operation tasks.

[0025] In one embodiment, step S2 further includes:

[0026] Step S24: When the sum of the SM utilization rate of the second group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold among the multiple groups, and the second sum value of other parameters of the second group and other parameters of the first computing power operation task meets the preset parameter condition, add the first computing power operation task to the second group;

[0027] After adding the first computing power operation task to the second group, set the second group as the group that stops adding computing power operation tasks.

[0028] In one embodiment, step S2 further includes:

[0029] Step S26: When there is no group whose sum of the SM utilization rate and the SM utilization rate of the first computing power operation task is less than the first set threshold among the multiple groups, and there is no group whose sum of the SM utilization rate and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, and the second sum value of other parameters of the group and other parameters of the first computing power operation task meets the preset parameter condition, create a third group;

[0030] Step S27: Add the first computing power operation task to the third group.

[0031] In one embodiment, among the computing power operation tasks included in each group, there are first-type tasks and second-type tasks, where,

[0032] The first-type task is a computing power operation task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization rate is less than the first set video memory threshold;

[0033] The second-type task is a computing power operation task whose other parameters are less than the second set parameter threshold and whose SM utilization rate is greater than or equal to the second set video memory threshold;

[0034] Step S3 includes:

[0035] Step S31: Schedule the first type of tasks to the accelerator card where the second type of tasks is located.

[0036] In one embodiment, step S31 includes:

[0037] Step S311: Calculate the migration cost corresponding to each group based on the data volume of the first type of tasks;

[0038] Step S312: Weight the SM utilization rate, other parameters of the at least one computing power operation task, and the migration cost to obtain the scheduling benefit parameter of each group;

[0039] Step S313: When the scheduling benefit parameter is greater than or equal to the set scheduling threshold, schedule the first type of tasks to the accelerator card where the second type of tasks is located.

[0040] In the present invention, the operation parameters corresponding to each computing power operation task in the intelligent computing center are obtained. The operation parameters include the streaming processor SM utilization rate and other parameters, and the other parameters include at least one of the video memory utilization rate, unified computing device architecture (CUDA) core activity, memory occupancy rate, and bandwidth occupancy rate. Based on the operation parameters, the multiple computing power operation tasks are divided into at least one group, and the operation parameters corresponding to each group meet the preset parameter conditions. The computing power operation tasks included in different groups are scheduled to different accelerator cards in the intelligent computing center, and each accelerator card in the intelligent computing center is used to execute the computing power operation tasks within one group. In this way, by dividing multiple computing power operation tasks into different groups and having the accelerator cards in the intelligent computing center execute the computing power operation tasks within the groups, the accelerator cards in the intelligent computing center can execute different computing power operation tasks simultaneously, greatly improving the resource utilization rate. At the same time, when a user leases computing power services, the user only needs to lease part of the resources of the accelerator card, greatly reducing the leasing cost, thereby realizing the wide application of inclusive computing power. Description of the Drawings

[0041] To more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for the description of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 It is a flowchart of a method for asymmetric cooperative scheduling of computing power in an intelligent computing center for inclusive computing power provided by the present invention. Detailed Embodiments

[0043] The following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the protection scope of the present invention.

[0044] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and achieve the output of the target result, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0045] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and achieve the output of the result, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .

[0046] The "carrying capacity" (Network Power, NP) described in the present invention refers to: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive index to measure the network transmission scheduling ability.

[0047] The "storage power" (Storage Power, SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive index to measure the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0048] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized computing, storage, transmission, and application of information.

[0049] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology infrastructures such as artificial intelligence, blockchain, and quantum computing.

[0050] The "computing power" described in the present invention includes general computing power, intelligent computing power, and super computing power.

[0051] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0052] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled up and deployed for various artificial intelligence innovation applications based on special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing and machine vision.

[0053] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0054] The "intelligent computing center" described in the present invention refers to a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0055] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0056] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0057] The "Computing Power Center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0058] The "Supercomputing Center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0059] The "Computing Power Resources" described in the present invention refer to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0060] The "Inclusive Computing Power" described in the present invention refers to: based on the requirements of equal opportunity and the principle of commercial sustainability, providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost.

[0061] The "Models" described in the present invention include but are not limited to "Large Language Models" and "Multimodal Large Models".

[0062] The "Large Language Model" described in the present invention refers to the large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0063] The "Multimodal Large Models" described in the present invention refer to: models trained by jointly combining multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0064] The "asymmetric collaborative scheduling" described in the present invention refers to running tasks by scheduling computing power, so that different parameters of the acceleration card reach an asymmetric balance state after the scheduling computing power runs the tasks. Among them, the asymmetric balance state means that different parameters respectively meet the corresponding set parameter ranges, and the set parameter ranges corresponding to different parameters can be the same or different. Through asymmetric collaborative scheduling, different parameters in the acceleration card can reach the set parameter range, greatly improving the utilization rate of parameter resources.

[0065] The "computing power running task" described in the present invention refers to a specific workload or job that is executed on computing power resources and requires a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.

[0066] Please refer to Figure 1 , Figure 1 which is a flowchart of an asymmetric collaborative scheduling method for computing power of an intelligent computing center for inclusive computing power provided by the present invention. As Figure 1 shown, it includes the following steps:

[0067] Step S1: Obtain the operation parameters corresponding to each computing power running task among multiple computing power running tasks in the intelligent computing center. The operation parameters include the utilization rate of Streaming Multiprocessors (SM) and other parameters. The other parameters include at least one of the utilization rate of video memory, the activity of Compute Unified Device Architecture (CUDA) cores, the memory occupancy rate, and the bandwidth occupancy rate.

[0068] The above operation parameters are obtained by collecting the computing power running tasks running in each acceleration card in the intelligent computing center. In some embodiments, a real-time monitoring module is arranged in the intelligent computing center, and the different computing power running tasks running in different acceleration cards are monitored through the real-time monitoring module to obtain the operation parameters corresponding to each computing power running task.

[0069] It should be noted that there are differences in the operation parameters corresponding to different computing power running tasks, and there are also differences in different types of operation parameters corresponding to one computing power running task. For example, for a model training computing power running task with a large amount of data, its occupied SM utilization rate is large, and its video memory occupancy rate is also large; for a model training computing power running task with a small amount of data, its occupied SM utilization rate is small, and its video memory occupancy rate is also small. By obtaining the operation parameters corresponding to different computing power running tasks, the different computing power running tasks can be divided by the operation parameters.

[0070] The above operating parameters include SM utilization rate and other parameters, and the other parameters can be one or more of the video memory utilization rate, CUDA core activity, memory occupancy rate, and bandwidth occupancy rate. For example, the operating parameters include SM utilization rate and video memory utilization rate, and different computing power running tasks are divided based on the SM utilization rate and video memory utilization rate.

[0071] Step S2: Divide the multiple computing power running tasks based on the operating parameters to obtain at least one group, and the operating parameters corresponding to each group satisfy the preset parameter conditions.

[0072] It should be noted that the operating parameters include SM utilization rate and other parameters, and the other parameters include at least one of the video memory utilization rate, CUDA core activity, memory occupancy rate, and bandwidth occupancy rate. The preset condition parameters are specifically set according to the parameters included in the operating parameters, so that when the operating parameters corresponding to each group satisfy the preset parameter conditions, the SM utilization rate and other parameters included in the operating parameters can achieve a relatively high utilization.

[0073] Specifically, in one embodiment, the preset parameter conditions include:

[0074] The sum of the SM utilization rates corresponding to each computing power running task included in each group is greater than or equal to the second set threshold;

[0075] The preset parameter conditions further include at least one of the following:

[0076] The sum of the video memory utilization rates corresponding to each computing power running task included in each group is less than or equal to the first set threshold;

[0077] The sum of the CUDA core activities corresponding to each computing power running task included in each group is less than or equal to the third set threshold;

[0078] The sum of the memory occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fourth set threshold;

[0079] The sum of the bandwidth occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fifth set threshold.

[0080] In this way, through the above method, while the SM utilization rate of each group is highly utilized, other parameters do not exceed the bearing range of the acceleration card, and the resource utilization rate is improved as much as possible.

[0081] Each of the above groups includes at least one computing power operation task. Among them, when each group includes only one computing power operation task, the acceleration card has used up various resources as much as possible when executing this computing power operation task, or the remaining resources are not enough to execute other computing power operation tasks. In this case, it is not necessary to schedule this computing power operation task. And when each group includes at least two computing power operation tasks, the two computing power operation tasks can be executed by one acceleration card at the same time, which can greatly improve the resource utilization rate of the acceleration card.

[0082] Step S3: Schedule the computing power operation tasks included in different groups to different acceleration cards of the intelligent computing center, and each acceleration card of the intelligent computing center is used to execute the computing power operation tasks within one group.

[0083] In the present invention, obtain the operation parameters corresponding to each computing power operation task among multiple computing power operation tasks in the intelligent computing center. The operation parameters include the utilization rate of the streaming multiprocessor (SM) and other parameters. The other parameters include at least one of the video memory utilization rate, the activity of the Compute Unified Device Architecture (CUDA) cores, the memory occupancy rate, and the bandwidth occupancy rate. Based on the operation parameters, divide the multiple computing power operation tasks to obtain at least one group, and the operation parameters corresponding to each group meet the preset parameter conditions. Schedule the computing power operation tasks included in different groups to different acceleration cards of the intelligent computing center, and each acceleration card of the intelligent computing center is used to execute the computing power operation tasks within one group. In this way, by dividing multiple computing power operation tasks into different groups and having the acceleration cards of the intelligent computing center execute the computing power operation tasks within the groups, the acceleration cards of the intelligent computing center can execute different computing power operation tasks at the same time, greatly improving the resource utilization rate. At the same time, when users lease computing power services, they only need to lease part of the resources of the acceleration card, greatly reducing the leasing cost, thus realizing the wide application of inclusive computing power.

[0084] In one embodiment, step S2 includes:

[0085] Step S21: Traverse multiple groups in the intelligent computing center that allow adding computing power operation tasks;

[0086] Step S22: When the sum of the SM utilization rate of the first group and the SM utilization rate of the first computing power operation task is less than the first set threshold among the multiple groups, add the first computing power operation task to the first group, where the first computing power operation task is the computing power operation task to be added to the group among the multiple computing power operation tasks.

[0087] The above-mentioned multiple groups are groups that allow the addition of computing power operation tasks. It should be noted that the intelligent computing center includes groups that allow the addition of computing power operation tasks and groups that stop adding computing power operation tasks. Among them, the groups that allow the addition of computing power operation tasks are groups with unused resources, and at this time, computing power operation tasks can still be added to the group; the groups that stop adding computing power operation tasks are groups in which a certain resource has been used up. At this time, since a certain resource in the group has been used up, adding computing power operation tasks again will cause the system resources to be insufficient and lead to crashes, so computing power operation tasks cannot be added at this time.

[0088] In some embodiments, after the computing power operation tasks in the group are executed, the operation parameters of other computing power operation tasks in the group are recalculated, and based on the operation parameters of other computing power operation tasks, it is determined whether to add computing power operation tasks to the group.

[0089] The sum of the SM utilization rate of the first group and the SM utilization rate of the first computing power operation task is used to represent the SM utilization rate of the first group after the first computing power operation task is added to the first group. It should be noted that when the sum of the SM utilization rate of the first group and the SM utilization rate of the first computing power operation task is less than the first set threshold, at this time, the first group occupies less resources of the acceleration card, and the computing power operation task can be directly added to the first group, and the resources of the acceleration card of the intelligent computing center can effectively execute each computing power operation task included in the first group.

[0090] In one embodiment, step S2 further includes:

[0091] Step S24, when the sum of the SM utilization rate of the second group and the SM utilization rate of the first computing power operation task in the multiple groups is greater than or equal to the first set threshold, and the second sum value of the other parameters of the second group and the other parameters of the first computing power operation task meets the preset parameter conditions, add the first computing power operation task to the second group;

[0092] Step S25, after adding the first computing power operation task to the second group, set the second group as a group that stops adding computing power operation tasks.

[0093] It should be noted that when the sum of the SM utilization rate of the second group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, the SM utilization rate reaches the limit of the SM resources that the acceleration card can provide. At this time, it is necessary to determine whether other parameters reach the preset parameter conditions to confirm whether it is necessary to add the computing power operation task to the second group.

[0094] In the present invention, when the sum of the SM utilization rate of the second group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, and the second sum value of the other parameters of the second group and the other parameters of the first computing power operation task meets the preset parameter conditions, after adding the first computing power operation task to the second group, the SM utilization rate and other parameters of the second group can both meet the preset parameter conditions. At this time, when adding the first computing power operation task to the second group, the acceleration card of the intelligent computing center can execute all the computing power operation tasks within the second group. However, in this case, since the SM utilization rate and other parameters of the second group can both meet the preset parameter conditions, adding another computing power operation task to the second group will cause a certain parameter of the acceleration card executing the second group to exceed the limit, resulting in system resource shortage and crashing. Therefore, at this time, no more computing power operation tasks can be added, and the second group needs to be set as the group that stops adding computing power operation tasks to maintain the stability of the second group.

[0095] In one embodiment, step S2 further includes:

[0096] Step S26, when there is no group in the multiple groups where the sum of the SM utilization rate of the group and the SM utilization rate of the first computing power operation task is less than the first set threshold, and there is no group where the sum of the SM utilization rate of the group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, and the second sum value of the other parameters of the group and the other parameters of the first computing power operation task meets the preset parameter conditions, create a third group;

[0097] Step S27, add the first computing power operation task to the third group.

[0098] It should be noted that when the sum of the SM utilization rate of the group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, and the second sum value of the other parameters of the group and the other parameters of the first computing power operation task does not meet the preset parameter conditions, at this time, other computing power operation tasks can be added to the group. However, adding the first computing power operation task to the group will cause a certain parameter of the acceleration card executing the group to exceed the limit, resulting in system resource shortage and crashing. Therefore, at this time, the first computing power operation task cannot be added to this group, and another group needs to be selected to add the first computing power operation task.

[0099] In the present invention, when there is no case where the sum of the SM utilization rate of a group and the SM utilization rate of the first computing power operation task in the multiple groups is less than the first set threshold, and there is no case where the sum of the SM utilization rate of a group and the SM utilization rate of the first computing power operation task is greater than or equal to the first set threshold, and the second sum value of the other parameters of the group and the other parameters of the first computing power operation task meets the preset parameter condition, at this time, there is no suitable group for adding the first computing power operation task among the multiple groups for adding computing power operation tasks during the operation of the intelligent computing center. At this time, a new third group needs to be created, and the first computing power operation task is added through this third group.

[0100] In one embodiment, the computing power operation tasks included in each group include a first type of task and a second type of task, where

[0101] the first type of task is a computing power operation task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization rate is less than the first set video memory threshold;

[0102] the second type of task is a computing power operation task whose other parameters are less than the second set parameter threshold and whose SM utilization rate is greater than or equal to the second set video memory threshold;

[0103] Step S3 includes:

[0104] Step S31: Schedule the first type of task to the acceleration card where the second type of task is located.

[0105] The parameters between the above-mentioned first type of task and the second type of task are complementary. By distinguishing the first type of task and the second type of task, the scheduling method of the computing power operation task is determined.

[0106] For example, the other parameters include the video memory occupancy rate. The first type of task is a computing power operation task with a video memory occupancy rate greater than 60% and an SM utilization rate less than 30%; the second type of task is a computing power operation task with a video memory occupancy rate less than 40% and an SM utilization rate greater than 65%. At this time, the first type of task is a video memory-intensive task, and the second type of task is a computing-intensive task. At this time, scheduling the first type of task to the acceleration card where the second type of task is located can switch the traffic of the first type of task to the new acceleration card through load balancing and gray release; take offline the first type of task in the old acceleration card to complete the scheduling of the computing power operation task.

[0107] In some other embodiments, the second type of task can also be scheduled to the acceleration card where the first type of task is located.

[0108] In one embodiment, step S31 includes:

[0109] Step S311: Calculate the migration cost corresponding to each group based on the data volume of the first type of task;

[0110] Step S312: Weight the SM utilization rate, other parameters, and the migration cost of the at least one computing power operation task to obtain the scheduling benefit parameter for each group;

[0111] Step S313: When the scheduling benefit parameter is greater than or equal to the set scheduling threshold, schedule the first type of task to the acceleration card where the second type of task is located.

[0112] It should be noted that scheduling the computing power operation task consumes the resources of the intelligent computing center. If the consumed resources are greater than the resources saved after scheduling the computing power operation task, then scheduling the computing power operation task will instead reduce the resource utilization rate of the intelligent computing center. Therefore, after dividing the groups, it is also necessary to calculate the scheduling benefit parameter to determine whether scheduling needs to be performed.

[0113] In some embodiments, the migration cost can be calculated based on parameters such as the data volume, number of executions, and transmission bandwidth of the first type of task. For example, the migration cost of the first type of task can be calculated through the following formula:

[0114] Migration cost = k1 × data volume + k2 × number of executions + k3 × transmission bandwidth;

[0115] Among them, k1, k2, and k3 are weight coefficients.

[0116] The above scheduling benefit parameter can be calculated through the following formula:

[0117] F = α × (1 - other parameters) + β × SM utilization rate + γ × migration cost;

[0118] Among them, F is the scheduling benefit parameter, and α, β, and γ are weight coefficients. Preferably, α = 0.6, β = 0.3, and γ = 0.1.

[0119] In the present invention, the migration cost corresponding to each group is calculated based on the data volume of the first type of task; the SM utilization rate, other parameters, and the migration cost of the at least one computing power operation task are weighted to obtain the scheduling benefit parameter for each group; when the scheduling benefit parameter is greater than or equal to the set scheduling threshold, the first type of task is scheduled to the acceleration card where the second type of task is located. In this way, when the scheduling benefit parameter is greater than or equal to the set scheduling threshold, the computing power operation task is scheduled, so that the intelligent computing center can reduce resource waste after scheduling the computing power operation task, and greatly improve the resource utilization rate.

[0120] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In addition, the terms "comprising", "having" and any of their variations are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including the 7 cases of A alone, B alone, C alone, A and B both present, B and C both present, A and C both present, and A, B and C all present.

[0121] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.

[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of this application.

[0123] The above describes the embodiments of this application in conjunction with the accompanying drawings, but this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them belong to the protection scope of this application.

Claims

1. A method for asymmetric collaborative scheduling of computing power in intelligent computing centers for universal computing power, characterized in that: include: Step S1, obtaining the operation parameters corresponding to each computing power operation task in the multiple computing power operation tasks in the intelligent computing center, wherein the operation parameters include the streaming processor SM utilization and other parameters, and the other parameters include at least one of the video memory utilization, the unified computing device architecture CUDA core activity, the memory occupancy rate, and the bandwidth occupancy rate; Step S2: dividing the plurality of computing power operation tasks based on the operation parameters to obtain at least one group, wherein the operation parameters corresponding to each group meet the preset parameter conditions; Step S3: dispatching computing power running tasks included in different groups to different accelerator cards of the intelligent computing center, where each accelerator card of the intelligent computing center is used to execute computing power running tasks in one group; The computing power running tasks included in each group include first type tasks and second type tasks, wherein: The first type of task is a computing power running task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization is less than the first set video memory threshold; The second type of task is a computing power running task whose other parameters are less than the second set parameter threshold and whose SM utilization is greater than or equal to the second set video memory threshold; The step S3 comprises: Step S31: Schedule the first type of tasks to the accelerator card where the second type of tasks are located.

2. The method according to claim 1, characterized in that The preset parameter conditions include: The sum of the SM utilization rates corresponding to each computing power running task included in each group is greater than or equal to the second set threshold; The preset parameter conditions also include at least one of the following: The sum of the video memory utilization rates corresponding to each computing power running task included in each group is less than or equal to the first set threshold; The sum of the CUDA core activities corresponding to each computing power running task included in each group is less than or equal to the third set threshold; The sum of the memory occupancy rates corresponding to each computing power running task included in each group is less than or equal to a fourth set threshold; The sum of the bandwidth occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fifth set threshold.

3. The method according to claim 2, characterized in that The step S2 comprises: Step S21, traversing multiple groups in the intelligent computing center that are allowed to add computing power running tasks; Step S22: When the sum of the SM utilization of the first group and the SM utilization of the first computing power running task among the multiple groups is less than the first set threshold, the first computing power running task is added to the first group. The first computing power running task is the computing power running task to be added to the group among the multiple computing power running tasks.

4. The method according to claim 3, characterized in that The step S2 further comprises: Step S24: if the sum of the SM utilization rate of a second group and the SM utilization rate of the first computing power running task in the multiple groups is greater than or equal to the first set threshold, and the second sum of other parameters of the second group and other parameters of the first computing power running task meets the preset parameter conditions, add the first computing power running task to the second group; Step S25: After adding the first computing power running task to the second group, the second group is set as a group to which computing power running tasks are stopped being added.

5. The method according to claim 3, characterized in that The step S2 further comprises: Step S26: If the sum of the SM utilization of no group in the multiple groups and the SM utilization of the first computing power running task is less than the first set threshold, and the sum of the SM utilization of no group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task meets the preset parameter conditions, create a third group; Step S27: Add the first computing power running task to the third group.

6. The method according to claim 1, characterized in that The step S31 comprises: Step S311: calculating the migration cost corresponding to each group based on the data volume of the first type of tasks; Step S312: weighting the SM utilization rate, other parameters and the migration cost of the at least one computing power running task to obtain a scheduling benefit parameter of each group; Step S313: When the scheduling benefit parameter is greater than or equal to a set scheduling threshold, scheduling the first type of task to the accelerator card where the second type of task is located.

Citation Information

Patent Citations

  • Load-balanced power grid computing power network resource scheduling method and system

    CN117762640A

  • Automatic model parallel scheduling strategy generation method and device based on heterogeneous computing power

    CN118939391A