Intelligent computing center computing power asymmetric collaborative scheduling method oriented to common computing power
By obtaining and dividing the packets of computing power operation tasks in the intelligent computing center and scheduling them to different acceleration cards, the problems of low resource utilization and high leasing costs are solved, and the wide application of universal computing power is achieved.
Patent Information
- Application Number
- CN202510439992.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing intelligent computing center has low resource utilization rate and high computing power leasing costs, which leads to the limitation of the widespread application of universal computing power.
By obtaining the operating parameters of the computing power operation tasks in the intelligent computing center, different groups are divided, and the operating parameters of each group meet the preset conditions, the computing power operation tasks of different groups are scheduled to different acceleration cards of the intelligent computing center.
It improves the resource utilization rate of the acceleration card of the intelligent computing center, reduces the cost of users' leasing computing power services, and realizes the widespread application of universal computing power.
Smart Images

Figure CN119938284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to an asymmetric collaborative scheduling method for computing power of intelligent computing centers for universal computing power. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have come into being.
[0003] "Intelligent computing center" refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing needs. It is the computing power to process information data and achieve target result output. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] The current intelligent computing center provides acceleration cards to provide users with accelerated computing services. In the prior art, the intelligent computing center is equipped with multiple acceleration cards, and each acceleration card performs different computing power operation tasks. However, different users require different resources. For example, some computing power operation tasks require more video memory resources, while other computing power operation tasks require more computing power resources. In the process of the accelerator card of the intelligent computing center providing computing power services to users, there is a situation where the video memory resources or computing power resources of some accelerator cards are wasted, resulting in low resource utilization of the intelligent computing center. At the same time, since the computing power services provided by the intelligent computing center are limited by the number of acceleration cards, users usually need to rent the entire acceleration card when renting computing power services, resulting in high rental costs and difficulty in achieving widespread application of inclusive computing power.
[0008] It can be seen that the existing technology has the problem of low resource utilization of intelligent computing centers and high computing power leasing costs. Summary of the invention
[0009] The present invention provides an asymmetric collaborative scheduling method for computing power of an intelligent computing center for universal computing power, so as to solve the problems in the prior art of low resource utilization of the intelligent computing center and high computing power leasing cost.
[0010] To solve the above problems, the present invention is achieved as follows: In a first aspect, the present invention provides a method for asymmetric collaborative scheduling of computing power of intelligent computing centers for universal computing power, comprising: Step S1, obtaining the operation parameters corresponding to each computing power operation task in the multiple computing power operation tasks in the intelligent computing center, wherein the operation parameters include the streaming processor SM utilization and other parameters, and the other parameters include at least one of the video memory utilization, the unified computing device architecture CUDA core activity, the memory occupancy rate, and the bandwidth occupancy rate; Step S2: dividing the plurality of computing power operation tasks based on the operation parameters to obtain at least one group, wherein the operation parameters corresponding to each group meet the preset parameter conditions; Step S3: dispatch computing power running tasks included in different groups to different accelerator cards of the intelligent computing center, and each accelerator card of the intelligent computing center is used to execute computing power running tasks within a group.
[0011] In one embodiment, the preset parameter conditions include: The sum of the SM utilization rates corresponding to each computing power running task included in each group is greater than or equal to the second set threshold; The preset parameter conditions also include at least one of the following: The sum of the video memory utilization rates corresponding to each computing power running task included in each group is less than or equal to the first set threshold; The sum of the CUDA core activities corresponding to each computing power running task included in each group is less than or equal to the third set threshold; The sum of the memory occupancy rates corresponding to each computing power running task included in each group is less than or equal to a fourth set threshold; The sum of the bandwidth occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fifth set threshold.
[0012] In one embodiment, step S2 includes: Step S21, traversing multiple groups in the intelligent computing center that are allowed to add computing power running tasks; Step S22: When the sum of the SM utilization of the first group and the SM utilization of the first computing power running task among the multiple groups is less than the first set threshold, the first computing power running task is added to the first group. The first computing power running task is the computing power running task to be added to the group among the multiple computing power running tasks.
[0013] In one embodiment, the step S2 further includes: Step S24: if the sum of the SM utilization rate of a second group and the SM utilization rate of the first computing power running task in the multiple groups is greater than or equal to the first set threshold, and the second sum of other parameters of the second group and other parameters of the first computing power running task meets the preset parameter conditions, add the first computing power running task to the second group; Step S25: After adding the first computing power running task to the second group, the second group is set as a group to which computing power running tasks are stopped being added.
[0014] In one embodiment, the step S2 further includes: Step S26: If the sum of the SM utilization of no group in the multiple groups and the SM utilization of the first computing power running task is less than the first set threshold, and the sum of the SM utilization of no group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task meets the preset parameter conditions, create a third group; Step S27: Add the first computing power running task to the third group.
[0015] In one embodiment, the computing power running tasks included in each group include first type tasks and second type tasks, wherein: The first type of task is a computing power running task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization is less than the first set video memory threshold; The second type of task is a computing power running task whose other parameters are less than the second set parameter threshold and whose SM utilization is greater than or equal to the second set video memory threshold; The step S3 comprises: Step S31: Schedule the first type of tasks to the accelerator card where the second type of tasks are located.
[0016] In one embodiment, step S31 includes: Step S311: calculating the migration cost corresponding to each group based on the data volume of the first type of tasks; Step S312: weighting the SM utilization rate, other parameters and the migration cost of the at least one computing power running task to obtain a scheduling benefit parameter of each group; Step S313: When the scheduling benefit parameter is greater than or equal to a set scheduling threshold, scheduling the first type of task to the accelerator card where the second type of task is located.
[0017] In the present invention, the operating parameters corresponding to each computing power operation task in the multiple computing power operation tasks in the intelligent computing center are obtained, and the operating parameters include the streaming processor SM utilization and other parameters, and the other parameters include at least one of the video memory utilization, unified computing device architecture CUDA core activity, memory occupancy, and bandwidth occupancy; the multiple computing power operation tasks are divided based on the operating parameters to obtain at least one group, and the operating parameters corresponding to each group meet the preset parameter conditions; the computing power operation tasks included in different groups are dispatched to different accelerator cards of the intelligent computing center, and each accelerator card of the intelligent computing center is used to execute the computing power operation tasks in a group. In this way, by dividing multiple computing power operation tasks into different groups, and the accelerator card of the intelligent computing center executes the computing power operation tasks in the group, the accelerator card of the intelligent computing center can execute different computing power operation tasks at the same time, which greatly improves the resource utilization. At the same time, when renting computing power services, users only need to rent part of the resources of the accelerator card, which greatly reduces the rental cost, thereby realizing the wide application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0019] Figure 1 It is a flow chart of an asymmetric collaborative scheduling method for computing power of an intelligent computing center for universal computing power provided by the present invention. DETAILED DESCRIPTION
[0020] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data center to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and it mainly provides services to the society through computing power infrastructure.
[0022] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .
[0023] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capacity of the computing power facilities, including the comprehensive capabilities of network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0024] The "Storage Power" (SP) mentioned in the present invention refers to: the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB=2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.
[0025] The "computing power infrastructure" mentioned in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, which can realize centralized computing, storage, transmission and application of information.
[0026] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0027] The “computing power” mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0028] The “general computing power” mentioned in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0029] The "intelligent computing power" mentioned in the present invention refers to: a computing platform based on GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) and other dedicated chips for various innovative artificial intelligence applications, such as natural language processing, machine vision, etc.
[0030] The "super computing power" mentioned in the present invention refers to: the computing power mainly provided by high-performance computing clusters such as supercomputers. It uses the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0031] The "intelligent computing center" mentioned in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0032] The “intelligent computing center” mentioned in the present invention includes but is not limited to the “intelligent computing center”.
[0033] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0034] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0035] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide functions such as large-scale computing, storage and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0036] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0037] The "inclusive computing power" mentioned in the present invention refers to providing appropriate and effective computing power services at an affordable cost to all social classes and groups that have computing power service needs based on the requirements of equal opportunity and the principle of commercial sustainability.
[0038] The “model” mentioned in the present invention includes but is not limited to a “large language model” and a “multimodal large model”.
[0039] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale, designed to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0040] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, videos, audio, etc. for training, including but not limited to multimodal large language models.
[0041] The "asymmetric collaborative scheduling" described in the present invention refers to scheduling computing power to run tasks so that different parameters of the accelerator card reach an asymmetric balance state after the computing power is scheduled to run tasks. The asymmetric balance state is that different parameters respectively meet the corresponding set parameter ranges, and the set parameter ranges corresponding to different parameters can be the same or different. Through asymmetric collaborative scheduling, different parameters in the accelerator card can reach the set parameter range, which greatly improves the utilization rate of parameter resources.
[0042] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.
[0043] See also Figure 1 , Figure 1 : is a flow chart of an asymmetric collaborative scheduling method for intelligent computing center computing power for universal computing power provided by the present invention, such as Figure 1 As shown, the following steps are included: Step S1, obtaining the operating parameters corresponding to each computing power running task in the multiple computing power running tasks in the intelligent computing center, wherein the operating parameters include the streaming multiprocessor (SM) utilization and other parameters, and the other parameters include at least one of the video memory utilization, the unified computing device architecture (CUDA) core activity, the memory occupancy rate, and the bandwidth occupancy rate.
[0044] The above operating parameters are obtained by collecting the computing power operating tasks running in each accelerator card in the intelligent computing center. In some implementations, a real-time monitoring module is arranged in the intelligent computing center to monitor different computing power operating tasks running in different accelerator cards through the real-time monitoring module to obtain the operating parameters corresponding to each computing power operating task.
[0045] It should be noted that the operating parameters corresponding to different computing power running tasks are different, and the operating parameters corresponding to different types of computing power running tasks are also different. For example, for the model training computing power running tasks with a large amount of data, the SM utilization rate is large and the video memory occupancy rate is also large; for the model training computing power running tasks with a small amount of data, the SM utilization rate is small and the video memory occupancy rate is also small. By obtaining the operating parameters corresponding to different computing power running tasks, the division of different computing power running tasks by operating parameters can be realized.
[0046] The above-mentioned operating parameters include SM utilization and other parameters, and the other parameters may be one or more of video memory utilization, CUDA core activity, memory occupancy, and bandwidth occupancy. For example, the operating parameters include SM utilization and video memory utilization, and different computing power operating tasks are divided based on SM utilization and video memory utilization.
[0047] Step S2: divide the multiple computing power operation tasks based on the operation parameters to obtain at least one group, and the operation parameters corresponding to each group meet the preset parameter conditions.
[0048] It should be noted that the operating parameters include SM utilization and other parameters, and the other parameters include at least one of video memory utilization, CUDA core activity, memory occupancy, and bandwidth occupancy. The preset condition parameters are specifically set according to the parameters included in the operating parameters, so that when the operating parameters corresponding to each group meet the preset parameter conditions, the SM utilization and other parameters included in the operating parameters can achieve maximum utilization.
[0049] Specifically, in one embodiment, the preset parameter conditions include: The sum of the SM utilization rates corresponding to each computing power running task included in each group is greater than or equal to the second set threshold; The preset parameter conditions also include at least one of the following: The sum of the video memory utilization rates corresponding to each computing power running task included in each group is less than or equal to the first set threshold; The sum of the CUDA core activities corresponding to each computing power running task included in each group is less than or equal to the third set threshold; The sum of the memory occupancy rates corresponding to each computing power running task included in each group is less than or equal to a fourth set threshold; The sum of the bandwidth occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fifth set threshold.
[0050] In this way, the SM utilization rate of each group is made higher through the above method, while other parameters do not exceed the carrying range of the accelerator card, thereby improving the resource utilization rate as much as possible.
[0051] Each of the above groups includes at least one computing power running task. Among them, when each group includes only one computing power running task, the accelerator card has used up all kinds of resources as much as possible to execute the computing power running task, or the remaining resources are insufficient to execute other computing power running tasks. In this case, there is no need to schedule the computing power running task. When each group includes at least two computing power running tasks, the two computing power running tasks can be executed by one accelerator card at the same time, which can greatly improve the resource utilization rate of the accelerator card.
[0052] Step S3: dispatch computing power running tasks included in different groups to different accelerator cards of the intelligent computing center, and each accelerator card of the intelligent computing center is used to execute computing power running tasks within a group.
[0053] In the present invention, the operating parameters corresponding to each computing power operation task in the multiple computing power operation tasks in the intelligent computing center are obtained, and the operating parameters include the streaming processor SM utilization and other parameters, and the other parameters include at least one of the video memory utilization, unified computing device architecture CUDA core activity, memory occupancy, and bandwidth occupancy; the multiple computing power operation tasks are divided based on the operating parameters to obtain at least one group, and the operating parameters corresponding to each group meet the preset parameter conditions; the computing power operation tasks included in different groups are dispatched to different accelerator cards of the intelligent computing center, and each accelerator card of the intelligent computing center is used to execute the computing power operation tasks in a group. In this way, by dividing multiple computing power operation tasks into different groups, and the accelerator card of the intelligent computing center executes the computing power operation tasks in the group, the accelerator card of the intelligent computing center can execute different computing power operation tasks at the same time, which greatly improves the resource utilization. At the same time, when renting computing power services, users only need to rent part of the resources of the accelerator card, which greatly reduces the rental cost, thereby realizing the wide application of inclusive computing power.
[0054] In one embodiment, step S2 includes: Step S21, traversing multiple groups in the intelligent computing center that are allowed to add computing power running tasks; Step S22: When the sum of the SM utilization of the first group and the SM utilization of the first computing power running task among the multiple groups is less than the first set threshold, the first computing power running task is added to the first group. The first computing power running task is the computing power running task to be added to the group among the multiple computing power running tasks.
[0055] The above multiple groups are groups that allow adding computing power running tasks. It should be noted that the intelligent computing center includes groups that allow adding computing power running tasks and groups that stop adding computing power running tasks. Among them, the group that allows adding computing power running tasks is the group whose resources have not been used up. At this time, computing power running tasks can still be added to the group; the group that stops adding computing power running tasks is the group where a certain resource has been used up. At this time, since a certain resource in the group has been used up, adding computing power running tasks will cause insufficient system resources and cause a crash, so computing power running tasks cannot be added at this time.
[0056] In some embodiments, after the computing power running task in a group is executed, the running parameters of other computing power running tasks in the group are recalculated, and based on the running parameters of other computing power running tasks, it is determined whether the group should run the added computing power running task.
[0057] The sum of the SM utilization of the first group and the SM utilization of the first computing power running task is used to characterize the SM utilization of the first group after the first computing power running task is added to the first group. It should be noted that when the sum of the SM utilization of the first group and the SM utilization of the first computing power running task is less than the first set threshold, the first group occupies less resources of the accelerator card, and the computing power running task can be directly added to the first group, and the resources of the intelligent computing center accelerator card can effectively execute the various computing power running tasks included in the first group.
[0058] In one embodiment, the step S2 further includes: Step S24: if the sum of the SM utilization rate of a second group and the SM utilization rate of the first computing power running task in the multiple groups is greater than or equal to the first set threshold, and the second sum of other parameters of the second group and other parameters of the first computing power running task meets the preset parameter conditions, add the first computing power running task to the second group; Step S25: After adding the first computing power running task to the second group, the second group is set as a group to which computing power running tasks are stopped being added.
[0059] It should be noted that when the sum of the SM utilization of the second group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, the SM utilization reaches the limit of the SM resources that the accelerator card can provide. At this time, it is necessary to determine whether other parameters reach the preset parameter conditions to confirm whether the computing power running task needs to be added to the second group.
[0060] In the present invention, when the sum of the SM utilization of the second group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the second group and other parameters of the first computing power running task meets the preset parameter conditions, after the first computing power running task is added to the second group, the SM utilization and other parameters of the second group can meet the preset parameter conditions. At this time, the first computing power running task is added to the second group, and the accelerator card of the intelligent computing center can execute all computing power running tasks in the second group. However, in this case, since the SM utilization and other parameters of the second group can meet the preset parameter conditions, adding computing power running tasks to the second group will cause a parameter of the accelerator card executing the second group to exceed the limit, resulting in insufficient system resources and crash. Therefore, it is impossible to add computing power running tasks at this time, and the second group needs to be set as a group that stops adding computing power running tasks to maintain the stability of the second group.
[0061] In one embodiment, the step S2 further includes: Step S26: If the sum of the SM utilization of no group in the multiple groups and the SM utilization of the first computing power running task is less than the first set threshold, and the sum of the SM utilization of no group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task meets the preset parameter conditions, create a third group; Step S27: Add the first computing power running task to the third group.
[0062] It should be noted that when the sum of the SM utilization of the group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task does not meet the preset parameter conditions, the group can add other computing power running tasks at this time, but adding the first computing power running task to the group will cause a parameter of the accelerator card executing the group to exceed the limit, resulting in insufficient system resources and a crash. Therefore, the first computing power running task cannot be added to the group at this time, and it is necessary to select other groups to add the first computing power running task.
[0063] In the present invention, when the sum of the SM utilization of the non-existent group in the multiple groups and the SM utilization of the first computing power running task is less than the first set threshold, and the sum of the SM utilization of the non-existent group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task meets the preset parameter conditions, at this time, there is no suitable group for adding the first computing power running task among the multiple groups for adding computing power running tasks in the intelligent computing center, and at this time, it is necessary to create a new third group to add the first computing power running task through the third group.
[0064] In one embodiment, the computing power running tasks included in each group include first type tasks and second type tasks, wherein: The first type of task is a computing power running task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization is less than the first set video memory threshold; The second type of task is a computing power running task whose other parameters are less than the second set parameter threshold and whose SM utilization is greater than or equal to the second set video memory threshold; The step S3 comprises: Step S31: Schedule the first type of tasks to the accelerator card where the second type of tasks are located.
[0065] The parameters between the first type of tasks and the second type of tasks are complementary, and the scheduling method of the computing power running tasks is determined by distinguishing the first type of tasks from the second type of tasks.
[0066] For example, other parameters include video memory occupancy rate. The first type of task is a computing power running task with a video memory occupancy rate greater than 60% and an SM utilization rate less than 30%; the second type of task is a computing power running task with a video memory occupancy rate less than 40% and an SM utilization rate greater than 65%. In this case, the first type of task is a video memory intensive task, and the second type of task is a computing intensive task. At this time, the first type of task is scheduled to the accelerator card where the second type of task is located. The traffic of the first type of task can be switched to the new accelerator card through load balancing and grayscale release; the first type of task in the old accelerator card is offline to complete the scheduling of the computing power running task.
[0067] In some other implementations, the second type of tasks may also be scheduled to the accelerator card where the first type of tasks are located.
[0068] In one embodiment, step S31 includes: Step S311: calculating the migration cost corresponding to each group based on the data volume of the first type of tasks; Step S312: weighting the SM utilization rate, other parameters and the migration cost of the at least one computing power running task to obtain a scheduling benefit parameter of each group; Step S313: When the scheduling benefit parameter is greater than or equal to a set scheduling threshold, scheduling the first type of task to the accelerator card where the second type of task is located.
[0069] It should be noted that scheduling computing power to run tasks will consume resources of the intelligent computing center. If the consumed resources are greater than the resources saved after scheduling computing power to run tasks, then scheduling computing power to run tasks will reduce the resource utilization of the intelligent computing center. Therefore, after dividing the groups, it is also necessary to calculate the scheduling benefit parameters to determine whether scheduling is required.
[0070] In some implementations, the migration cost may be calculated by parameters such as the data volume, execution times, and transmission bandwidth of the first type of task. For example, the migration cost of the first type of task may be calculated by the following formula: Migration cost = k 1 × amount of data + k 2 × number of executions + k 3 ×Transmission bandwidth; Among them, k 1 , k 2 , k 3 is the weight coefficient.
[0071] The above scheduling benefit parameters can be calculated by the following formula: F = α × (1-other parameters) + β × SM utilization + γ × migration cost; Wherein, F is the scheduling benefit parameter, α, β, and γ are weight coefficients, preferably, α=0.6, β=0.3, and γ=0.1.
[0072] In the present invention, the migration cost corresponding to each group is calculated based on the data volume of the first type of task; the SM utilization rate, other parameters and the migration cost of the at least one computing power running task are weighted to obtain the scheduling benefit parameter of each group; when the scheduling benefit parameter is greater than or equal to the set scheduling threshold, the first type of task is scheduled to the accelerator card where the second type of task is located. In this way, the computing power running task is scheduled when the scheduling benefit parameter is greater than or equal to the set scheduling threshold, so that after scheduling the computing power running task, the intelligent computing center can reduce the waste of resources and greatly improve the resource utilization.
[0073] The terms "first", "second" etc. in the present invention are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment. In addition, "and / or" is used in the present application to represent at least one of the connected objects, such as A and / or B and / or C, which means to include 7 situations including single A, single B, single C, and A and B all exist, B and C all exist, A and C all exist, and A, B and C all exist.
[0074] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0075] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a second terminal device, etc.) to execute the methods of each embodiment of the present application.
[0076] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for asymmetric collaborative scheduling of computing power in intelligent computing centers for universal computing power, characterized in that: include: Step S1, obtaining the operation parameters corresponding to each computing power operation task in the multiple computing power operation tasks in the intelligent computing center, wherein the operation parameters include the streaming processor SM utilization and other parameters, and the other parameters include at least one of the video memory utilization, the unified computing device architecture CUDA core activity, the memory occupancy rate, and the bandwidth occupancy rate; Step S2: dividing the plurality of computing power operation tasks based on the operation parameters to obtain at least one group, wherein the operation parameters corresponding to each group meet the preset parameter conditions; Step S3: dispatch computing power running tasks included in different groups to different accelerator cards of the intelligent computing center, and each accelerator card of the intelligent computing center is used to execute computing power running tasks within a group.
2. The method according to claim 1, characterized in that The preset parameter conditions include: The sum of the SM utilization rates corresponding to each computing power running task included in each group is greater than or equal to the second set threshold; The preset parameter conditions also include at least one of the following: The sum of the video memory utilization rates corresponding to each computing power running task included in each group is less than or equal to the first set threshold; The sum of the CUDA core activities corresponding to each computing power running task included in each group is less than or equal to the third set threshold; The sum of the memory occupancy rates corresponding to each computing power running task included in each group is less than or equal to a fourth set threshold; The sum of the bandwidth occupancy rates corresponding to each computing power running task included in each group is less than or equal to the fifth set threshold.
3. The method according to claim 2, characterized in that The step S2 comprises: Step S21, traversing multiple groups in the intelligent computing center that are allowed to add computing power running tasks; Step S22: When the sum of the SM utilization of the first group and the SM utilization of the first computing power running task among the multiple groups is less than the first set threshold, the first computing power running task is added to the first group. The first computing power running task is the computing power running task to be added to the group among the multiple computing power running tasks.
4. The method according to claim 3, characterized in that The step S2 further comprises: Step S24: if the sum of the SM utilization rate of a second group and the SM utilization rate of the first computing power running task in the multiple groups is greater than or equal to the first set threshold, and the second sum of other parameters of the second group and other parameters of the first computing power running task meets the preset parameter conditions, add the first computing power running task to the second group; Step S25: After adding the first computing power running task to the second group, the second group is set as a group to which computing power running tasks are stopped being added.
5. The method according to claim 3, characterized in that The step S2 further comprises: Step S26: If the sum of the SM utilization of no group in the multiple groups and the SM utilization of the first computing power running task is less than the first set threshold, and the sum of the SM utilization of no group and the SM utilization of the first computing power running task is greater than or equal to the first set threshold, and the second sum of other parameters of the group and other parameters of the first computing power running task meets the preset parameter conditions, create a third group; Step S27: Add the first computing power running task to the third group.
6. The method according to any one of claims 1 to 5, characterized in that The computing power running tasks included in each group include first type tasks and second type tasks, wherein: The first type of task is a computing power running task whose other parameters are greater than or equal to the first set parameter threshold and whose SM utilization is less than the first set video memory threshold; The second type of task is a computing power running task whose other parameters are less than the second set parameter threshold and whose SM utilization is greater than or equal to the second set video memory threshold; The step S3 comprises: Step S31: Schedule the first type of tasks to the accelerator card where the second type of tasks are located.
7. The method according to claim 6, characterized in that The step S31 comprises: Step S311: calculating the migration cost corresponding to each group based on the data volume of the first type of tasks; Step S312: weighting the SM utilization rate, other parameters and the migration cost of the at least one computing power running task to obtain a scheduling benefit parameter of each group; Step S313: When the scheduling benefit parameter is greater than or equal to a set scheduling threshold, scheduling the first type of task to the accelerator card where the second type of task is located.
Citation Information
Patent Citations
Method and device for server task scheduling
CN106897132A
Inter-board multi-operation chip calculation power balance system applied to electric power instrument equipment
CN113641468A
Load-balanced power grid computing power network resource scheduling method and system
CN117762640A
Automatic model parallel scheduling strategy generation method and device based on heterogeneous computing power
CN118939391A
Organizing Task Placement Based On Workload Characterizations
US20120180061A1
Cited By
Multi-agent computing power scheduling method and device for intelligent computing center cloud platform
CN120256148A
Method and device for multi-agent scheduling computing power of intelligent computing center cloud platform
CN120256148B