Intelligent computing power scheduling method and device for intelligent computing center of common computing power

By receiving computing power operation tasks and the number of required resources in the intelligent computing center, determining the target nodes and scheduling tasks, the problems of low computing power resource utilization and high leasing costs are solved, and efficient utilization and universal application of computing power resources are achieved.

CN120029739APending Publication Date: 2025-05-23DATACANVAS LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510502648.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The low utilization rate of computing power resources and the high cost of leasing computing power resources in the intelligent computing center makes it difficult to achieve the widespread application of universal computing power.

Method used

By receiving computing power to run tasks and the number of required resources, determine multiple first nodes that meet the number of required resources, obtain the remaining resources of each node in NUMA channel, select the target node to schedule tasks, and reduce cross-channel scheduling losses.

Benefits of technology

It improves the utilization rate of computing power resources, reduces the economic cost of users leasing computing power resources services, and realizes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029739A_ABST
    Figure CN120029739A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent computing power scheduling method and device for a common computing power-oriented intelligent computing center, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the steps: S1, receiving a computing power operation task, and the number of required resources corresponding to the computing power operation task; s2, determining a plurality of first nodes based on the quantity of the required resources, wherein the plurality of first nodes are nodes meeting the quantity of the required resources in the intelligent computing center; s3, obtaining the number of residual resources corresponding to a non-uniform memory access (NUMA) channel included in each first node; step S4, determining a target node from the plurality of first nodes based on the number of residual resources; and S5, scheduling the computing power operation task to the target node, wherein the target node is used for executing the computing power operation task. According to the method, the utilization rate of computing power resources can be greatly improved, and wide application of the general computing power is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for intelligent scheduling of computing power in an intelligent computing center for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process parameters, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing parameter data, and a new type of productive force integrating parameter computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] An intelligent computing center is used to provide computing power resources to users to execute computing power operation tasks. In the prior art, an intelligent computing center has multiple nodes, and each node is also provided with multiple channels to provide computing power resources through different channels. However, in the prior art, when a node executes a computing power operation task, the node may schedule the computing power resources of different channels to execute the computing power operation task, and scheduling the computing power resources of different channels to execute the task will cause a large amount of loss of computing power resources during the scheduling process, resulting in a very low utilization rate of computing power resources. At the same time, due to a large amount of loss of computing power resources when the intelligent computing center provides computing power services, users need to lease more computing power services to complete the execution of computing power operation tasks, resulting in a very high leasing cost and making it difficult to achieve the wide application of inclusive computing power.

[0008] It can be seen that in the prior art, there are problems of very low utilization rate of computing power resources in the intelligent computing center and very high economic cost of leasing computing power resources. Summary of the invention

[0009] The present invention provides a computing power intelligent scheduling method and device for a universal computing power intelligent computing center, so as to solve the problems in the prior art that the computing power resource utilization rate of the intelligent computing center is very low and the computing power resource leasing cost is very high.

[0010] To solve the above problems, the present invention is achieved as follows: In a first aspect, the present invention provides a method for intelligent scheduling of computing power for an intelligent computing center for universal computing power, comprising: Step S1: receiving a computing power operation task and the required resource quantity corresponding to the computing power operation task; Step S2: determining a plurality of first nodes based on the required number of resources, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required number of resources; Step S3, obtaining the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node; Step S4, determining a target node from the plurality of first nodes based on the amount of remaining resources; Step S5: dispatch the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

[0011] In one embodiment, the required resource quantity includes the required number of graphics processors GPU, and step S2 includes: Step S21, traversing multiple second nodes of the intelligent computing center to obtain the number of first GPUs corresponding to the multiple second nodes; Step S22: Determine the multiple first nodes from the multiple second nodes, and the number of first GPUs corresponding to each first node matches the required number of GPUs.

[0012] In one embodiment, the required resource quantity also includes the required other resource quantity, and step S3 includes: Step S31, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channels included in each first node; The step S4 comprises: Step S41: if there is a target NUMA channel, set the first node where the target NUMA channel is located as the target node; The number of second GPUs corresponding to the target NUMA channel matches the required number of GPUs, and the number of remaining other resources corresponding to the target NUMA channel matches the required number of other resources.

[0013] In one embodiment, step S31 includes: Step S311: Obtain a fragmentation degree parameter corresponding to each first node, where the fragmentation degree parameter is used to characterize whether different GPUs in the first node are in use; Step S312: sorting the plurality of first nodes based on the corresponding fragmentation degree parameters; Step S313, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each of the sorted plurality of first nodes; The step S41 comprises: Step S411: After obtaining the second GPU quantity and the remaining other resource quantity corresponding to the target NUMA channel, set the first node where the target NUMA channel is located as the target node, and stop obtaining the second GPU quantity and the remaining other resource quantity corresponding to other channels.

[0014] In one embodiment, step S4 further includes: Step S42: if the target NUMA channel does not exist, obtain an intermediate NUMA channel included in each first node, where the intermediate NUMA channel is a NUMA channel with the most GPU resources represented by the second number of GPUs; Step S43, calculating the number of scheduling computing resources corresponding to each first node based on the number of second GPUs and the number of remaining other resources corresponding to the middle NUMA channel, as well as the required number of GPUs and the required number of other resources; Step S44: Set the first node with the smallest amount of scheduling computing resources as the target node.

[0015] In one embodiment, step S43 includes: Step S431, calculating a first difference between the required number of GPUs and the second number of GPUs; Step S432, calculating a second difference between the required quantity of other resources and the remaining quantity of other resources; Step S433: weight the first difference and the second difference to obtain the number of scheduling computing resources corresponding to each first node.

[0016] In a second aspect, the present invention further provides a computing power intelligent scheduling device for a universal computing power intelligent computing center, comprising: A receiving module, used to receive a computing power operation task and the quantity of required resources corresponding to the computing power operation task; A first determination module is used to determine a plurality of first nodes based on the required resource quantity, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required resource quantity; An acquisition module, used for acquiring the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node; A second determination module, configured to determine a target node from the plurality of first nodes based on the remaining resource quantity; The scheduling module is used to schedule the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

[0017] In a third aspect, the present invention further provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps in the method for intelligent scheduling of computing power for an intelligent computing center for universal computing power as described in the first aspect above are implemented.

[0018] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method for intelligent scheduling of computing power for an intelligent computing center for universal computing power as described in the first aspect above are implemented.

[0019] In a fifth aspect, the present invention further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps in the method for intelligent scheduling of computing power for an intelligent computing center for universal computing power as described in the first aspect above.

[0020] In the present invention, a computing power operation task and the number of required resources corresponding to the computing power operation task are received; multiple first nodes are determined based on the number of required resources, and the multiple first nodes are nodes in the intelligent computing center that meet the number of required resources; the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node is obtained; the target node is determined from the multiple first nodes based on the number of remaining resources; the computing power operation task is scheduled to the target node, and the target node is used to execute the computing power operation task. In this way, the target node is determined by the number of remaining resources, so that the target node does not need to schedule computing power resources across channels, or schedules computing power resources as little as possible, thereby greatly improving the utilization rate of computing power resources. At the same time, it can also greatly reduce the economic cost of users leasing computing power resource services, thereby realizing the widespread application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0022] Figure 1 It is a flow chart of a computing power intelligent scheduling method for a universal computing power intelligent computing center provided by the present invention; Figure 2 It is a structural diagram of a computing power intelligent scheduling device for a universal computing power intelligent computing center provided by the present invention; Figure 3 It is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data center to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and it mainly provides services to the society through computing power infrastructure.

[0025] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0026] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capacity of the computing power facilities, including the comprehensive capabilities of network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0027] The "Storage Power" (SP) mentioned in the present invention refers to: the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB=2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0028] The "computing power infrastructure" mentioned in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, which can realize centralized computing, storage, transmission and application of information.

[0029] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0030] The “computing power” mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0031] The “general computing power” mentioned in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0032] The "intelligent computing power" mentioned in the present invention refers to: a computing platform based on GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) and other dedicated chips for various innovative artificial intelligence applications, such as natural language processing, machine vision, etc.

[0033] The "super computing power" mentioned in the present invention refers to: the computing power mainly provided by high-performance computing clusters such as supercomputers. It uses the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0034] The "intelligent computing center" mentioned in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0035] The “intelligent computing center” mentioned in the present invention includes but is not limited to the “intelligent computing center”.

[0036] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0037] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0038] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide functions such as large-scale computing, storage and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0039] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0040] The "inclusive computing power" mentioned in the present invention refers to providing appropriate and effective computing power services at an affordable cost to all social classes and groups that have computing power service needs based on the requirements of equal opportunity and the principle of commercial sustainability.

[0041] The “model” mentioned in the present invention includes but is not limited to a “large language model” and a “multimodal large model”.

[0042] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale, designed to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0043] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, videos, audio, etc. for training, including but not limited to multimodal large language models.

[0044] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.

[0045] See also Figure 1 , Figure 1 It is a flowchart of a method for intelligent scheduling of computing power for an intelligent computing center of universal computing power provided by the present invention. Figure 1 As shown, the following steps are included: Step S1: Receive a computing power operation task and the amount of required resources corresponding to the computing power operation task.

[0046] The above computing power operation tasks are tasks sent by users to the intelligent computing center through the terminal. After receiving the computing power operation tasks, the intelligent computing center needs to dispatch computing power resources to process the computing power operation tasks. Among them, when the user sends the computing power operation tasks through the terminal, the required resource quantity needs to be sent at the same time, so that the intelligent computing center can determine the computing power resources that need to be dispatched according to the required resource quantity.

[0047] Among them, the required number of resources may include at least one of the operating environment parameters, the required number of GPU resources, the required number of CPU resources, the required number of memory, the required number of disk storage and the required network parameters. The intelligent computing center receives the required number of resources corresponding to the computing power running task to schedule the corresponding resources based on the operating environment parameters, the required number of GPU resources, the required number of CPU resources, the required number of memory, the required number of disk storage and / or the required network parameters to execute the computing power running task.

[0048] Step S2: determine a plurality of first nodes based on the required number of resources, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required number of resources.

[0049] The above-mentioned multiple first nodes are nodes used to execute computing power operation tasks in the intelligent computing center. It should be noted that there are multiple nodes in the intelligent computing center, each of which can be used to execute computing power operation tasks, and computing power services are provided to different users through nodes. However, when a user applies for computing power services from the intelligent computing center, some nodes in the intelligent computing center are in the state of processing computing power operation tasks of other users. At this time, these nodes cannot provide sufficient computing power resources to process new computing power operation tasks. Therefore, when the intelligent computing center receives a new computing power operation task, it is necessary to first determine multiple first nodes that meet the required number of resources, and then determine the target node for processing the computing power operation task from the multiple first nodes, thereby improving the efficiency of determining the target node.

[0050] Step S3: Obtain the remaining resource quantity corresponding to the Non-Uniform Memory Access (NUMA) channel included in each first node.

[0051] It should be noted that each first node includes at least one NUMA channel, each NUMA channel corresponds to multiple GPUs and other resources, and the execution of computing power running tasks is achieved by scheduling the resources of the NUMA channels in the node. In the specific scheduling process, if the computing power resources are scheduled across channels, there will be more scheduling losses, so it is necessary to determine the number of remaining resources corresponding to different NUMA channels, and determine the target node based on the number of remaining resources, so as to achieve the scheduling of the computing power resources of a NUMA channel to process the computing power running tasks as much as possible, and reduce the losses caused by cross-channel scheduling.

[0052] The remaining resource quantity is the resource quantity corresponding to the required resource quantity. The remaining resource quantity can be used to determine whether the NUMA channel can provide sufficient computing resources. For example, if the required resource quantity includes the required GPU quantity and the required CPU quantity, the remaining resource quantity is the remaining GPU quantity and the remaining CPU quantity. The remaining GPU quantity and the remaining CPU quantity are used to determine whether there is a single NUMA channel that meets the required resource quantity.

[0053] Step S4: determine a target node from the multiple first nodes based on the remaining resource quantity.

[0054] In some embodiments, a target node is determined from multiple first nodes based on the number of remaining resources. There is a NUMA channel in the target node whose number of remaining resources corresponds to the required number of resources. The computing power running task can be executed by scheduling the computing power resources of the NUMA channel, thereby avoiding cross-channel scheduling of computing power resources and improving the utilization of computing power resources.

[0055] In some embodiments, a target node is determined from multiple first nodes based on the number of remaining resources. The target node is the node that requires the least cross-scheduling computing resources when multiple first nodes all need to schedule computing resources across channels. That is, the target node is the node of the NUMA channel with the largest number of remaining resources. By scheduling as few computing resources as possible, the loss in scheduling computing resources is reduced and the utilization rate of computing resources is improved.

[0056] Step S5: dispatch the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

[0057] In the present invention, a computing power operation task and the number of required resources corresponding to the computing power operation task are received; multiple first nodes are determined based on the number of required resources, and the multiple first nodes are nodes in the intelligent computing center that meet the number of required resources; the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node is obtained; the target node is determined from the multiple first nodes based on the number of remaining resources; the computing power operation task is scheduled to the target node, and the target node is used to execute the computing power operation task. In this way, the target node is determined by the number of remaining resources, so that the target node does not need to schedule computing power resources across channels, or schedules computing power resources as little as possible, thereby greatly improving the utilization rate of computing power resources. At the same time, improving the utilization rate of computing power resources can greatly reduce the cost of users renting computing power services, thereby realizing the widespread application of inclusive computing power.

[0058] In one embodiment, the required resource quantity includes the required number of graphics processors GPU, and step S2 includes: Step S21, traversing multiple second nodes of the intelligent computing center to obtain the number of first GPUs corresponding to the multiple second nodes; Step S22: Determine the multiple first nodes from the multiple second nodes, and the number of first GPUs corresponding to each first node matches the required number of GPUs.

[0059] The above-mentioned multiple second nodes are nodes used to execute computing power running tasks in the intelligent computing center. Among the multiple second nodes, some nodes are in the state of executing computing power running tasks and cannot meet the needs of executing new computing power running tasks. At this time, it is necessary to filter these second nodes first to obtain multiple first nodes, and then filter the target nodes from the multiple first nodes to improve the screening efficiency.

[0060] In the present invention, multiple second nodes of the intelligent computing center are traversed to obtain the number of first GPUs corresponding to the multiple second nodes; the multiple first nodes are determined from the multiple second nodes, and the number of first GPUs corresponding to each first node matches the required number of GPUs. In this way, multiple first nodes can be quickly determined from the multiple second nodes of the intelligent computing center by the number of GPUs.

[0061] It should be noted that computing power resources are provided by GPU acceleration cards in the intelligent computing center. When executing computing power operation tasks, the GPU resources of the node need to be occupied. Therefore, in the present invention, multiple first nodes are quickly determined by the required number of GPUs and the number of first GPUs.

[0062] In one embodiment, the required resource quantity also includes the required other resource quantity, and step S3 includes: Step S31, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channels included in each first node; The step S4 comprises: Step S41: if there is a target NUMA channel, set the first node where the target NUMA channel is located as the target node; The number of second GPUs corresponding to the target NUMA channel matches the required number of GPUs, and the number of remaining other resources corresponding to the target NUMA channel matches the required number of other resources.

[0063] The first number of GPUs mentioned above is used to characterize the GPU resources that the entire node can provide, while the second number of GPUs is used to characterize the GPU resources that the NUMA channel can provide. It should be noted that the GPU resources that the first node can provide are jointly provided by the included NUMA channels. There is a situation where the first node can provide GPU resources that meet the required number of GPUs, but there is no NUMA channel in the first node that can meet the required number of GPUs. Therefore, it is necessary to screen from multiple first nodes to determine whether there is a NUMA channel that meets the required number of resources. If there is a target NUMA channel that meets the required number of resources, the first node where the target NUMA channel is located is set as the target node, so that the target node does not need to schedule computing resources across channels to execute computing power operation tasks, which greatly improves the utilization rate of computing power resources.

[0064] Furthermore, the required resource quantity also includes the required other resource quantity, such as CPU resources or memory resources, etc. When selecting the target node, it is also necessary to consider whether the NUMA channel meets the required other resource quantity.

[0065] Specifically, in the present invention, the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each first node are obtained; in the case of a target NUMA channel, the first node where the target NUMA channel is located is set as the target node; wherein the number of second GPUs corresponding to the target NUMA channel matches the required number of GPUs, and the number of remaining other resources corresponding to the target NUMA channel matches the required number of other resources. In this way, the target NUMA channel is obtained by screening, and then the first node where the target NUMA channel is located is set as the target node, so that the target node can directly process computing power running tasks through the computing power resources of the target NUMA channel, without the need to schedule computing power resources across channels, thereby greatly improving the utilization rate of computing power resources.

[0066] In one embodiment, step S31 includes: Step S311: Obtain a fragmentation degree parameter corresponding to each first node, where the fragmentation degree parameter is used to characterize whether different GPUs in the first node are in use; Step S312: sorting the plurality of first nodes based on the corresponding fragmentation degree parameters; Step S313, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each of the sorted plurality of first nodes; The step S41 comprises: Step S411: After obtaining the second GPU quantity and the remaining other resource quantity corresponding to the target NUMA channel, set the first node where the target NUMA channel is located as the target node, and stop obtaining the second GPU quantity and the remaining other resource quantity corresponding to other channels.

[0067] It should be noted that different nodes have different degrees of fragmentation. The node's degree of fragmentation parameter is used to characterize whether the GPU is in use. For example, the larger the degree of fragmentation parameter, the fewer GPUs are in use in the node, and the smaller the degree of fragmentation parameter, the more GPUs are in use in the node. When determining the target node, it is necessary to use the node with a low degree of fragmentation to perform the computing power operation task as much as possible, so that other nodes can perform the computing power operation tasks that require more computing power resources.

[0068] In the present invention, the fragmentation degree parameter corresponding to each first node is obtained, and the fragmentation degree parameter is used to characterize whether different GPUs in the first node are used; the multiple first nodes are sorted based on the corresponding fragmentation degree parameter; the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each of the sorted multiple first nodes are obtained; after the number of second GPUs and the number of remaining other resources corresponding to the target NUMA channel are obtained, the first node where the target NUMA channel is located is set as the target node, and the number of second GPUs and the number of remaining other resources corresponding to other channels are stopped. In this way, the first nodes are sorted by the fragmentation degree parameter, and it is first determined whether the first node with a small fragmentation degree parameter has a target NUMA channel, so that the target node finally determined is the node with the smallest fragmentation degree parameter when the NUMA channel meets the required resource quantity.

[0069] In one embodiment, the step S4 further includes: Step S42: if the target NUMA channel does not exist, obtain an intermediate NUMA channel included in each first node, where the intermediate NUMA channel is a NUMA channel with the most GPU resources represented by the second number of GPUs; Step S43, calculating the number of scheduling computing resources corresponding to each first node based on the number of second GPUs and the number of remaining other resources corresponding to the middle NUMA channel, as well as the required number of GPUs and the required number of other resources; Step S44: Set the first node with the smallest amount of scheduling computing resources as the target node.

[0070] It should be noted that when there is no target NUMA channel, that is, when no NUMA channel meets the required number of resources, it is necessary to schedule computing resources across channels. In order to reduce the loss caused by scheduling computing resources across channels, it is necessary to schedule computing resources across channels as little as possible to improve the utilization of computing resources.

[0071] Specifically, in the present invention, in the absence of the target NUMA channel, the intermediate NUMA channel included in each first node is obtained, and the intermediate NUMA channel is the NUMA channel with the most GPU resources represented by the second GPU number; based on the second GPU number and the remaining other resource number corresponding to the intermediate NUMA channel, as well as the required GPU number and the required other resource number, the number of scheduling computing resources corresponding to each first node is calculated; and the first node with the smallest number of scheduling computing resources is set as the target node. In this way, by calculating the number of scheduling computing resources, the number of scheduling computing resources is used to represent the situation where scheduling resources are required, and the first node with the smallest number of scheduling computing resources is set as the target node, so that the target node can perform computing operation tasks with the least computing resource loss when scheduling computing resources across channels, thereby improving the utilization rate of computing resources.

[0072] For example, the computing power operation task requires the resources of 4 GPUs, one channel in node A includes 3 GPU resources, and the other includes 1 GPU resource, and one channel in node B includes 2 GPU resources, and the other also includes 2 GPU resources. The method of the present invention calculates that the number of scheduled computing power resources of node A is 1 GPU resource, and the number of scheduled computing power resources of node B is 2 GPU resources. At this time, the number of scheduled computing power resources of node A is smaller, and the computing power resource loss caused is less. Node A can be set as the target node to improve the utilization rate of computing power resources.

[0073] In one embodiment, step S43 includes: Step S431, calculating a first difference between the required number of GPUs and the second number of GPUs; Step S432, calculating a second difference between the required quantity of other resources and the remaining quantity of other resources; Step S433: weight the first difference and the second difference to obtain the number of scheduling computing resources corresponding to each first node.

[0074] In the present invention, the first difference between the required number of GPUs and the second number of GPUs is calculated; the second difference between the required number of other resources and the remaining number of other resources is calculated; the first difference and the second difference are weighted to obtain the number of scheduling computing resources corresponding to each first node. In this way, the number of GPUs and the number of other resources are combined to calculate the number of scheduling computing resources corresponding to each node.

[0075] Furthermore, the first difference and the second difference are weighted to obtain the number of scheduling computing resources corresponding to each first node, and the weight coefficients of the first difference and the second difference can be set according to demand. For example, when other resources of the intelligent computing center are sufficient and only GPU resources are underutilized, the weight coefficient of the first difference can be increased, and / or the weight coefficient of the second difference can be reduced.

[0076] See also Figure 2 , Figure 2 This is a structural diagram of a computing power intelligent scheduling device for a universal computing power intelligent computing center provided by the present invention, such as Figure 2 As shown, the computing power intelligent scheduling device 200 for the inclusive computing power intelligent computing center includes: The receiving module 201 is used to receive a computing power operation task and the required resource quantity corresponding to the computing power operation task; A first determination module 202 is used to determine a plurality of first nodes based on the required number of resources, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required number of resources; An acquisition module 203 is used to acquire the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node; A second determination module 204, configured to determine a target node from the plurality of first nodes based on the remaining resource quantity; The scheduling module 205 is used to schedule the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

[0077] In one embodiment, the required resource quantity includes the required number of graphics processors GPU, and the first determination module 202 includes: A traversal unit, used to traverse multiple second nodes of the intelligent computing center and obtain the number of first GPUs corresponding to the multiple second nodes; The first determining unit is used to determine the multiple first nodes from the multiple second nodes, and the number of first GPUs corresponding to each first node matches the required number of GPUs.

[0078] In one embodiment, the required resource quantity also includes the required other resource quantity, and the acquisition module 203 includes: A first acquisition unit, configured to acquire the number of second GPUs and the number of remaining other resources corresponding to the NUMA channels included in each first node; The second determining module 204 includes: A second determining unit is configured to set a first node where the target NUMA channel is located as the target node when there is a target NUMA channel; The number of second GPUs corresponding to the target NUMA channel matches the required number of GPUs, and the number of remaining other resources corresponding to the target NUMA channel matches the required number of other resources.

[0079] In one embodiment, the first acquiring unit includes: A first acquisition subunit is used to acquire a fragmentation degree parameter corresponding to each first node, where the fragmentation degree parameter is used to characterize whether different GPUs in the first node are used; A sorting subunit, configured to sort the plurality of first nodes based on the corresponding fragmentation degree parameters; A second acquisition subunit is used to acquire the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each of the sorted plurality of first nodes; The second determining unit includes: The processing subunit is used to, after obtaining the second GPU quantity and the remaining other resource quantity corresponding to the target NUMA channel, set the first node where the target NUMA channel is located as the target node, and stop obtaining the second GPU quantity and the remaining other resource quantity corresponding to other channels.

[0080] In one embodiment, the second determining module 204 further includes: A second acquisition unit is used to acquire an intermediate NUMA channel included in each of the first nodes when the target NUMA channel does not exist, where the intermediate NUMA channel is a NUMA channel with the most GPU resources represented by the second number of GPUs; A computing unit, configured to calculate the number of scheduling computing resources corresponding to each first node based on the number of second GPUs and the number of remaining other resources corresponding to the middle NUMA channel, as well as the required number of GPUs and the required number of other resources; A setting unit is used to set the first node with the smallest amount of scheduling computing resources as the target node.

[0081] In one embodiment, the computing unit comprises: A first calculation subunit, configured to calculate a first difference between the required number of GPUs and the second number of GPUs; A second calculation subunit, used for calculating a second difference between the required quantity of other resources and the remaining quantity of other resources; The third calculation subunit is used to weight the first difference and the second difference to obtain the number of scheduling computing resources corresponding to each first node.

[0082] The intelligent computing power scheduling device for a universal computing power intelligent computing center provided by the present invention can realize the various processes of the various embodiments of the above-mentioned intelligent computing power scheduling method for a universal computing power intelligent computing center. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0083] It should be noted that the computing power intelligent scheduling device for the universal computing power intelligent computing center in the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0084] The present invention also provides an electronic device, see Figure 3 , Figure 3 is a schematic diagram of the structure of an electronic device provided by the present invention, the electronic device includes a memory 301, a processor 302 and a program or instruction stored in the memory 301 and running thereon, and when the program or instruction is executed by the processor 302, it can be realized Figure 1 Any steps in the corresponding embodiments of the method for intelligent scheduling of computing power for an intelligent computing center with universal computing power and the same beneficial effects achieved will not be repeated here.

[0085] The processor 302 may be a CPU, an ASIC, an FPGA or a GPU.

[0086] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiment of the intelligent scheduling method for computing power of an intelligent computing center for universal computing power can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0087] The present invention also provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above Figure 1 The corresponding steps in the embodiment of the intelligent scheduling method for computing power of the intelligent computing center for universal computing power can achieve the same technical effect. To avoid repetition, they are not repeated here. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0088] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The corresponding processes of the implementation method of the intelligent scheduling method for the universal computing power intelligent computing center can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0089] The terms "first", "second" etc. in the present invention are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment. In addition, "and / or" is used in the present application to represent at least one of the connected objects, such as A and / or B and / or C, which means to include 7 situations including single A, single B, single C, and A and B all exist, B and C all exist, A and C all exist, and A, B and C all exist.

[0090] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0091] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a second terminal device, etc.) to execute the methods of each embodiment of the present application.

[0092] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for intelligent computing power scheduling for universal computing power intelligent computing center, characterized in that: include: Step S1: receiving a computing power operation task and the required resource quantity corresponding to the computing power operation task; Step S2: determining a plurality of first nodes based on the required number of resources, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required number of resources; Step S3, obtaining the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node; Step S4, determining a target node from the plurality of first nodes based on the amount of remaining resources; Step S5: dispatch the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

2. The method according to claim 1, characterized in that The required resource quantity includes the required number of graphics processors GPU, and step S2 includes: Step S21, traversing multiple second nodes of the intelligent computing center to obtain the number of first GPUs corresponding to the multiple second nodes; Step S22: Determine the multiple first nodes from the multiple second nodes, and the number of first GPUs corresponding to each first node matches the required number of GPUs.

3. The method according to claim 2, characterized in that The required resource quantity also includes the required other resource quantity, and the step S3 includes: Step S31, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channels included in each first node; The step S4 comprises: Step S41: if there is a target NUMA channel, set the first node where the target NUMA channel is located as the target node; The number of second GPUs corresponding to the target NUMA channel matches the required number of GPUs, and the number of remaining other resources corresponding to the target NUMA channel matches the required number of other resources.

4. The method according to claim 3, characterized in that The step S31 comprises: Step S311: Obtain a fragmentation degree parameter corresponding to each first node, where the fragmentation degree parameter is used to characterize whether different GPUs in the first node are in use; Step S312: sorting the plurality of first nodes based on the corresponding fragmentation degree parameters; Step S313, obtaining the number of second GPUs and the number of remaining other resources corresponding to the NUMA channel included in each of the sorted plurality of first nodes; The step S41 comprises: Step S411: After obtaining the second GPU quantity and the remaining other resource quantity corresponding to the target NUMA channel, set the first node where the target NUMA channel is located as the target node, and stop obtaining the second GPU quantity and the remaining other resource quantity corresponding to other channels.

5. The method according to claim 3, characterized in that The step S4 further comprises: Step S42: if the target NUMA channel does not exist, obtain an intermediate NUMA channel included in each first node, where the intermediate NUMA channel is a NUMA channel with the most GPU resources represented by the second number of GPUs; Step S43, calculating the number of scheduling computing resources corresponding to each first node based on the number of second GPUs and the number of remaining other resources corresponding to the middle NUMA channel, as well as the required number of GPUs and the required number of other resources; Step S44: Set the first node with the smallest amount of scheduling computing resources as the target node.

6. The method according to claim 5, characterized in that The step S43 comprises: Step S431, calculating a first difference between the required number of GPUs and the second number of GPUs; Step S432, calculating a second difference between the required quantity of other resources and the remaining quantity of other resources; Step S433: weight the first difference and the second difference to obtain the number of scheduling computing resources corresponding to each first node.

7. A computing power intelligent scheduling device for a universal computing power intelligent computing center, characterized in that: include: A receiving module, used to receive a computing power operation task and the quantity of required resources corresponding to the computing power operation task; A first determination module is used to determine a plurality of first nodes based on the required resource quantity, wherein the plurality of first nodes are nodes in the intelligent computing center that meet the required resource quantity; An acquisition module, used for acquiring the number of remaining resources corresponding to the non-uniform memory access NUMA channel included in each first node; A second determination module, configured to determine a target node from the plurality of first nodes based on the remaining resource quantity; The scheduling module is used to schedule the computing power operation task to the target node, and the target node is used to execute the computing power operation task.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for intelligent scheduling of computing power for an intelligent computing center for universal computing power are implemented as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for intelligent scheduling of computing power for an intelligent computing center for universal computing power as described in any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for intelligent scheduling of computing power for a universal computing power intelligent computing center as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Cluster resource scheduling method and device, computer equipment and storage medium

    CN113535332A

  • Task processing method and device, electronic equipment and storage medium

    CN116828350A

  • Resource allocation method based on computing power network and related equipment

    CN117834560A

  • Computing power directional scheduling method and device of intelligent computing center

    CN119088569A