Computing Power Intelligent Scheduling Method and Device for Inclusive Computing Power Intelligent Computing Center

By obtaining perceptual data in the intelligent computing center and using the computing model to dynamically schedule computing resources, the problem of low computing resource utilization is solved, the resource utilization is improved and the cost is reduced, and the wide application of universal computing power is achieved.

CN119938340BActive Publication Date: 2025-07-18DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510389932.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The low utilization rate of computing resources in intelligent computing centers leads to shortage during peak hours and idle during low peak hours, increasing economic costs and making it difficult to achieve widespread application of inclusive computing power.

Method used

By obtaining perceptual data, the pre-trained computing model is used to determine the computing resource ratio of each resource pool within the target time period, and dynamically schedule it, including the computing resource allocation of the inference resource pool, the training resource pool and the elastic resource pool, to adapt to the variability and timeliness of computing resource requirements.

Benefits of technology

It improves the utilization rate of computing resources, reduces idleness or overload caused by static allocation of inference tasks and training tasks, reduces model development costs, and realizes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938340B_ABST
    Figure CN119938340B_ABST
Patent Text Reader

Abstract

The present invention provides a computing power intelligent scheduling method and device for an inclusive computing power intelligent computing center, which relates to the technical field of computing power infrastructure. The method includes step S1: obtaining first perception data, which is used to calculate the load conditions of various computing resources in the computing power resource pool during a target time period. The computing power resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool; step S2: inputting the first perception data into a pre-trained computing model to determine a first target ratio during the target time period; step S3: scheduling the computing resources corresponding to each resource pool based on the first target ratio. In this way, the idleness or overload of inference tasks and training tasks caused by static allocation of computing resources is reduced, and the utilization rate of the computing resources in the intelligent computing center is improved; the redundant computing resources reserved for coping with peaks are reduced, and the cost of model development is greatly reduced by dynamically adjusting the allocation of hardware resources, realizing the wide application of inclusive computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a computing power intelligent scheduling method and device for an intelligent computing center for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result through processing information data, and a new type of productive force that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] At present, the computing resources of the intelligent computing center are usually used in two main scenarios, namely, model training scenario and model inference scenario. Based on these two scenarios, computing resources can be divided into independent training clusters and inference clusters. Among them, the training cluster is used for model development and optimization, while the inference cluster is used to process online service requests. However, in actual business, the load of the inference cluster shows an obvious periodic change pattern, namely the "computing power tide" phenomenon. For example, during peak hours (such as 8:00-24:00), the number of online service requests is large, resulting in a high load on the inference cluster, which is prone to slow service request response; during off-peak hours (such as 24:00-6:00), the number of online service requests is small, so the load on the inference cluster is low, and more computing resources are idle. At the same time, training tasks may be queued due to insufficient computing resources. The existence of the "computing power tide" phenomenon causes the inference cluster to be in short supply during peak hours and idle and wasted during off-peak hours. Whether it is increasing the deployment of inference clusters or idling inference clusters, it has invisibly increased the economic cost significantly, making it difficult to achieve the widespread application of inclusive computing power.

[0008] It can be seen that there is a problem of low utilization of computing resources in the existing technology. Summary of the invention

[0009] The embodiments of the present invention provide a computing power intelligent scheduling method and device for a universal computing power intelligent computing center to solve the problem of low utilization rate of computing power resources in the prior art.

[0010] To solve the above problems, the present invention is achieved as follows:

[0011] In a first aspect, an embodiment of the present invention provides a computing power intelligent scheduling method for a universal computing power intelligent computing center, including:

[0012] Step S1, obtaining first perception data, where the first perception data is used to calculate the load of each computing resource in a computing power resource pool within a target time period, where the computing power resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform reasoning tasks and training tasks;

[0013] Step S2: input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool;

[0014] Step S3: Schedule the computing resources corresponding to each resource pool based on the first target ratio.

[0015] In one embodiment, after the step S3, the method further includes:

[0016] Step S4: Count the number of first actual values outside the first effective interval, and determine a first result. The first actual value is the actual value obtained by monitoring a preset metric during a first sub-time period, the first effective interval is the effective value interval of the preset metric calculated according to the ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period;

[0017] Step S5: In the case where the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, input the obtained second perception data into the calculation model, and determine a second target ratio for the sub-time period after the first sub-time period in the target time period. The second target ratio includes at least one ratio between the computing resources corresponding to each resource pool in the computing power resource pool. The second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of the computing resources in each resource pool within the sub-time period after the first sub-time period;

[0018] Step S6: Schedule the computing resources corresponding to each resource pool based on the second target ratio.

[0019] In one embodiment, before the step S4, the method further includes:

[0020] Step S7: Obtain an actual value of monitoring the preset metric at each time interval, and obtain a plurality of first actual values. The first sub-time period includes a plurality of the time intervals, and each time interval corresponds to one of the first actual values;

[0021] Wherein, when there are at least N consecutive time intervals in the first sub-time period whose corresponding first actual values are outside the first effective interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or,

[0022] Wherein, when there are at least M first actual values in the first sub-time period outside the first effective interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

[0023] In one embodiment, after the step S6, the method further includes:

[0024] Step S8: Count the number of second actual values outside the second effective range, and determine a second result. The second actual value is the actual value obtained by monitoring the preset metric in the second sub - time period. The second effective range is the effective value range of the preset metric calculated according to the ratio corresponding to the second sub - time period in the second target ratio. The second sub - time period is the sub - time period after the first sub - time period;

[0025] Step S9: Retrain the calculation model when the second result indicates that the error of the ratio corresponding to the second sub - time period is greater than the preset value.

[0026] In one embodiment, step S3 includes:

[0027] Step S31: When the ratio of the sum of the computing resources in the inference resource pool and the elastic resource pool to the computing resources in the training resource pool in the computing resource pool is less than the proportion of the computing resources in the inference resource pool in the first target ratio, based on the priority and / or progress of the training task, determine the first computing resource in the training resource pool. The priority of the training task of the first computing resource is less than the priority of the training tasks of the other computing resources in the training resource pool except the first computing resource, and / or the progress of the training task of the first computing resource is greater than the preset progress;

[0028] Step S32: After the training task of the first computing resource is paused or ended, determine the first computing resource as the computing resource in the inference resource pool to jointly execute the inference task. The proportion of the computing resources jointly executing the inference task in the computing resource pool is equal to the proportion of the computing resources in the inference resource pool in the first target ratio.

[0029] In one embodiment, step S3 includes:

[0030] Step S33: When the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the inference resource pool in the computing resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, determine the second computing resource in the inference resource pool based on the load condition of the computing resources. The load rate of the second computing resource is less than the load rates of the other computing resources in the inference resource pool;

[0031] Step S34: After the inference task of the second computing resource is ended, determine the second computing resource as the computing resource in the training resource pool to jointly execute the training task. The proportion of the computing resources jointly executing the training task in the computing resource pool is equal to the proportion of the computing resources in the training resource pool in the first target ratio.

[0032] In a second aspect, an embodiment of the present invention further provides a computing power intelligent scheduling device for a general computing power intelligent computing center, including:

[0033] A first acquisition module, configured to acquire first perception data, where the first perception data is used to calculate the load conditions of the computing resources in the computing power resource pool during a target time period. The computing power resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool. Calculation resources are preset in the inference resource pool, the training resource pool, and the elastic resource pool respectively. The calculation resources in the inference resource pool are used to perform inference tasks, the calculation resources in the training resource pool are used to perform training tasks, and the calculation resources in the elastic resource pool can perform both inference tasks and training tasks;

[0034] A first determination module, configured to input the first perception data into a pre-trained calculation model to determine a first target ratio during the target time period. The first target ratio includes at least one ratio among the calculation resources corresponding to the respective resource pools in the computing power resource pool;

[0035] A first scheduling module, configured to schedule the calculation resources corresponding to the respective resource pools based on the first target ratio.

[0036] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the computing power intelligent scheduling method for the general computing power intelligent computing center described in the first aspect above are implemented.

[0037] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the computing power intelligent scheduling method for the general computing power intelligent computing center described in the first aspect above are implemented.

[0038] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the computing power intelligent scheduling method for the general computing power intelligent computing center described in the first aspect above are implemented.

[0039] In an embodiment of the present invention, the current load conditions of each computing resource in the computing resource pool are sensed by obtaining first sensing data; then the first sensing data is input into a pre-trained computing model to determine a first target ratio in a target time period, so as to achieve intelligent computing; finally, the computing resources corresponding to each resource pool are scheduled based on the first target ratio, so as to achieve elastic scheduling. In this way, by combining data sensing, intelligent computing and elastic scheduling, it is possible to adapt to the characteristics of variable and time-sensitive computing resource requirements in the intelligent computing center, reduce the idle or overload of inference tasks and training tasks caused by static allocation of computing resources, and improve the utilization rate of the computing resources in the intelligent computing center; and the computing resources in the elastic resource pool can quickly respond to sudden demands, reduce the redundant computing resources reserved in the inference resource pool and the training resource pool to cope with peaks, and greatly reduce the model development cost through dynamic adjustment of hardware resource allocation, which is conducive to the wide application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0041] Figure 1 is one of the flowcharts of a computing power intelligent scheduling method for an inclusive computing power intelligent computing center provided by an embodiment of the present invention;

[0042] Figure 2 is the second flowchart of a computing power intelligent scheduling method for an inclusive computing power intelligent computing center provided by an embodiment of the present invention;

[0043] Figure 3 is the statistical chart corresponding to the first target ratio in the target time period provided by an embodiment of the present invention;

[0044] Figure 4 is the structural diagram of a computing power intelligent scheduling device for an inclusive computing power intelligent computing center provided by an embodiment of the present invention;

[0045] Figure 5 is the structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and achieve the output of the target result, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0048] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .

[0049] The "carrying capacity" (Network Power, NP) referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive index for measuring network transmission scheduling ability.

[0050] The "storage power" (Storage Power, SP) referred to in the present invention means: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive index for measuring the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0051] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.

[0052] The "new type of information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology infrastructures such as artificial intelligence, blockchain, and quantum computing.

[0053] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0054] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0055] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.

[0056] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0057] The "intelligent computing center" described in the present invention refers to: a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0058] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0059] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0060] The "Computing Power Center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0061] The "Supercomputing Center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0062] The "Computing Power Resources" described in the present invention refer to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0063] The "Inclusive Computing Power" described in the present invention refers to providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost based on the requirements of equal opportunity and the principle of commercial sustainability.

[0064] The "Model" described in the present invention includes but is not limited to "Large Language Model" and "Multimodal Large Model".

[0065] The "Large Language Model" described in the present invention refers to a large-scale language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0066] The "Multimodal Large Model" described in the present invention (Multimodal Large Models) refers to: a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0067] The "Agent" described in the present invention refers to an entity that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, and has autonomy, adaptability, and interaction capabilities. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal. Agents are widely used in the field of artificial intelligence and are commonly found in automated systems, robots, virtual assistants, and game characters, etc. The core lies in the ability to learn autonomously and evolve continuously to better complete tasks and adapt to complex environments.

[0068] Please refer to Figure 1 , Figure 1 FIG. is one of the flowcharts of a computing power intelligent scheduling method for a general computing power intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:

[0069] Step S1: Obtain first sensing data, which is used to calculate the load conditions of each computing resource in the computing power resource pool during a target time period. The computing power resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool. There are preset computing resources in the inference resource pool, the training resource pool, and the elastic resource pool respectively. The computing resources in the inference resource pool are used for inference tasks, the computing resources in the training resource pool are used for training tasks, and the computing resources in the elastic resource pool can perform both inference tasks and training tasks;

[0070] In this step, the first sensing data may include a first index and a second index. The first index may include the utilization rate of the Graphics Processing Unit (GPU) collected in real time, the task queue depth, the Queries-per-second (QPS) of inference requests, etc.; the second index may include environmental factor data, time factor data, event factor data, and other indicators collected. Both the first index and the second index are factors affecting the load of each computing resource in the computing power resource pool. For example, a high GPU utilization rate, a large task queue, and a high QPS can directly reflect a relatively high load of the corresponding computing resource; while environmental factor data, time factor data, and event factor data indirectly affect the load of computing resources. For example, when there are more users in a certain area, during working hours, and when a hot event occurs, it can be predicted that the load of the computing resources corresponding to these environments, times, and events will be relatively high. Therefore, through the first sensing data, the load conditions of each computing resource in the computing power resource pool during the target time period can be calculated.

[0071] Among them, the target time period can be a period of time after the time point when the first perception data is obtained. For example, the first perception data is obtained at 8:00 to calculate the load conditions of the computing resources in the computing power resource pool within the target time period, so as to achieve reasonable allocation of computing resources within the target time period and improve resource utilization. Exemplarily, the target time period can be the four hours from 9:00 to 13:00. Of course, the target time period can also be other time periods, and can also be longer or shorter time, which is not limited here.

[0072] Among them, in this embodiment, the computing power resource pool can be divided into three categories, namely the inference resource pool, the training resource pool, and the elastic resource pool. Each type of resource pool corresponds to computing resources. The sum of the computing resources in the inference resource pool, the training resource pool, and the elastic resource pool is the number of computing resources in the computing power resource pool. The computing resources in the inference resource pool are used for inference tasks, the computing resources in the training resource pool are used for training tasks, and the computing resources in the elastic resource pool can perform inference tasks and training tasks. When the priority of the inference task is higher than that of the training task, the computing resources in the elastic resource pool can be called to perform the inference task; in some cases, the computing resources in the training resource pool can also be called to perform the inference task. When the priority of the training task is higher than that of the inference task, the computing resources in the elastic resource pool can be called to perform the training task; in some cases, the computing resources in the inference resource pool can also be called to perform the training task. In this way, according to the load conditions of the computing resources in the computing power resource pool, the idle computing resources can be flexibly scheduled to adapt to the periodic changes in the load corresponding to the inference task (i.e., the "computing power tide" phenomenon), and during the low peak period of the inference task, reduce the situation where the training task queues up due to insufficient computing resources in the training resource pool.

[0073] Step S2: Input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, where the first target ratio includes at least one ratio between the computing resources corresponding to each resource pool in the computing power resource pool;

[0074] In this step, the computing model can be a multi-classification model pre-trained with sample perception data as features. The sample perception data can include the load information of each computing resource recorded in history at different times, as well as whether it is a holiday, whether there is a special event, and indicators such as GPU utilization rate, task queue depth, and QPS in the computing resources. Exemplarily, the data can be divided into a training set (historical data) and a test set (recent data) in chronological order to avoid data leakage caused by chronological order.

[0075] Among them, the output of the computing model can be the ratio between the computing resources corresponding to each resource pool in the computing power resource pool, such as [inference:training:flexible]. The target time period can be a relatively long time period. Therefore, there can be multiple different ratios in the target time period. In this way, by inputting the first perception data into the pre-trained computing model to determine the first target ratio in the target time period, the long-term trend prediction of the computing resource allocation corresponding to each resource pool in the computing power resource pool is realized.

[0076] Exemplarily, the target time period can be from 9:00 to 13:00. Inputting the first perception data into the pre-trained computing model to determine the first target ratio from 9:00 to 13:00 can include [5:4:1], [6:4:0], and [7:3:0]. Among them, [5:4:1] can be the quantity ratio between the computing resources corresponding to the inference resource pool, the training resource pool, and the flexible resource pool respectively from 9:00 to 10:00; [6:4:0] can be the quantity ratio between the computing resources corresponding to the inference resource pool, the training resource pool, and the flexible resource pool respectively from 10:00 to 12:00; [7:3:0] can be the quantity ratio between the computing resources corresponding to the inference resource pool, the training resource pool, and the flexible resource pool respectively from 12:00 to 13:00. It can be seen from the calculated first target ratio that:

[0077] According to the first perception data, it is predicted that there are not many inference tasks from 9:00 to 10:00. At this time, the computing resource allocation in the inference resource pool is 5, which can handle the inference tasks from 9:00 to 10:00. At the same time, the computing resource allocation in the training resource pool is 4 for normal training tasks, and the computing resource allocation in the flexible resource pool is 1 to handle emergencies;

[0078] According to the first perception data, it is predicted that the inference tasks increase from 10:00 to 12:00. At this time, the computing resource allocation in the inference resource pool is set to 6, which can relieve the load pressure caused by the increased inference tasks from 10:00 to 12:00; at the same time, the computing resource allocation in the training resource pool remains 4 for normal training tasks; and the computing resources in the flexible resource pool will be scheduled to the inference resource pool, so the computing resource allocation in the flexible resource pool is 0.

[0079] According to the first perception data, it is predicted that the inference tasks will further increase from 12:00 to 13:00. At this time, the computing resource ratio in the inference resource pool is set to 7, which can relieve the load pressure caused by the further increased inference tasks from 12:00 to 13:00. While the training tasks are being carried out, 1 unit of computing resources in the training resource pool can be scheduled to the inference resource pool to meet the requirements of the inference tasks. Therefore, the computing resource ratio in the training resource pool is reduced to 3; and the computing resources in the elastic resource pool will be scheduled to the inference resource pool. Therefore, the computing resource ratio in the elastic resource pool is 0.

[0080] Among them, the first target ratio can also be other situations. For example, the computing resources in the inference resource pool and / or the elastic resource pool can be scheduled to the training resource pool. During the low peak period of the inference tasks, the situation that the training tasks queue up waiting due to insufficient computing resources in the training resource pool is reduced.

[0081] In this way, the current load is monitored in real time through the first perception data, and the demands of each resource pool in the target time period are predicted, so as to determine the appropriate first target ratio, and then the corresponding ratio is scheduled. By combining data perception, intelligent computing and elastic scheduling, it can adapt to the characteristics of the variable and time-sensitive computing resource requirements of the intelligent computing center, reduce the idle or overload caused by the static allocation of computing resources for the inference tasks and training tasks, and improve the utilization rate of the computing resources of the intelligent computing center; and an elastic resource pool is set up, and the computing resources in the elastic resource pool can quickly respond to sudden demands (such as temporary tasks), reduce the redundant computing resources reserved in the inference resource pool and the training resource pool to cope with the peak value, and greatly reduce the cost of model development through dynamic adjustment of the hardware resource allocation, realizing the wide application of inclusive computing power.

[0082] Step S3: Schedule the computing resources corresponding to each resource pool based on the first target ratio.

[0083] In this step, after determining the first target ratio, the computing resources corresponding to each resource pool are scheduled based on the first target ratio, realizing the dynamic adjustment of the hardware resource allocation, improving the task throughput, and thus improving the utilization rate of the computing resources of the intelligent computing center.

[0084] Specifically, as Figure 2 shown, Figure 2It is the second flowchart of a computing power intelligent scheduling method for an inclusive computing power intelligent computing center provided by an embodiment of the present invention. This method can be applied to an intelligent agent. A computing power resource pool is deployed in the intelligent agent, and the computing power resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool. In addition, a tide perception layer, an intelligent computing layer, and an elastic scheduling layer are also set in the intelligent agent. The tide perception layer is used to obtain first perception data, which includes a first indicator and a second indicator. The first indicator may include the GPU utilization rate collected in real time, the task queue depth, the QPS of inference requests, etc. The second indicator may include indicator data such as environmental factor data, time factor data, and event factor data collected. Data perception is achieved through the tide perception layer. The intelligent computing layer is used to determine the first target ratio in a target time period according to the first perception data, so as to achieve intelligent computing. The elastic scheduling layer is used to schedule the computing resources corresponding to each resource pool based on the first target ratio, so as to achieve elastic scheduling of the computing resource ratio. Among them, the computing power resource pool can schedule the computing resources corresponding to each resource pool based on the scheduling instruction sent by the elastic scheduling layer, and the scheduling instruction includes the first target ratio information. By dynamically adjusting the hardware resource allocation, the idle or overload caused by the static allocation of computing resources for inference tasks and training tasks is reduced, the utilization rate of the computing resources in the intelligent computing center is improved, the cost of model development is greatly reduced, and it is beneficial to realize the wide application of inclusive computing power.

[0085] In an embodiment of the present invention, the current load conditions of each computing resource in the computing power resource pool are perceived by obtaining the first perception data; then the first perception data is input into a pre-trained computing model to determine the first target ratio in the target time period, and intelligent computing is achieved; finally, the computing resources corresponding to each resource pool are scheduled based on the first target ratio to achieve elastic scheduling. In this way, by combining data perception, intelligent computing, and elastic scheduling, it can adapt to the characteristics of the changing and time-sensitive computing resource requirements of the intelligent computing center, reduce the idle or overload caused by the static allocation of computing resources for inference tasks and training tasks, and improve the utilization rate of the computing resources in the intelligent computing center; and the computing resources in the elastic resource pool can quickly respond to sudden demands, reduce the redundant computing resources reserved in the inference resource pool and the training resource pool to cope with peaks, and greatly reduce the cost of model development by dynamically adjusting the hardware resource allocation, realizing the wide application of inclusive computing power.

[0086] In one embodiment, after the step S3, the method further includes:

[0087] Step S4: Count the number of first actual values outside the first valid range, and determine a first result. The first actual value is the actual value obtained by monitoring a preset metric during a first sub-time period. The first valid range is the valid value range of the preset metric calculated according to the ratio corresponding to the first sub-time period in the first target ratio. The first sub-time period is any sub-time period within the target time period.

[0088] Step S5: When the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, input the obtained second perception data into the calculation model to determine a second target ratio for the sub-time period after the first sub-time period in the target time period. The second target ratio includes at least one ratio between the computing resources corresponding to each resource pool in the computing power resource pool. The second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of the computing resources in each resource pool within the sub-time period after the first sub-time period.

[0089] Step S6: Schedule the computing resources corresponding to each resource pool based on the second target ratio.

[0090] In this embodiment, during the process of scheduling the computing resources corresponding to each resource pool according to the first target ratio in the target time period, considering that if there is a large error in the calculated first target ratio, there will still be situations where some computing resources in the computing power resource pool are idle or overloaded. To reduce this calculation error, the calculation ratio can be corrected according to the short-term trend within the target time period. Specifically, refer to the following description:

[0091] Exemplarily, as Figure 3 shown, the target time period can include multiple sub-time periods. Calculate the first target ratio among the computing resources in the inference resource pool, the computing resources in the training resource pool, and the computing resources in the elastic resource pool to be executed within the target time period (taking 9:00 to 13:00 as an example): the ratio corresponding to 9:00 to 10:00 is [5:4:1]; the ratio corresponding to 10:00 to 12:00 is [6:4:0]; the ratio corresponding to 12:00 to 13:00 is [7:3:0].

[0092] An actual value of monitoring the preset metric can be obtained every preset time interval, for example, every 15 minutes. Among them, in a scenario dominated by inference tasks, the preset metric can include the QPS of inference requests; in a scenario dominated by training tasks, the preset metric can include the queue depth of training tasks. The following takes the scenario dominated by inference tasks as an example for exemplary illustration. It should be understood that it is also applicable in a scenario dominated by training tasks.

[0093] After step S3, the scheduling of the computing resources corresponding to each resource pool is completed based on the first target ratio. Therefore, the ratio corresponding to 9:00 to 10:00 is [5:4:1]. Assuming that the first sub-time period is the sub-time period from 9:00 to 10:00 within 9:00 to 13:00, then the QPS of the inference requests collected every 15 minutes starting from 9:00 can obtain 4 actual values corresponding to the first sub-time period, that is, 4 first actual values. Among them, according to the ratio corresponding to 9:00 to 10:00 being [5:4:1], the theoretical upper limit value and the theoretical lower limit value of the QPS of the inference requests under this ratio can be calculated, so as to determine the effective value range of the QPS of the inference requests. At this time, the first effective range is the effective value range of the QPS of the inference requests when the computing resource ratios are [5:4:1] during 9:00 to 10:00. In this way, the number of the first actual values located outside the first effective range can be counted to determine the first result.

[0094] Among them, in this example, the sub-time period from 9:00 to 10:00 within 9:00 to 13:00 is used as the first sub-time period for illustration. It should be understood that the first sub-time period can be any sub-time period within the target time period. In this example, if the first result indicates that the error of the ratio corresponding to the first sub-time period, that is, [5:4:1], is less than or equal to the preset value, it means that the first target ratio calculated by the computing model based on the first perception data is relatively accurate. In other sub-time periods after the first sub-time period (that is, 10:00 to 13:00), the scheduling can still be executed according to the ratios corresponding to each sub-time period, thereby improving the utilization rate of the computing resources of the intelligent computing center. If the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, it means that the first target ratio calculated by the computing model based on the first perception data is relatively inaccurate. When the scheduling is still executed according to the ratios corresponding to each sub-time period in other sub-time periods after the first sub-time period (that is, 10:00 to 13:00), errors will still occur. Therefore, it is necessary to correct the calculated ratio: input the obtained second perception data into the computing model to determine the second target ratio of the sub-time period after the first sub-time period in the target time period. The second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of the computing resources in each computing resource pool within the sub-time period after the first sub-time period. Among them, the sub-time period after the first sub-time period can be a sub-time period within the target time period, such as 10:00 to 13:00; it can also be a sub-time period outside the target time period, such as 10:00 to 14:00. In this way, the calculated ratio is corrected through the short-term trend within the target time period to obtain the second target ratio again; then the computing resources corresponding to each resource pool are scheduled based on the second target ratio. The accuracy of calculating the computing resource ratios in each resource pool is improved.

[0095] In one embodiment, before the step S4, the method further includes:

[0096] Step S7: Obtain an actual value for monitoring the preset metric at each time interval, obtaining a plurality of first actual values. The first sub - time period includes a plurality of such time intervals, and each time interval corresponds to one of the first actual values;

[0097] Wherein, when there are at least N consecutive time intervals in the first sub - time period whose corresponding first actual values are outside the first effective interval, the first result indicates that the error of the proportion corresponding to the first sub - time period is greater than the preset value, and N is a positive integer; or,

[0098] Wherein, when there are at least M first actual values in the first sub - time period that are outside the first effective interval, the first result indicates that the error of the proportion corresponding to the first sub - time period is greater than the preset value, and M is a positive integer.

[0099] In one example, it is determined that the first target proportion calculated based on the first perception data by the calculation model is relatively inaccurate. In other words, the first result indicates that the error of the proportion corresponding to the first sub - time period is greater than the preset value, which can be determined based on the fact that there are at least N consecutive time intervals in the first sub - time period whose corresponding first actual values are outside the first effective interval. That is, it can be judged based on the number of consecutive error occurrences. For example, an actual value for monitoring the preset metric is obtained every preset time, such as every 15 minutes. If there are at least 3 consecutive time intervals in the period from 9:00 to 10:00 whose corresponding first actual values are outside the first effective interval, then it can be determined that the proportion [5:4:1] does not conform to the actual situation and needs to be recalculated. For the allocation ratio of computing resources in each resource pool at the current time, it can be temporarily scheduled through the conventional ratio to promptly respond to the inference request.

[0100] In another example, it can also be determined whether the first target proportion is accurate based on the fact that there are at least M first actual values in the first sub - time period that are outside the first effective interval. In other words, it can be judged based on the total number of error occurrences. For example, an actual value for monitoring the preset metric is obtained every preset time, such as every 15 minutes. If the total number of error occurrences reaches 3 times in the period from 9:00 to 10:00, that is, there are at least 3 first actual values outside the first effective interval in total, then it can be determined that the proportion [5:4:1] does not conform to the actual situation and needs to be recalculated. For the allocation ratio of computing resources in each resource pool at the current time, it can also be temporarily scheduled through the conventional ratio to promptly respond to the inference request.

[0101] In one embodiment, after the step S6, the method further includes:

[0102] Step S8, counting the number of second actual values outside the second effective interval, and determining a second result, where the second actual value is the actual value obtained by monitoring the preset index in the second sub-time period, the second effective interval is the effective value range of the preset index calculated according to the ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is the sub-time period after the first sub-time period;

[0103] Step S9, retraining the calculation model when the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value.

[0104] In this embodiment, during the process of scheduling the computing resources corresponding to each resource pool according to the first target ratio in the target time period, according to the first result, it can be determined that the error of the ratio corresponding to the first sub-time period is relatively large. Therefore, the sub-time period after the first sub-time period is recalculated to obtain the second target ratio, and the computing resources corresponding to each resource pool are rescheduled based on the second target ratio. In this way, the calculation ratio is corrected according to the short-term trend in the target time period, improving the accuracy of the calculation result. Further, during the process of scheduling the computing resources corresponding to each resource pool according to the second target ratio in the sub-time period after the first sub-time period, considering that if there is a large error in the recalculated second target ratio, there will still be idle or overloaded computing resources in some computing resource pools. To reduce this calculation error, the calculation ratio can be corrected according to the long-term trend of the first sub-time period and other sub-time periods after the first sub-time period. Specifically, the following description can be referred to:

[0105] Exemplarily, assume that the second sub-time period is the sub-time period from 10:00 to 11:00 in the period from 10:00 to 13:00. Then, the QPS of the inference requests collected every 15 minutes starting from 10:00 can be obtained, and 4 actual values corresponding to the second sub-time period can be obtained, that is, 4 second actual values. Among them, according to the ratio corresponding to 10:00 to 12:00 being [6:4:0], the theoretical upper limit value and the theoretical lower limit value of the inference request QPS corresponding to this ratio can be calculated, so as to determine the effective value range of the inference request QPS. At this time, the second effective interval is the effective value range of the inference request QPS when the computing resource ratios are all [6:4:0] during the period from 10:00 to 11:00. In this way, the number of second actual values outside the second effective interval can be counted to determine the second result.

[0106] Among them, in this example, the sub-time period from 9:00 to 11:00 within the time period from 10:00 to 13:00 is used as the second sub-time period for illustration. It should be understood that the second sub-time period can be any sub-time period after the first sub-time period (from 9:00 to 10:00). In this example, if the second result represents the ratio corresponding to the second sub-time period, that is, when the error of [6:4:0] is less than or equal to the preset value, it indicates that the second target ratio calculated by the calculation model based on the second sensing data is relatively accurate. In other sub-time periods after the second sub-time period (i.e., from 11:00 to 13:00), the scheduling can still be performed according to the ratio corresponding to each sub-time period, thereby improving the utilization rate of the computing resources of the intelligent computing center. If the error of the ratio corresponding to the second sub-time period represented by the second result is greater than the preset value, it indicates that the second target ratio calculated by the calculation model based on the second sensing data is relatively inaccurate, and errors will still occur when the scheduling is still performed according to the ratio corresponding to each sub-time period in other sub-time periods after the second sub-time period. Therefore, it is necessary to correct the calculated ratio: Considering that previously, the first target ratio calculated by the calculation model based on the first sensing data was also inaccurate. If the correction at this time is still to update the sensing data input to the calculation model, the subsequent calculation results will still be inaccurate. Therefore, in this case, the calculation model is retrained. Thus, the accuracy of calculating the computing resource ratio in each resource pool is ensured.

[0107] It should be understood that considering the existence of accidental factors, it can be when the subsequent calculation results are still inaccurate after updating the sensing data input to the calculation model multiple times, that the calculation model is retrained.

[0108] In some alternative embodiments, step S3 includes:

[0109] Step S31: When the ratio of the sum of the computing resources in the inference resource pool and the elastic resource pool to the computing resources in the training resource pool in the computing power resource pool is less than the proportion of the computing resources in the inference resource pool in the first target ratio, based on the priority and / or progress of the training task, determine a first computing resource in the training resource pool, where the priority of the training task of the first computing resource is less than the priority of the training tasks of the other computing resources in the training resource pool except the first computing resource, and / or, the progress of the training task of the first computing resource is greater than the preset progress;

[0110] Step S32: After the training task of the first computing resource is paused or ended, determine the first computing resource as the computing resource in the inference resource pool to jointly execute the inference task, and the proportion of the computing resources jointly executing the inference task in the computing power resource pool is equal to the proportion of the computing resources in the inference resource pool in the first target ratio.

[0111] In this embodiment, when the ratio of the total computing resources of the inference resource pool plus the elastic resource pool to the computing resources of the training resource pool is less than the proportion of the computing resources in the inference resource pool in the first target ratio, it indicates that the current inference resources are insufficient and resources need to be allocated from the training resource pool. For example, assume that the first target ratio requires the inference resources to account for 60%, and the current inference + elasticity only accounts for 40%, then 20% of the resources need to be released from the training resource pool. The 20% of the computing resources (i.e., the first computing resources) that need to be released from the training resource pool can be: select the computing resources with a lower training task priority (for example, the priority of the offline training task is lower than that of the online model fine-tuning task); select the computing resources corresponding to the tasks whose training progress has exceeded the preset threshold (such as 80%), because the recovery cost after suspension is lower. Among them, the task priority, progress, and computing power usage of each computing resource can be collected through the monitoring system, and sorted in ascending order of priority and descending order of progress to determine the candidate list that can be suspended.

[0112] Then, issue a Save Checkpoint instruction to the selected first computing resources to record the training status (such as model parameters, optimizer status); release the first computing resources after stopping the training task, and add the released first computing resources to the inference resource pool for performing inference tasks; determine the first computing resources as the computing resources in the inference resource pool through a service mesh (such as Istio) or a load balancer (such as Nginx) to jointly execute the inference tasks. The proportion of the computing resources jointly executing the inference tasks in the computing power resource pool is equal to the proportion of the computing resources in the inference resource pool in the first target ratio.

[0113] Among them, when the proportion of the inference resource pool reaches the first target ratio, stop resource migration. If there are still remaining released resources, they can be returned to the elastic resource pool for subsequent tasks to dynamically apply. If the inference load drops, resources can be transferred back to the training pool according to the reverse logic. In this way, without adding hardware, the resource reuse of training and inference is realized. The utilization rate of the computing resources of the intelligent computing center is improved.

[0114] In some alternative embodiments, step S3 includes:

[0115] Step S33: In the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the inference resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, determine second computing resources in the inference resource pool based on the load conditions of the computing resources, and the load rate of the second computing resources is less than the load rates of other computing resources in the inference resource pool;

[0116] Step S34: After the inference task of the second computing resource ends, determine the second computing resource as the computing resource in the training resource pool to jointly execute the training task. The proportion of the computing resources jointly executing the training task in the computing power resource pool is equal to the proportion of the computing resources in the training resource pool in the first target ratio.

[0117] In this embodiment, when the ratio of the total computing resources of the training resource pool plus the elastic resource pool to the computing resources of the inference resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, it indicates that the current training resources are insufficient and resources need to be allocated from the inference resource pool. For example: Suppose the first target ratio requires that the training resources account for 40%, and the current training + elasticity only accounts for 30%, then 10% of the resources need to be released from the inference resource pool. The 10% of the computing resources (i.e., the second computing resource) that need to be released from the inference resource pool can be determined as follows: Select the computing resources with a lower inference task load rate (such as the CPU / GPU utilization rate is lower than the threshold). Among them, the load rate calculation can combine indicators such as the task queue depth and response latency; preferentially release general-purpose inference resources (such as non-real-time inference services) to avoid affecting key services. Among them, tools such as Prometheus can be used to collect indicators such as the load rate and task queue length of each inference node in real time, and sort them in ascending order of the load rate, and select the cluster group with the lowest load as the release target.

[0118] Then, issue an instruction to the selected second computing resource to stop receiving new requests, and wait for the current request to be processed; after closing the inference service instance, release the second computing resource, and add the released second computing resource to the training resource pool for training tasks; determine the second computing resource as the computing resource in the training resource pool to jointly execute the training task. The proportion of the computing resources jointly executing the training task in the computing power resource pool is equal to the proportion of the computing resources in the training resource pool in the first target ratio.

[0119] Among them, when the proportion of the training resource pool reaches the target ratio, stop resource migration. If there are still remaining released resources, they can be retained in the elastic resource pool for subsequent dynamic allocation. If the training load decreases, resources can be transferred back to the inference pool according to the reverse logic. In this way, without adding hardware, the resource reuse of training and inference is realized. The utilization rate of the computing resources of the intelligent computing center is improved.

[0120] Please refer to Figure 4 , Figure 4 which is the structural diagram of a computing power intelligent scheduling device for an intelligent computing center for inclusive computing power provided by an embodiment of the present invention. As Figure 4 shown, the computing power intelligent scheduling device 400 for an intelligent computing center for inclusive computing power includes:

[0121] The first acquisition module 401 is configured to acquire first perception data, where the first perception data is used to calculate the load conditions of the computing resources in the computing resource pool during the target time period. The computing resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool. Computing resources are preset in the inference resource pool, the training resource pool, and the elastic resource pool respectively. The computing resources in the inference resource pool are used to perform inference tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform both inference tasks and training tasks;

[0122] The first determination module 402 is configured to input the first perception data into a pre-trained computing model to determine a first target ratio during the target time period. The first target ratio includes at least one ratio among the computing resources corresponding to the respective resource pools in the computing resource pool;

[0123] The first scheduling module 403 is configured to schedule the computing resources corresponding to the respective resource pools based on the first target ratio.

[0124] In one embodiment, the apparatus further includes:

[0125] The first statistics module is configured to count the number of first actual values outside the first effective interval to determine a first result. The first actual value is the actual value obtained by monitoring a preset index during a first sub-time period. The first effective interval is the effective value interval of the preset index calculated according to the ratio corresponding to the first sub-time period in the first target ratio. The first sub-time period is any sub-time period within the target time period;

[0126] The second determination module is configured to, when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, input the acquired second perception data into the computing model to determine a second target ratio for the sub-time period after the first sub-time period during the target time period. The second target ratio includes at least one ratio among the computing resources corresponding to the respective resource pools in the computing resource pool. The second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of the computing resources in the computing resource pool during the sub-time period after the first sub-time period;

[0127] The second scheduling module is configured to schedule the computing resources corresponding to the respective resource pools based on the second target ratio.

[0128] In one embodiment, the apparatus further includes:

[0129] A second acquisition module, configured to acquire an actual value of monitoring the preset metric in each time interval, obtaining a plurality of first actual values. The first sub-time period includes a plurality of the time intervals, and each time interval corresponds to one of the first actual values;

[0130] Wherein, when there are at least N consecutive time intervals in the first sub-time period, and the corresponding first actual values are outside the first valid interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or,

[0131] Wherein, when there are at least M of the first actual values in the first sub-time period outside the first valid interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

[0132] In one embodiment, the apparatus further includes:

[0133] A second statistics module, configured to count the number of second actual values outside the second valid interval, determining a second result. The second actual value is the actual value obtained by monitoring the preset metric in the second sub-time period. The second valid interval is the valid value interval of the preset metric calculated according to the ratio corresponding to the second sub-time period in the second target ratio. The second sub-time period is the sub-time period after the first sub-time period;

[0134] A retraining module, configured to retrain the calculation model when the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value.

[0135] In one embodiment, the first scheduling module 403 is specifically configured to:

[0136] In the computing power resource pool, when the ratio of the sum of the computing resources in the inference resource pool and the elastic resource pool to the computing resources in the training resource pool is less than the proportion of the computing resources in the inference resource pool in the first target ratio, based on the priority and / or progress of the training task, determine first computing resources in the training resource pool;

[0137] After the training task of the first computing resources is paused or ended, determine the first computing resources as the computing resources in the inference resource pool to jointly execute the inference task. The proportion of the computing resources jointly executing the inference task in the computing power resource pool is equal to the proportion of the computing resources in the inference resource pool in the first target ratio.

[0138] In one embodiment, the first scheduling module 403 is specifically configured to:

[0139] In the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the inference resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, a second computing resource is determined in the inference resource pool based on the load condition of the computing resources.

[0140] After the inference task of the second computing resource ends, the second computing resource is determined as the computing resource in the training resource pool to jointly execute the training task, and the proportion of the computing resources jointly executing the training task in the computing power resource pool is equal to the proportion of the computing resources in the training resource pool in the first target ratio.

[0141] The computing power intelligent scheduling device for the inclusive computing power intelligent computing center provided by the embodiments of the present invention can implement each process of the above-mentioned computing power intelligent scheduling method for the inclusive computing power intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they are not elaborated here.

[0142] It should be noted that the computing power intelligent scheduling device for the inclusive computing power intelligent computing center in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0143] The embodiments of the present invention also provide an electronic device. Refer to Figure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction running on the memory 501. When the program or instruction is executed by the processor 502, it can implement Figure 1 any step in the corresponding embodiment of the computing power intelligent scheduling method for the inclusive computing power intelligent computing center and achieve the same beneficial effects, which are not elaborated here.

[0144] Among them, the processor 502 can be a CPU, an ASIC, an FPGA, or a GPU.

[0145] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned embodiment of the computing power intelligent scheduling method for the inclusive computing power intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0146] The embodiments of the present invention also provide a readable storage medium. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1Any step in the corresponding embodiment of the intelligent computing power scheduling method for the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. The storage medium, such as Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disc, etc.

[0147] The present invention also provides a computer program product, including computer instructions, which when executed by a processor implement the above Figure 1 Each process of the corresponding embodiment of the intelligent computing power scheduling method for the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0148] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, and A and B both exist, B and C both exist, A and C both exist, and A, B, and C all exist, a total of 7 cases.

[0149] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element.

[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of the present application.

[0151] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A computing power intelligent scheduling method for an inclusive computing power intelligent computing center, characterized in that, Including: Step S1: Obtain first perception data, which is used to calculate the load conditions of various computing resources in the computing resource pool during the target time period. The computing resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool. There are preset computing resources in the inference resource pool, the training resource pool, and the elastic resource pool respectively. The computing resources in the inference resource pool are used for inference tasks, the computing resources in the training resource pool are used for training tasks, and the computing resources in the elastic resource pool can perform both inference tasks and training tasks. Step S2: Input the first perception data into a pre-trained computing model to determine a first target ratio during the target time period. The first target ratio includes at least one ratio among the computing resources corresponding to each resource pool in the computing resource pool. Step S3: Schedule the computing resources corresponding to each resource pool based on the first target ratio. After the step S3, the method further includes: Step S4: Count the number of first actual values outside the first effective interval to determine a first result. The first actual value is the actual value obtained by monitoring a preset index during a first sub-time period. The first effective interval is the effective value interval of the preset index calculated according to the ratio corresponding to the first sub-time period in the first target ratio. The first sub-time period is any sub-time period within the target time period. Step S5: In the case where the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, input the obtained second perception data into the computing model to determine a second target ratio for the sub-time period after the first sub-time period in the target time period. The second target ratio includes at least one ratio among the computing resources corresponding to each resource pool in the computing resource pool. The second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of various computing resources in the computing resource pool during the sub-time period after the first sub-time period. Step S6: Schedule the computing resources corresponding to each resource pool based on the second target ratio; wherein, both the first target ratio and the second target ratio are the ratios of the computing resources in the inference resource pool, the computing resources in the training resource pool, and the computing resources in the elastic resource pool.

2. The method according to claim 1, characterized in that, Before the step S4, the method further includes: Step S7: Obtain an actual value of monitoring the preset index at each time interval to obtain a plurality of first actual values. The first sub-time period includes a plurality of the time intervals, and each time interval corresponds to one of the first actual values. Wherein, when there are at least N consecutive time intervals in the first sub-time period corresponding to the first actual values located outside the first effective interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or, Wherein, when at least M of the first actual values in the first sub-time period are outside the first valid interval, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

3. The method according to claim 1, characterized in that, After the step S6, the method further includes: Step S8, counting the number of second actual values outside the second valid interval to determine a second result, where the second actual value is the actual value obtained by monitoring the preset index in a second sub-time period, the second valid interval is the valid value interval of the preset index calculated according to the ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is the sub-time period after the first sub-time period; Step S9, retraining the calculation model when the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value.

4. The method according to any one of claims 1 to 3, characterized in that, The step S3 includes: Step S31, when the ratio of the sum of the computing resources in the inference resource pool and the elastic resource pool to the computing resources in the training resource pool in the computing resource pool is less than the proportion of the computing resources in the inference resource pool in the first target ratio, determining first computing resources in the training resource pool based on the priority and / or progress of the training task; Step S32, after the training task of the first computing resources is paused or ended, determining the first computing resources as the computing resources in the inference resource pool to jointly execute the inference task, and the proportion of the computing resources jointly executing the inference task in the computing resource pool is equal to the proportion of the computing resources in the inference resource pool in the first target ratio.

5. The method according to any one of claims 1 to 3, characterized in that The step S3 includes: Step S33, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the inference resource pool in the computing resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, determining second computing resources in the inference resource pool based on the load condition of the computing resources; Step S34, after the inference task of the second computing resources ends, determining the second computing resources as the computing resources in the training resource pool to jointly execute the training task, and the proportion of the computing resources jointly executing the training task in the computing resource pool is equal to the proportion of the computing resources in the training resource pool in the first target ratio.

6. A computing power intelligent scheduling device for an inclusive computing power intelligent computing center, characterized in that, Including: A first acquisition module, configured to acquire first perception data, where the first perception data is used to calculate the load conditions of the computing resources in the computing resource pool within a target time period, the computing resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool, computing resources are preset in the inference resource pool, the training resource pool, and the elastic resource pool respectively, the computing resources in the inference resource pool are used to perform inference tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform inference tasks and training tasks; A first determination module, configured to input the first perception data into a pre-trained computing model, and determine a first target ratio in the target time period, where the first target ratio includes at least one ratio between the computing resources corresponding to each resource pool in the computing power resource pool; A first scheduling module, configured to schedule the computing resources corresponding to each resource pool based on the first target ratio; The apparatus further includes: A first statistics module, configured to count the number of first actual values outside the first valid interval, and determine a first result, where the first actual value is an actual value obtained by monitoring a preset index in a first sub-time period, the first valid interval is a valid value interval of the preset index calculated according to the ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period; A second determination module, configured to, when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, input the obtained second perception data into the computing model, and determine a second target ratio in the sub-time period after the first sub-time period in the target time period, where the second target ratio includes at least one ratio between the computing resources corresponding to each resource pool in the computing power resource pool, the second perception data includes the first actual value and the first perception data, and the second perception data is used to calculate the load conditions of the computing resources in each resource pool in the sub-time period after the first sub-time period; A second scheduling module, configured to schedule the computing resources corresponding to each resource pool based on the second target ratio; Wherein, both the first target ratio and the second target ratio are ratios of the computing resources in the inference resource pool, the computing resources in the training resource pool, and the computing resources in the elastic resource pool.

7. An electronic device, characterized in that, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the computing power intelligent scheduling method for a general-purpose computing power intelligent computing center according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the computing power intelligent scheduling method for a general-purpose computing power intelligent computing center according to any one of claims 1 to 5 are implemented.

9. A computer program product, characterized in that, Comprising computer instructions, where when the computer instructions are executed by a processor, the steps of the computing power intelligent scheduling method for a general-purpose computing power intelligent computing center according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Resource scheduling method and device, electronic equipment and storage medium

    CN112148468A

  • Resource management method, system and device, processor and electronic equipment

    CN116483558A