Intelligent computing power scheduling method and device for intelligent common computing power computing center

Through the intelligent computing power scheduling method of the intelligent computing center, the calculation resource ratio of the inference resource pool, training resource pool and elastic resource pool is dynamically adjusted, which solves the problem of low utilization of computing power resources in the intelligent computing center, realizes efficient resource utilization and cost reduction, and promotes the application of universal computing power.

CN119938340AActive Publication Date: 2025-05-06DATACANVAS LTD

Patent Information

Application Number
CN202510389932.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-06
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

In the prior art, the computing power resource utilization rate of intelligent computing centers is very low, resulting in shortage of inference clusters during peak hours and idle and wasted during low peak hours, increasing economic costs, and it is difficult to achieve widespread application of inclusive computing power.

Method used

A computing power intelligent scheduling method for the universal computing power intelligent computing center is proposed. By obtaining perceptual data, inputting a pre-trained computing model, determining the computing resource ratio in the target time period, and dynamically scheduling the computing resources of the inference resource pool, training resource pool and elastic resource pool based on this ratio.

Benefits of technology

Through the combination of data perception, intelligent computing and elastic scheduling, it adapts to the variability and timeliness of the computing resource requirements of the intelligent computing center, reduces the idleness or overload caused by static allocation of tasks, improves the utilization rate of computing resources, reduces the cost of model development, and realizes the wide application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938340A_ABST
    Figure CN119938340A_ABST
Patent Text Reader

Abstract

The invention provides a computing power intelligent scheduling method and device oriented to a common computing power intelligent computing center, and relates to the technical field of computing power infrastructure, and the method comprises the steps: S1, obtaining first sensing data which is used for calculating the load condition of each computing resource in a computing power resource pool in a target time period, the computing power resource pool comprises a reasoning resource pool, a training resource pool and an elastic resource pool; s2, inputting the first perception data into a pre-trained calculation model, and determining a first target proportion in a target time period; and S3, scheduling the computing resources corresponding to the resource pools based on the first target proportion. In this way, idle or overload of the reasoning task and the training task caused by static allocation of the computing resources is reduced, and the utilization rate of the computing resources of the intelligent computing center is increased; according to the method, redundant computing resources reserved for coping with peak values are reduced, the cost of model development is greatly reduced by dynamically adjusting hardware resource allocation, and wide application of the cost benefit is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a computing power intelligent scheduling method and device for a universal computing power intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have come into being.

[0003] "Intelligent computing center" refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing needs. It is the computing power to process information data and achieve target result output. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] At present, the computing resources of the intelligent computing center are usually used in two main scenarios, namely, model training scenario and model inference scenario. Based on these two scenarios, computing resources can be divided into independent training clusters and inference clusters. Among them, the training cluster is used for model development and optimization, while the inference cluster is used to process online service requests. However, in actual business, the load of the inference cluster shows an obvious periodic change pattern, namely the "computing power tide" phenomenon. For example, during peak hours (such as 8:00-24:00), the number of online service requests is large, resulting in a high load on the inference cluster, which is prone to slow service request response; during off-peak hours (such as 24:00-6:00), the number of online service requests is small, so the load on the inference cluster is low, and more computing resources are idle. At the same time, training tasks may be queued due to insufficient computing resources. The existence of the "computing power tide" phenomenon causes the inference cluster to be in short supply during peak hours and idle and wasted during off-peak hours. Whether it is increasing the deployment of inference clusters or idling inference clusters, it has invisibly increased the economic cost significantly, making it difficult to achieve the widespread application of inclusive computing power.

[0008] It can be seen that there is a problem of low utilization of computing resources in the existing technology. Summary of the invention

[0009] The embodiments of the present invention provide a computing power intelligent scheduling method and device for a universal computing power intelligent computing center to solve the problem of low utilization rate of computing power resources in the prior art.

[0010] To solve the above problems, the present invention is achieved as follows: In a first aspect, an embodiment of the present invention provides a computing power intelligent scheduling method for a universal computing power intelligent computing center, including: Step S1, obtaining first perception data, where the first perception data is used to calculate the load of each computing resource in a computing power resource pool within a target time period, where the computing power resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform reasoning tasks and training tasks; Step S2: input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; Step S3: Scheduling computing resources corresponding to each resource pool based on the first target ratio.

[0011] In one embodiment, after step S3, the method further includes: Step S4, counting the number of first actual values ​​outside the first valid interval to determine a first result, wherein the first actual value is an actual value obtained by monitoring the preset indicator in the first sub-time period, the first valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period; Step S5: when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, the acquired second perception data is input into the computing model to determine a second target ratio of a sub-time period after the first sub-time period in the target time period, the second target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool, the second perception number includes the first actual value and the first perception data, and the second perception number is used to calculate the load of each computing resource in the computing power resource pool in the sub-time period after the first sub-time period; Step S6: Scheduling the computing resources corresponding to each resource pool based on the second target ratio.

[0012] In one embodiment, before step S4, the method further includes: Step S7: acquiring an actual value for monitoring the preset indicator at each time interval to obtain a plurality of first actual values, wherein the first sub-time period includes a plurality of the time intervals, and each of the time intervals corresponds to a first actual value; Wherein, when there are at least N consecutive time intervals in the first sub-time period whose corresponding first actual values ​​are outside the first valid interval, the first result indicates that the error of the proportion corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or, When there are at least M first actual values ​​outside the first valid interval in the first sub-time period, the first result indicates that an error of a proportion corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

[0013] In one embodiment, after step S6, the method further includes: Step S8, counting the number of second actual values ​​outside the second valid interval to determine a second result, where the second actual value is an actual value obtained by monitoring the preset indicator in the second sub-time period, the second valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is a sub-time period after the first sub-time period; Step S9: When the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value, retrain the calculation model.

[0014] In one embodiment, step S3 includes: Step S31: In the computing power resource pool, when the ratio of the sum of the computing resources in the reasoning resource pool and the elastic resource pool to the computing resources in the training resource pool is less than the proportion of the computing resources in the reasoning resource pool in the first target ratio, a first computing resource is determined in the training resource pool based on the priority of the training task and / or the progress of the training task, the priority of the training task of the first computing resource is less than the priority of the training tasks of other computing resources in the training resource pool except the first computing resource, and / or the progress of the training task of the first computing resource is greater than a preset progress; Step S32: After the training task of the first computing resource is paused or ended, the first computing resource is determined as the computing resource in the reasoning resource pool to jointly perform the reasoning task, and the proportion of computing resources that jointly perform the reasoning task in the computing power resource pool is equal to the proportion of computing resources in the reasoning resource pool in the first target ratio.

[0015] In one embodiment, step S3 includes: Step S33: In the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the reasoning resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, a second computing resource is determined in the reasoning resource pool based on the load of the computing resources, and the load rate of the second computing resource is less than the load rate of other computing resources in the reasoning resource pool; Step S34: After the inference task of the second computing resource is completed, the second computing resource is determined as the computing resource in the training resource pool to jointly perform the training task, and the proportion of computing resources that jointly perform the training task in the computing resource pool is equal to the proportion of computing resources in the training resource pool in the first target ratio.

[0016] In a second aspect, an embodiment of the present invention further provides a computing power intelligent scheduling device for a universal computing power intelligent computing center, including: A first acquisition module is used to acquire first perception data, where the first perception data is used to calculate the load of each computing resource in a computing resource pool within a target time period, where the computing resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform both reasoning tasks and training tasks; A first determination module, configured to input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; The first scheduling module is used to schedule the computing resources corresponding to each resource pool based on the first target ratio.

[0017] In a third aspect, the present invention further provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps in the method for intelligent scheduling of computing power for a universal computing power intelligent computing center as described in the first aspect above are implemented.

[0018] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method for intelligent scheduling of computing power for a universal computing power intelligent computing center as described in the first aspect above are implemented.

[0019] In a fifth aspect, the present invention further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps in the method for intelligent scheduling of computing power for a universal computing power intelligent computing center as described in the first aspect above.

[0020] In an embodiment of the present invention, the first perception data is obtained to perceive the current load of each computing resource in the computing power resource pool; then the first perception data is input into a pre-trained computing model to determine the first target ratio in the target time period to achieve intelligent computing; finally, the computing resources corresponding to each resource pool are scheduled based on the first target ratio to achieve elastic scheduling. In this way, combining data perception, intelligent computing and elastic scheduling can adapt to the characteristics of variable and time-sensitive computing resource requirements of the intelligent computing center, reduce the idleness or overload of reasoning tasks and training tasks caused by static allocation of computing resources, and improve the utilization rate of computing resources in the intelligent computing center; and through the computing resources in the elastic resource pool, it can quickly respond to sudden demands, reduce the redundant computing resources reserved in the reasoning resource pool and the training resource pool to cope with peaks, and greatly reduce the cost of model development by dynamically adjusting the allocation of hardware resources, which is conducive to the widespread application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0022] Figure 1 This is one of the flow charts of a computing power intelligent scheduling method for a universal computing power intelligent computing center provided by an embodiment of the present invention; Figure 2 This is the second flowchart of a computing power intelligent scheduling method for a universal computing power intelligent computing center provided by an embodiment of the present invention; Figure 3 is a statistical chart corresponding to the first target ratio of the target time period provided by an embodiment of the present invention; Figure 4 It is a structural diagram of a computing power intelligent scheduling device for a universal computing power intelligent computing center provided by an embodiment of the present invention; Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data center to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and it mainly provides services to the society through computing power infrastructure.

[0025] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0026] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capacity of the computing power facilities, including the comprehensive capabilities of network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0027] The "Storage Power" (SP) mentioned in the present invention refers to: the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB=2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0028] The "computing power infrastructure" mentioned in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, which can realize centralized computing, storage, transmission and application of information.

[0029] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0030] The “computing power” mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0031] The “general computing power” mentioned in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0032] The "intelligent computing power" mentioned in the present invention refers to: a computing platform based on GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) and other dedicated chips for various innovative artificial intelligence applications, such as natural language processing, machine vision, etc.

[0033] The "super computing power" mentioned in the present invention refers to: the computing power mainly provided by high-performance computing clusters such as supercomputers. It uses the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0034] The "intelligent computing center" mentioned in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0035] The “intelligent computing center” mentioned in the present invention includes but is not limited to the “intelligent computing center”.

[0036] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0037] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0038] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide functions such as large-scale computing, storage and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0039] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0040] The "inclusive computing power" mentioned in the present invention refers to providing appropriate and effective computing power services at an affordable cost to all social classes and groups that have computing power service needs based on the requirements of equal opportunity and the principle of commercial sustainability.

[0041] The “model” mentioned in the present invention includes but is not limited to a “large language model” and a “multimodal large model”.

[0042] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale, designed to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0043] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, videos, audio, etc. for training, including but not limited to multimodal large language models.

[0044] The "agent" described in the present invention refers to an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware or a system with autonomy, adaptability and interaction capabilities. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then performs actions to affect the environment or achieve predetermined goals. Agents are widely used in the field of artificial intelligence and are commonly seen in automation systems, robots, virtual assistants, and game characters. The core of agents is the ability to learn autonomously and continuously evolve to better complete tasks and adapt to complex environments.

[0045] See also Figure 1 , Figure 1 This is one of the flow charts of a computing power intelligent scheduling method for a universal computing power intelligent computing center provided by an embodiment of the present invention, such as Figure 1 As shown, the following steps are included: Step S1, obtaining first perception data, where the first perception data is used to calculate the load of each computing resource in a computing power resource pool within a target time period, where the computing power resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform reasoning tasks and training tasks; In this step, the first perception data may include a first indicator and a second indicator. The first indicator may include the real-time collected graphics processing unit (GPU) utilization, task queue depth, and query rate per second (QPS) of inference requests; the second indicator may include collected environmental factor data, time factor data, and event factor data. The first indicator and the second indicator are factors that affect the load of each computing resource in the computing resource pool. For example, high GPU utilization, many task queues, and high QPS can directly reflect that the corresponding computing resources have a large load; while environmental factor data, time factor data, and event factor data indirectly affect the load of computing resources. For example, when there are many users in a certain area, during working hours, and when hot events occur, it can be foreseen that the load of computing resources corresponding to these environments, times, and times is large. Therefore, through the first perception data, the load of each computing resource in the computing resource pool within the target time period can be calculated.

[0046] Among them, the target time period can be a period of time after the time point of obtaining the first perception data. For example, the first perception data is obtained at 8:00 to calculate the load of each computing resource in the computing resource pool within the target time period, so as to achieve a reasonable allocation of computing resources within the target time period and improve resource utilization. Exemplarily, the target time period can be four hours from 9:00 to 13:00. Of course, the target time period can also be other time periods, or longer or shorter times, which are not limited here.

[0047] Among them, in this embodiment, the computing resource pool can be divided into three categories, namely, the reasoning resource pool, the training resource pool and the elastic resource pool. Each type of resource pool corresponds to computing resources, and the sum of the computing resources in the reasoning resource pool, the training resource pool and the elastic resource pool is the number of computing resources in the computing resource pool. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform reasoning tasks and training tasks. When the priority of the reasoning task is higher than that of the training task, the computing resources in the elastic resource pool can be called to perform the reasoning task; in some cases, the computing resources in the training resource pool can also be called to perform the reasoning task. When the priority of the training task is higher than that of the reasoning task, the computing resources in the elastic resource pool can be called to perform the training task; in some cases, the computing resources in the reasoning resource pool can also be called to perform the training task. In this way, the idle computing resources can be flexibly scheduled according to the load of each computing resource in the computing resource pool to adapt to the periodic changes in the corresponding load of the reasoning task (that is, the "computing power tide" phenomenon), and during the off-peak period of the reasoning task, reduce the situation where the training task is queued due to insufficient computing resources in the training resource pool.

[0048] Step S2: input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; In this step, the computing model can be a multi-classification model pre-trained with sample perception data as features. Sample perception data can include historical records of the load information of various computing resources at different times, whether it is a holiday, whether there is a special event, GPU utilization in computing resources, task queue depth, QPS and other indicators. Exemplarily, the data can be divided into a training set (historical data) and a test set (recent data) in chronological order to avoid data leakage caused by chronological order.

[0049] The output of the computing model may be the ratio of computing resources corresponding to each resource pool in the computing resource pool, such as [reasoning: training: elasticity]. The target time period may be a longer time period, so there may be multiple different ratios in the target time period. In this way, the first perception data is input into the pre-trained computing model to determine the first target ratio of the target time period, thereby realizing the prediction of the long-term trend of the computing resource ratio corresponding to each resource pool in the computing resource pool.

[0050] Exemplarily, the target time period may be 9:00 to 13:00. The first perception data is input into a pre-trained computing model to determine that the first target ratio from 9:00 to 13:00 may include [5:4:1], [6:4:0], and [7:3:0]. Among them, [5:4:1] may be the quantity ratio between the computing resources corresponding to the reasoning resource pool, the training resource pool, and the elastic resource pool from 9:00 to 10:00; [6:4:0] may be the quantity ratio between the computing resources corresponding to the reasoning resource pool, the training resource pool, and the elastic resource pool from 10:00 to 12:00; and [7:3:0] may be the quantity ratio between the computing resources corresponding to the reasoning resource pool, the training resource pool, and the elastic resource pool from 12:00 to 13:00. It can be seen from the calculated first target ratio that: According to the first perception data, it is predicted that there will not be many inference tasks from 9:00 to 10:00. At this time, the computing resource ratio in the inference resource pool is 5, which can handle the inference tasks from 9:00 to 10:00. At the same time, the computing resource ratio in the training resource pool is 4 for normal training tasks, and the computing resource ratio in the elastic resource pool is 1 to cope with emergencies; According to the first perception data, it is predicted that the reasoning tasks will increase from 10:00 to 12:00. At this time, the computing resources in the reasoning resource pool are allocated to 6, which can alleviate the load pressure caused by the increased reasoning tasks from 10:00 to 12:00. At the same time, the computing resources in the training resource pool are still allocated to 4 for normal training tasks. The computing resources in the elastic resource pool will be scheduled to the reasoning resource pool, so the computing resources in the elastic resource pool are allocated to 0.

[0051] According to the first perception data, it is predicted that the reasoning tasks will further increase from 12:00 to 13:00. At this time, the computing resources in the reasoning resource pool are allocated to 7, which can alleviate the load pressure caused by the further increase in reasoning tasks from 12:00 to 13:00; while performing training tasks, the computing resources with a ratio of 1 in the training resource pool can be scheduled to the reasoning resource pool to meet the needs of the reasoning tasks, so the computing resources ratio in the training resource pool is reduced to 3; and the computing resources in the elastic resource pool will be scheduled to the reasoning resource pool, so the computing resources ratio in the elastic resource pool is 0.

[0052] The first target ratio may also be other situations, for example, computing resources in the reasoning resource pool and / or the elastic resource pool may be dispatched to the training resource pool. During the off-peak period of reasoning tasks, the situation where training tasks have to wait in line due to insufficient computing resources in the training resource pool is reduced.

[0053] In this way, the current load is monitored in real time through the first perception data, and the demand of each resource pool in the target time period is predicted, so as to determine the appropriate first target ratio and then schedule the corresponding ratio. By combining data perception, intelligent computing and elastic scheduling, it can adapt to the characteristics of variable and time-sensitive computing resource requirements of the intelligent computing center, reduce the idleness or overload of reasoning tasks and training tasks caused by static allocation of computing resources, and improve the utilization rate of computing resources in the intelligent computing center; and set up an elastic resource pool, through which the computing resources in the elastic resource pool can quickly respond to sudden demands (such as temporary tasks), reduce the redundant computing resources reserved in the reasoning resource pool and training resource pool to cope with peaks, and greatly reduce the cost of model development by dynamically adjusting the allocation of hardware resources, so as to realize the widespread application of inclusive computing power.

[0054] Step S3: Scheduling computing resources corresponding to each resource pool based on the first target ratio.

[0055] In this step, after determining the first target ratio, the computing resources corresponding to each resource pool are scheduled based on the first target ratio, thereby dynamically adjusting the hardware resource allocation, improving the task throughput, and thus improving the utilization rate of the computing resources of the intelligent computing center.

[0056] Specifically, Figure 2 As shown, Figure 2This is the second flowchart of a computing power intelligent scheduling method for a universal computing power intelligent computing center provided by an embodiment of the present invention, which can be applied to an intelligent body. A computing power resource pool is deployed in the intelligent body, and the computing power resource pool includes an inference resource pool, a training resource pool, and an elastic resource pool; in addition, a tidal perception layer, an intelligent computing layer, and an elastic scheduling layer are also provided in the intelligent body. The tidal sensing layer is used to obtain the first perception data, and the first perception data includes a first indicator and a second indicator. The first indicator may include real-time collected GPU utilization, task queue depth, QPS of inference requests, etc., and the second indicator may include collected environmental factor data, time factor data, event factor data, and other indicators, and data perception is realized through the tidal sensing layer. The intelligent computing layer is used to determine the first target ratio in the target time period according to the first perception data, thereby realizing intelligent computing. The elastic scheduling layer is used to schedule the computing resources corresponding to each resource pool based on the first target ratio, thereby realizing elastic scheduling of computing resource ratios. Among them, the computing power resource pool can schedule the computing resources corresponding to each resource pool based on the scheduling instruction sent by the elastic scheduling layer, and the scheduling instruction includes the first target ratio information. By dynamically adjusting the allocation of hardware resources, the idleness or overload of inference and training tasks caused by the static allocation of computing resources is reduced, the utilization rate of computing resources in the intelligent computing center is improved, the cost of model development is greatly reduced, and it is conducive to the widespread application of inclusive computing power.

[0057] In an embodiment of the present invention, the first perception data is obtained to perceive the current load of each computing resource in the computing power resource pool; then the first perception data is input into a pre-trained computing model to determine the first target ratio in the target time period to achieve intelligent computing; finally, the computing resources corresponding to each resource pool are scheduled based on the first target ratio to achieve elastic scheduling. In this way, combining data perception, intelligent computing and elastic scheduling can adapt to the characteristics of variable and time-sensitive computing resource requirements of the intelligent computing center, reduce the idleness or overload of reasoning tasks and training tasks caused by static allocation of computing resources, and improve the utilization rate of computing resources in the intelligent computing center; and through the computing resources in the elastic resource pool, it can quickly respond to sudden demands, reduce the redundant computing resources reserved in the reasoning resource pool and the training resource pool to cope with peaks, and greatly reduce the cost of model development by dynamically adjusting the allocation of hardware resources, and realize the wide application of inclusive computing power.

[0058] In one embodiment, after step S3, the method further includes: Step S4, counting the number of first actual values ​​outside the first valid interval to determine a first result, wherein the first actual value is an actual value obtained by monitoring the preset indicator in the first sub-time period, the first valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period; Step S5: when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, the acquired second perception data is input into the computing model to determine a second target ratio of a sub-time period after the first sub-time period in the target time period, the second target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool, the second perception number includes the first actual value and the first perception data, and the second perception number is used to calculate the load of each computing resource in the computing power resource pool in the sub-time period after the first sub-time period; Step S6: Scheduling the computing resources corresponding to each resource pool based on the second target ratio.

[0059] In this embodiment, during the process of scheduling the computing resources corresponding to each resource pool according to the first target ratio in the target time period, it is considered that if there is a large error in the calculated first target ratio, some computing resources in the computing resource pool will still be idle or overloaded. In order to reduce this calculation error, the calculation ratio can be corrected according to the short-term trend in the target time period, which can be specifically described as follows: For example, Figure 3 As shown, the target time period may include multiple sub-time periods. The first target ratio between the computing resources in the reasoning resource pool, the computing resources in the training resource pool, and the computing resources in the elastic resource pool to be executed within the target time period (taking 9:00 to 13:00 as an example) is calculated as follows: the corresponding ratio from 9:00 to 10:00 is [5:4:1]; the corresponding ratio from 10:00 to 12:00 is [6:4:0]; the corresponding ratio from 12:00 to 13:00 is [7:3:0].

[0060] An actual value for monitoring a preset indicator can be obtained at preset time intervals, for example, every 15 minutes. In scenarios dominated by reasoning tasks, the preset indicators may include the QPS of reasoning requests; in scenarios dominated by training tasks, the preset indicators may include the queue depth of training tasks. The following is an illustrative explanation using the scenario dominated by reasoning tasks as an example. It should be understood that the same is applicable to scenarios dominated by training tasks.

[0061] After step S3, the computing resources corresponding to each resource pool are scheduled based on the first target ratio. Therefore, the corresponding ratio from 9:00 to 10:00 is [5:4:1]. Assuming that the first sub-time period is the sub-time period from 9:00 to 10:00 in 9:00 to 13:00, then the QPS of the reasoning request collected every 15 minutes starting from 9:00 can obtain 4 actual values ​​corresponding to the first sub-time period, that is, 4 first actual values. Among them, according to the corresponding ratio of [5:4:1] from 9:00 to 10:00, the theoretical upper limit and theoretical lower limit corresponding to the reasoning request QPS under this ratio can be calculated, thereby determining the effective value range of the reasoning request QPS. At this time, the first effective interval is the effective value range of the reasoning request QPS under the ratio of [5:4:1] of each computing resource from 9:00 to 10:00. In this way, the number of first actual values ​​outside the first effective interval can be counted to determine the first result.

[0062] Among them, in this example, the sub-time period from 9:00 to 10:00 in 9:00 to 13:00 is used as the first sub-time period for explanation. It should be understood that the first sub-time period can be any sub-time period within the target time period. In this example, if the first result indicates that the ratio corresponding to the first sub-time period, that is, the error of [5:4:1] is less than or equal to the preset value, it means that the first target ratio calculated by the calculation model based on the first perception data is relatively accurate, and the other sub-time periods after the first sub-time period (i.e., 10:00 to 13:00) can still be scheduled according to the ratio corresponding to each sub-time period, thereby improving the utilization rate of the computing resources of the intelligent computing center. If the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value, it means that the first target ratio calculated by the calculation model based on the first perception data is relatively inaccurate, and the other sub-time periods after the first sub-time period (i.e., 10:00 to 13:00) still perform scheduling according to the ratio corresponding to each sub-time period, and errors will still occur when scheduling is performed according to the ratio corresponding to each sub-time period. Therefore, it is necessary to correct the computing ratio: input the acquired second perception data into the computing model, determine the second target ratio of the sub-time period after the first sub-time period in the target time period, and the second perception number includes the first actual value and the first perception data. The second perception number is used to calculate the load of each computing resource in the computing resource pool in the sub-time period after the first sub-time period. Among them, the sub-time period after the first sub-time period can be a sub-time period within the target time period, such as 10:00 to 13:00; it can also be a sub-time period outside the target time period, such as 10:00 to 14:00. In this way, the computing ratio is corrected by the short-term trend in the target time period, and the second target ratio is obtained again; then the computing resources corresponding to each resource pool are scheduled based on the second target ratio. The accuracy of calculating the computing resource ratio in each resource pool is improved.

[0063] In one embodiment, before step S4, the method further includes: Step S7: acquiring an actual value for monitoring the preset indicator at each time interval to obtain a plurality of first actual values, wherein the first sub-time period includes a plurality of the time intervals, and each of the time intervals corresponds to a first actual value; Wherein, when there are at least N consecutive time intervals in the first sub-time period whose corresponding first actual values ​​are outside the first valid interval, the first result indicates that the error of the proportion corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or, When there are at least M first actual values ​​outside the first valid interval in the first sub-time period, the first result indicates that an error of a proportion corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

[0064] In one example, it is judged that the first target ratio calculated by the calculation model based on the first perception data is relatively inaccurate. In other words, the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than the preset value. It can be determined based on the first actual value corresponding to at least N consecutive time intervals in the first sub-time period being outside the first valid interval. That is, it can be judged based on the number of consecutive errors. For example, an actual value for monitoring a preset indicator is obtained at every preset time interval, such as every 15 minutes. If there are at least 3 consecutive time intervals from 9:00 to 10:00 and the first actual value corresponding to the time interval is outside the first valid interval, it can be determined that the ratio [5:4:1] is not in line with reality and needs to be recalculated. The ratio of computing resources in each resource pool at the current time can be temporarily scheduled through conventional ratios to respond to inference requests in a timely manner.

[0065] In another example, it is also possible to judge whether the first target ratio is accurate based on the fact that there are at least M first actual values ​​outside the first valid interval in the first sub-time period. In other words, it can be judged based on the total number of times errors occur. For example, at every preset time interval, such as every 15 minutes, an actual value for monitoring a preset indicator is obtained. If the number of errors reached 3 times between 9:00 and 10:00, that is, at least 3 first actual values ​​were outside the first valid interval, it can be determined that the ratio [5:4:1] is not in line with reality and needs to be recalculated. The ratio of computing resources in each resource pool at the current time can also be temporarily scheduled through conventional ratios to respond to inference requests in a timely manner.

[0066] In one embodiment, after step S6, the method further includes: Step S8, counting the number of second actual values ​​outside the second valid interval to determine a second result, where the second actual value is an actual value obtained by monitoring the preset indicator in the second sub-time period, the second valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is a sub-time period after the first sub-time period; Step S9: When the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value, retrain the calculation model.

[0067] In this embodiment, in the process of scheduling the computing resources corresponding to each resource pool according to the first target ratio in the target time period, it can be determined based on the first result that the error of the ratio corresponding to the first sub-time period is large, so the sub-time period after the first sub-time period is recalculated to obtain the second target ratio, and the computing resources corresponding to each resource pool are rescheduled based on the second target ratio. In this way, the computing ratio is corrected according to the short-term trend in the target time period, which improves the accuracy of the calculation result. Furthermore, in the process of scheduling the computing resources corresponding to each resource pool according to the second target ratio in the sub-time period after the first sub-time period, it is considered that if there is a large error in the recalculated second target ratio, some computing resources in the computing power resource pool will still be idle or overloaded. In order to reduce this calculation error, the calculation ratio can be corrected according to the long-term trend of the first sub-time period and other sub-time periods after the first sub-time period. For details, please refer to the following description: Exemplarily, assuming that the second sub-time period is the sub-time period from 10:00 to 11:00 in 10:00 to 13:00, then the QPS of the inference request collected every 15 minutes starting from 10:00 can obtain 4 actual values ​​corresponding to the second sub-time period, that is, 4 second actual values. Among them, according to the ratio [6:4:0] corresponding to 10:00 to 12:00, the theoretical upper limit and theoretical lower limit corresponding to the inference request QPS under this ratio can be calculated, thereby determining the effective value range of the inference request QPS. At this time, the second effective interval is the effective value range of the inference request QPS under the ratio of [6:4:0] of each computing resource during the period from 10:00 to 11:00. In this way, the number of second actual values ​​outside the second effective interval can be counted to determine the second result.

[0068] In this example, the sub-time period from 9:00 to 11:00 in 10:00 to 13:00 is used as the second sub-time period for explanation. It should be understood that the second sub-time period can be any sub-time period after the first sub-time period (9:00 to 10:00). In this example, if the second result indicates that the ratio corresponding to the second sub-time period, that is, the error of [6:4:0] is less than or equal to the preset value, it means that the second target ratio calculated by the calculation model based on the second perception data is relatively accurate, and the other sub-time periods after the second sub-time period (i.e., 11:00 to 13:00) can still be scheduled according to the ratio corresponding to each sub-time period, thereby improving the utilization rate of the computing resources of the intelligent computing center. If the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value, it means that the second target ratio calculated by the calculation model based on the second perception data is relatively inaccurate, and the other sub-time periods after the second sub-time period still perform scheduling according to the ratio corresponding to each sub-time period. Therefore, it is necessary to correct the calculation ratio: considering that before this, the first target ratio calculated by the calculation model based on the first perception data is also inaccurate. If the correction at this time is still to update the perception data input into the computing model, the subsequent computing results will still be inaccurate. Therefore, in this case, the computing model is retrained to ensure the accuracy of the computing resource allocation in each resource pool.

[0069] It should be understood that, in consideration of the existence of accidental factors, the computing model may be retrained if the perception data input into the computing model is updated multiple times and the subsequent computing results are still inaccurate.

[0070] In some optional embodiments, step S3 includes: Step S31: In the computing power resource pool, when the ratio of the sum of the computing resources in the reasoning resource pool and the elastic resource pool to the computing resources in the training resource pool is less than the proportion of the computing resources in the reasoning resource pool in the first target ratio, a first computing resource is determined in the training resource pool based on the priority of the training task and / or the progress of the training task, the priority of the training task of the first computing resource is less than the priority of the training tasks of other computing resources in the training resource pool except the first computing resource, and / or the progress of the training task of the first computing resource is greater than a preset progress; Step S32: After the training task of the first computing resource is paused or ended, the first computing resource is determined as the computing resource in the reasoning resource pool to jointly perform the reasoning task, and the proportion of computing resources that jointly perform the reasoning task in the computing power resource pool is equal to the proportion of computing resources in the reasoning resource pool in the first target ratio.

[0071] In this embodiment, when the ratio of the sum of the computing resources of the reasoning resource pool plus the elastic resource pool to the computing resources of the training resource pool is less than the proportion of computing resources in the reasoning resource pool in the first target ratio, it means that the current reasoning resources are insufficient and resources need to be allocated from the training resource pool. For example: Assuming that the first target ratio requires that the reasoning resources account for 60%, and the current reasoning + elasticity only accounts for 40%, 20% of the resources need to be released from the training resource pool. Determining the 20% of computing resources (i.e., the first computing resources) that need to be released from the training resource pool can be: selecting computing resources with lower training task priority (for example, offline training tasks have lower priority than online model fine-tuning tasks); selecting computing resources corresponding to tasks whose training progress has exceeded a preset threshold (such as 80%), because the cost of recovery after suspension is low. Among them, the task priority, progress, and computing power usage of each computing resource can be collected through the monitoring system, and sorted in ascending order of priority and descending order of progress to determine a candidate list that can be suspended.

[0072] Then, a checkpoint save instruction is issued to the selected first computing resource to record the training status (such as model parameters, optimizer status); after stopping the training task, the first computing resource is released, and the released first computing resource is added to the inference resource pool for performing the inference task; the first computing resource is determined as a computing resource in the inference resource pool through a service grid (such as Istio) or a load balancer (such as Nginx) to jointly perform the inference task, and the proportion of computing resources that jointly perform the inference task in the computing power resource pool is equal to the proportion of computing resources in the inference resource pool in the first target ratio.

[0073] When the inference resource pool reaches the first target ratio, resource migration stops. If there are still remaining released resources, they can be returned to the elastic resource pool for dynamic application by subsequent tasks. If the inference load decreases, the resources can be transferred back to the training pool according to the opposite logic. In this way, without adding hardware, resource reuse for training and inference is achieved. This improves the utilization rate of computing resources in the intelligent computing center.

[0074] In some optional embodiments, step S3 includes: Step S33: In the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the reasoning resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, a second computing resource is determined in the reasoning resource pool based on the load of the computing resources, and the load rate of the second computing resource is less than the load rate of other computing resources in the reasoning resource pool; Step S34: After the inference task of the second computing resource is completed, the second computing resource is determined as the computing resource in the training resource pool to jointly perform the training task, and the proportion of computing resources that jointly perform the training task in the computing resource pool is equal to the proportion of computing resources in the training resource pool in the first target ratio.

[0075] In this embodiment, when the ratio of the sum of the computing resources of the training resource pool plus the elastic resource pool to the computing resources of the reasoning resource pool is less than the proportion of computing resources in the training resource pool in the first target ratio, it means that the current training resources are insufficient and resources need to be allocated from the reasoning resource pool. For example: Assuming that the first target ratio requires that the training resources account for 40%, and the current training + elasticity only accounts for 30%, 10% of the resources need to be released from the reasoning resource pool. Determining the 10% of computing resources (i.e., the second computing resources) that need to be released from the reasoning resource pool can be: selecting computing resources with a low load rate of reasoning tasks (such as CPU / GPU utilization below a threshold), where the load rate calculation can be combined with indicators such as task queue depth and response delay; giving priority to releasing general reasoning resources (such as non-real-time reasoning services) to avoid affecting key businesses. Among them, the load rate, task queue length and other indicators of each reasoning node can be collected in real time through tools such as Prometheus, and sorted in ascending order by load rate, and the cluster group with the lowest load is selected as the release target.

[0076] Then, an instruction to stop receiving new requests is issued to the selected second computing resource, waiting for the current request to be processed; after closing the inference service instance, the second computing resource is released, and the released second computing resource is added to the training resource pool for training tasks; the second computing resource is determined as a computing resource in the training resource pool to jointly execute the training task, and the proportion of computing resources that jointly execute the training task in the computing power resource pool is equal to the proportion of computing resources in the training resource pool in the first target ratio.

[0077] When the training resource pool reaches the target ratio, resource migration stops. If there are still remaining released resources, they can be retained in the elastic resource pool for subsequent dynamic allocation. If the training load decreases, the resources can be transferred back to the inference pool according to the opposite logic. In this way, without adding hardware, resource reuse for training and inference is achieved, which improves the utilization rate of computing resources in the intelligent computing center.

[0078] See also Figure 4 , Figure 4 is a structural diagram of a computing power intelligent scheduling device for a universal computing power intelligent computing center provided by an embodiment of the present invention, such as Figure 4 As shown, the computing power intelligent scheduling device 400 for the inclusive computing power intelligent computing center includes: A first acquisition module 401 is used to acquire first perception data, where the first perception data is used to calculate the load of each computing resource in a computing resource pool within a target time period, where the computing resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform both reasoning tasks and training tasks; A first determination module 402 is used to input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, where the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; The first scheduling module 403 is used to schedule the computing resources corresponding to each resource pool based on the first target ratio.

[0079] In one embodiment, the apparatus further comprises: a first statistical module, configured to count the number of first actual values ​​outside a first valid interval to determine a first result, wherein the first actual value is an actual value obtained by monitoring a preset indicator in a first sub-time period, the first valid interval is a valid value interval of the preset indicator calculated according to a ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period; A second determination module is used to input the acquired second perception data into the computing model to determine a second target ratio of a sub-time period after the first sub-time period in the target time period when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, wherein the second target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool, the second perception number includes the first actual value and the first perception data, and the second perception number is used to calculate the load of each computing resource in the computing power resource pool in the sub-time period after the first sub-time period; The second scheduling module is used to schedule the computing resources corresponding to each resource pool based on the second target ratio.

[0080] In one embodiment, the apparatus further comprises: A second acquisition module is used to acquire an actual value for monitoring the preset indicator at each time interval to obtain multiple first actual values, wherein the first sub-time period includes multiple time intervals, and each time interval corresponds to one first actual value; Wherein, when there are at least N consecutive time intervals in the first sub-time period whose corresponding first actual values ​​are outside the first valid interval, the first result indicates that the error of the proportion corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or, When there are at least M first actual values ​​outside the first valid interval in the first sub-time period, the first result indicates that an error of a proportion corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

[0081] In one embodiment, the apparatus further comprises: a second statistical module, configured to count the number of second actual values ​​outside a second valid interval to determine a second result, wherein the second actual value is an actual value obtained by monitoring the preset indicator in a second sub-time period, the second valid interval is a valid value interval of the preset indicator calculated according to a ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is a sub-time period after the first sub-time period; A retraining module is used to retrain the calculation model when the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value.

[0082] In one embodiment, the first scheduling module 403 is specifically used to: In the computing power resource pool, when the ratio of the sum of the computing resources in the reasoning resource pool and the elastic resource pool to the computing resources in the training resource pool is less than the proportion of the computing resources in the reasoning resource pool in the first target ratio, a first computing resource is determined in the training resource pool based on the priority of the training task and / or the progress of the training task; After the training task of the first computing resource is paused or ended, the first computing resource is determined as the computing resource in the reasoning resource pool to jointly perform the reasoning task, and the proportion of computing resources that jointly perform the reasoning task in the computing resource pool is equal to the proportion of computing resources in the reasoning resource pool in the first target ratio.

[0083] In one embodiment, the first scheduling module 403 is specifically used to: In the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the reasoning resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, determining the second computing resource in the reasoning resource pool based on the load of the computing resources; After the inference task of the second computing resource is completed, the second computing resource is determined as the computing resource in the training resource pool to jointly perform the training task, and the proportion of computing resources that jointly perform the training task in the computing resource pool is equal to the proportion of computing resources in the training resource pool in the first target ratio.

[0084] The computing power intelligent scheduling device for the universal computing power intelligent computing center provided by the embodiment of the present invention is capable of realizing the various processes of the various embodiments of the above-mentioned computing power intelligent scheduling method for the universal computing power intelligent computing center. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0085] It should be noted that the computing power intelligent scheduling device for the inclusive computing power intelligent computing center in the embodiment of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0086] The embodiment of the present invention further provides an electronic device, see Figure 5 , Figure 5 is a schematic diagram of the structure of an electronic device provided by the present invention, the electronic device includes a memory 501, a processor 502 and a program or instruction stored in the memory 501 and running on the memory 501. When the program or instruction is executed by the processor 502, Figure 1 Any steps in the corresponding embodiment of the computing power intelligent scheduling method for the inclusive computing power intelligent computing center and the same beneficial effects achieved will not be repeated here.

[0087] The processor 502 may be a CPU, an ASIC, an FPGA or a GPU.

[0088] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiment of the intelligent computing power scheduling method for the universal computing power intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0089] The embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above Figure 1 Any steps in the corresponding embodiments of the computing power intelligent scheduling method for the inclusive computing power intelligent computing center can achieve the same technical effect, and will not be repeated here to avoid repetition. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0090] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The corresponding processes of the implementation method of the computing power intelligent scheduling method for the inclusive computing power intelligent computing center can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0091] The terms "first", "second" etc. in the embodiments of the present invention are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. In addition, the terms "include" and "have" and any variation thereof are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment. In addition, "and / or" is used in the present application to represent at least one of the connected objects, such as A and / or B and / or C, indicating that A alone, B alone, C alone, and A and B all exist, B and C all exist, A and C all exist, and 7 situations in which A, B and C all exist.

[0092] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a second terminal device, etc.) to execute the methods of each embodiment of the present application.

[0094] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A computing power intelligent scheduling method for a universal computing power intelligent computing center, characterized in that: include: Step S1, obtaining first perception data, where the first perception data is used to calculate the load of each computing resource in a computing power resource pool within a target time period, where the computing power resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform reasoning tasks and training tasks; Step S2: input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; Step S3: Scheduling computing resources corresponding to each resource pool based on the first target ratio.

2. The method according to claim 1, characterized in that After step S3, the method further includes: Step S4, counting the number of first actual values ​​outside the first valid interval to determine a first result, wherein the first actual value is an actual value obtained by monitoring the preset indicator in the first sub-time period, the first valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the first sub-time period in the first target ratio, and the first sub-time period is any sub-time period within the target time period; Step S5: when the first result indicates that the error of the ratio corresponding to the first sub-time period is greater than a preset value, the acquired second perception data is input into the computing model to determine a second target ratio of a sub-time period after the first sub-time period in the target time period, the second target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool, the second perception number includes the first actual value and the first perception data, and the second perception number is used to calculate the load of each computing resource in the computing power resource pool in the sub-time period after the first sub-time period; Step S6: Scheduling the computing resources corresponding to each resource pool based on the second target ratio.

3. The method according to claim 2, characterized in that Before step S4, the method further includes: Step S7: acquiring an actual value for monitoring the preset indicator at each time interval to obtain a plurality of first actual values, wherein the first sub-time period includes a plurality of the time intervals, and each of the time intervals corresponds to a first actual value; Wherein, when there are at least N consecutive time intervals in the first sub-time period whose corresponding first actual values ​​are outside the first valid interval, the first result indicates that the error of the proportion corresponding to the first sub-time period is greater than the preset value, and N is a positive integer; or, When there are at least M first actual values ​​outside the first valid interval in the first sub-time period, the first result indicates that an error of a proportion corresponding to the first sub-time period is greater than the preset value, and M is a positive integer.

4. The method according to claim 2, characterized in that After step S6, the method further includes: Step S8, counting the number of second actual values ​​outside the second valid interval to determine a second result, where the second actual value is an actual value obtained by monitoring the preset indicator in the second sub-time period, the second valid interval is a valid value interval of the preset indicator calculated according to the ratio corresponding to the second sub-time period in the second target ratio, and the second sub-time period is a sub-time period after the first sub-time period; Step S9: When the second result indicates that the error of the ratio corresponding to the second sub-time period is greater than the preset value, retrain the calculation model.

5. The method according to any one of claims 1 to 4, characterized in that The step S3 comprises: Step S31: in the computing power resource pool, when the ratio of the sum of the computing resources in the reasoning resource pool and the elastic resource pool to the computing resources in the training resource pool is less than the proportion of the computing resources in the reasoning resource pool in the first target ratio, determine the first computing resource in the training resource pool based on the priority of the training task and / or the progress of the training task; Step S32: After the training task of the first computing resource is paused or ended, the first computing resource is determined as the computing resource in the reasoning resource pool to jointly perform the reasoning task, and the proportion of computing resources that jointly perform the reasoning task in the computing power resource pool is equal to the proportion of computing resources in the reasoning resource pool in the first target ratio.

6. The method according to any one of claims 1 to 4, characterized in that The step S3 comprises: Step S33: in the computing power resource pool, when the ratio of the sum of the computing resources in the training resource pool and the elastic resource pool to the computing resources in the reasoning resource pool is less than the proportion of the computing resources in the training resource pool in the first target ratio, determine the second computing resource in the reasoning resource pool based on the load of the computing resources; Step S34: After the inference task of the second computing resource is completed, the second computing resource is determined as the computing resource in the training resource pool to jointly perform the training task, and the proportion of computing resources that jointly perform the training task in the computing resource pool is equal to the proportion of computing resources in the training resource pool in the first target ratio.

7. A computing power intelligent scheduling device for a universal computing power intelligent computing center, characterized in that: include: A first acquisition module is used to acquire first perception data, where the first perception data is used to calculate the load of each computing resource in a computing resource pool within a target time period, where the computing resource pool includes a reasoning resource pool, a training resource pool, and an elastic resource pool, where computing resources are preset in the reasoning resource pool, the training resource pool, and the elastic resource pool, respectively. The computing resources in the reasoning resource pool are used to perform reasoning tasks, the computing resources in the training resource pool are used to perform training tasks, and the computing resources in the elastic resource pool can perform both reasoning tasks and training tasks; A first determination module, configured to input the first perception data into a pre-trained computing model to determine a first target ratio in the target time period, wherein the first target ratio includes at least one ratio between computing resources corresponding to each resource pool in the computing power resource pool; The first scheduling module is used to schedule the computing resources corresponding to each resource pool based on the first target ratio.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the computing power intelligent scheduling method for a universal computing power intelligent computing center are implemented as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the computing power intelligent scheduling method for a universal computing power intelligent computing center as described in any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the computing power intelligent scheduling method for a universal computing power intelligent computing center as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Resource scheduling method and device, electronic equipment and storage medium

    CN112148468A

  • Training and reasoning integrated deep learning GPU cluster scheduling method

    CN116048802A

  • Resource management method, system and device, processor and electronic equipment

    CN116483558A

  • Resource adjustment method and device of edge computing node, equipment and medium

    CN119645618A

  • Resource control method for function computing, device, and medium

    US20230273830A1

Cited By

  • Resource allocation method, machine learning platform, equipment and storage medium

    CN120144328A

  • Computing system, computing power distribution method, electronic device, and readable storage medium

    CN120256133A

  • Method and device for intelligent computing center cloud platform to adjust large model training task based on computing power use state

    CN120407211A

  • Computing power routing addressing method and device of intelligent computing center cloud platform

    CN120547112A

  • Tidal computing power scheduling system and scheduling method oriented to heterogeneous computing power resource pool

    CN120950209A