Multi-region computing power scheduling method and device for inclusive computing power intelligent computing center

By performing multi-region scheduling of computing power resources between intelligent computing centers, the high economic cost problem of users when performing tasks on intelligent computing centers is solved, and efficient scheduling of computing power resources and wide application of universal computing power is achieved.

CN119902876BActive Publication Date: 2025-06-24DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510380490.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the prior art, when users rely on the computing power resources of the intelligent computing center to perform tasks, they face computing power resource scheduling bottlenecks and high economic costs, which restrict the promotion and application of inclusive computing power.

Method used

Provide a multi-region computing power scheduling method for the general computing power intelligent computing center. By determining the tasks to be migrated, filtering optional intelligent computing centers in multiple alternative intelligent computing centers, calculating the migration and continuing to execute costs, and deciding whether to migrate the tasks based on cost comparison, we can find the target intelligent computing center with the best electricity price to continue running the tasks.

Benefits of technology

It realizes efficient scheduling and optimized configuration of computing power resources, significantly reduces the economic cost of users using computing power resources in intelligent computing centers, promotes cross-regional sharing and collaboration of computing power resources, and realizes the wide application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902876B_ABST
    Figure CN119902876B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and provides a multi-region computing power scheduling method and device for an intelligent computing center for inclusive computing power, including: determining and based on the task to be migrated, determining at least one optional intelligent computing center among each alternative intelligent computing center, and determining whether there is a target intelligent computing center among the at least one optional intelligent computing center according to the task to be migrated; if so, migrating the current intermediate operation result of the task to be migrated to the target intelligent computing center; and continuing to run the task to be migrated based on the computing power of the target intelligent computing center and the intermediate operation result. Thus, through an intelligent task migration mechanism, efficient scheduling and optimal allocation of computing power resources are achieved. On the premise of ensuring the smooth execution of tasks, the economic cost of users using the computing power resources of the intelligent computing center is greatly reduced, the cross-region sharing and collaboration of computing power resources are promoted, and the wide application of inclusive computing power is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and specifically relates to a multi-region computing power scheduling method and device for an intelligent computing center for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", is the ability of computer devices or computing / data centers to process information, is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, is the computing ability to achieve the output of target results through processing information data, is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] However, currently, when tasks are executed based on the computing power resources of an intelligent computing center (such as a model training task), the tasks usually run fixed in a single intelligent computing center, resulting in users bearing huge electricity bills during peak electricity prices or in high-cost regions. At the same time, the lack of collaborative ability of cross-region computing power resources leads to some intelligent computing centers being overloaded for a long time while the computing power resources of other intelligent computing centers are idle, exacerbating the uneven distribution of computing power.

[0008] In summary, in the prior art, when users rely on the computing power resources of intelligent computing centers to execute tasks, they still face bottlenecks in computing power resource scheduling and high economic costs, which restricts the popularization and application of inclusive computing power. Summary of the Invention

[0009] The present invention provides a multi-region computing power scheduling method and device for an inclusive computing power intelligent computing center to solve the technical problems in the prior art that when users rely on the computing power resources of the intelligent computing center to execute tasks, they still face bottlenecks in computing power resource scheduling and high economic costs, which restricts the popularization and application of inclusive computing power.

[0010] To solve the above technical problems, the present invention is implemented as follows:

[0011] In a first aspect, the present invention provides a multi-region computing power scheduling method for an inclusive computing power intelligent computing center, the method comprising:

[0012] Step S1: Determine a task to be migrated, where the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center;

[0013] Step S2: According to the task to be migrated, determine at least one optional intelligent computing center among each alternative intelligent computing center, where each alternative intelligent computing center is located in a different preset region;

[0014] Step S3: According to the task to be migrated, determine whether there is a target intelligent computing center among the at least one optional intelligent computing center;

[0015] Among them, the step S3 includes:

[0016] Step S31: Calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center, where the migration cost is determined according to the data volume to be migrated of the task to be migrated and the traffic cost of the data volume to be migrated;

[0017] Step S32: Calculate the first continued execution cost of the task to be migrated on the source intelligent computing center, where the first continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price of the preset region where the source intelligent computing center is located;

[0018] Step S33: Calculate the second continued execution cost of the task to be migrated on each optional intelligent computing center respectively, where the second continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price of the corresponding preset region where the optional intelligent computing center is located;

[0019] Step S34: Determine the minimum second continued execution cost among the second continued execution costs;

[0020] Step S35: If the first continued execution cost is greater than the sum of the minimum second continued execution cost, the migration cost, and a preset threshold, it is determined that there is the target intelligent computing center; wherein, the target intelligent computing center is the optional intelligent computing center corresponding to the minimum second continued execution cost.

[0021] Step S36: If the first continued execution cost is less than or equal to the sum of the minimum second continued execution cost, the migration cost, and the preset threshold, it is determined that there is no target intelligent computing center.

[0022] Step S4: If there is the target intelligent computing center, migrate the current intermediate operation result of the task to be migrated to the target intelligent computing center.

[0023] Step S5: Based on the computing power of the target intelligent computing center and the intermediate operation result, continue to run the task to be migrated.

[0024] Optionally, after the step S3, the method further includes:

[0025] Step S6: If there is no target intelligent computing center, continue to run the task to be migrated on the source intelligent computing center, and repeat the step S3 every first preset time period until the task to be migrated finishes running on the source intelligent computing center, or until it is determined that there is the target intelligent computing center.

[0026] Optionally, the step S2 includes:

[0027] Step S21: Determine the requirement label of the task to be migrated, where the requirement label includes at least one of the following: GPU model, data volume to be migrated.

[0028] Step S22: According to the requirement label and the estimated computing power situation of each alternative intelligent computing center in the next second preset time period, determine the optional intelligent computing center, where the estimated computing power situation includes at least one of the following: GPU model available for use in the next second preset time period, GPU utilization rate in the next second preset time period.

[0029] Optionally, the step S5 includes:

[0030] Step S51: Determine the task content of the task to be migrated, where the task content includes: training script, inference code, and container image.

[0031] Step S52: On the target intelligent computing center, continue to run the task to be migrated based on the intermediate operation result and the task content.

[0032] Second aspect, the present invention provides a multi-region computing power scheduling device for an inclusive computing power intelligent computing center, and the device includes:

[0033] A determination module, configured to determine a task to be migrated, where the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center;

[0034] An execution module, configured to determine at least one optional intelligent computing center among each alternative intelligent computing center according to the task to be migrated, where each alternative intelligent computing center is located in a different preset region;

[0035] According to the task to be migrated, determine whether there is a target intelligent computing center among the at least one optional intelligent computing center;

[0036] If there is the target intelligent computing center, migrate the current intermediate operation result of the task to be migrated to the target intelligent computing center;

[0037] Based on the computing power of the target intelligent computing center and the intermediate operation result, continue to run the task to be migrated;

[0038] Wherein, the execution module is further configured to calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center, where the migration cost is determined according to the data volume to be migrated of the task to be migrated and the traffic cost of the data volume to be migrated;

[0039] Calculate the first continued execution cost of the task to be migrated on the source intelligent computing center, where the first continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price of the preset region where the source intelligent computing center is located;

[0040] Calculate the second continued execution cost of the task to be migrated on each optional intelligent computing center respectively, where the second continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price of the corresponding optional intelligent computing center located in the preset region;

[0041] Among the second continued execution costs, determine the minimum second continued execution cost;

[0042] If the first continued execution cost is greater than the sum of the minimum second continued execution cost, the migration cost and a preset threshold, determine that there is the target intelligent computing center; wherein, the target intelligent computing center is the optional intelligent computing center corresponding to the minimum second continued execution cost;

[0043] If the first continue execution cost is less than or equal to the sum of the minimum second continue execution cost, the migration cost, and the preset threshold, it is determined that there is no such target intelligent computing center.

[0044] Optionally, the execution module is further configured to, after determining whether there is a target intelligent computing center in the optional intelligent computing centers according to the task to be migrated, if there is no such target intelligent computing center, continue to run the task to be migrated on the source intelligent computing center, and repeat the step of determining whether there is a target intelligent computing center in the optional intelligent computing centers according to the task to be migrated every first preset time period until the task to be migrated is completed on the source intelligent computing center or until it is determined that there is a target intelligent computing center.

[0045] Optionally, the execution module is further configured to determine the requirement tags of the task to be migrated, where the requirement tags include at least one of the following: GPU model, data volume to be migrated.

[0046] According to the requirement tags and the estimated computing power of each alternative intelligent computing center in the next second preset time period, determine the optional intelligent computing center, where the estimated computing power includes at least one of the following: GPU model available for use in the next second preset time period, GPU utilization rate in the next second preset time period.

[0047] Optionally, the execution module is further configured to determine the task content of the task to be migrated, where the task content includes: training script, inference code, and container image; and continue to run the task to be migrated on the target intelligent computing center based on the intermediate running result and the task content.

[0048] In a third aspect, the present invention provides a server, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of a multi-region computing power scheduling method for an intelligent computing center for inclusive computing power as described in the first aspect above are implemented.

[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the steps of a multi-region computing power scheduling method for an intelligent computing center for inclusive computing power as described in the first aspect above are implemented.

[0050] In a fifth aspect, the present invention provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of a multi-region computing power scheduling method for an intelligent computing center for inclusive computing power as described in the first aspect above are implemented.

[0051] In the present invention, first, screening is carried out among the alternative intelligent computing centers in different regions to determine the optional intelligent computing centers. By quantifying the migration cost (determined according to the amount of data to be migrated and the traffic cost) and the continued execution cost of the optional intelligent computing centers (determined according to the remaining duration of task execution and the estimated electricity price in the region), and comparing it with the continued execution cost of the source intelligent computing center. When the continued execution cost of the source intelligent computing center is higher than the sum of the minimum execution cost, migration cost and preset threshold of the target intelligent computing center, the intermediate state of the task is migrated to the target intelligent computing center with the optimal electricity price for continued operation.

[0052] Thus, through the intelligent task migration mechanism, the efficient scheduling and optimal allocation of computing power resources are realized. On the premise of ensuring the smooth execution of tasks, the economic cost of users using the computing power resources of the intelligent computing center is significantly reduced, the cross-regional sharing and collaboration of computing power resources are promoted, and the wide application of inclusive computing power is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0054] Figure 1 is a flowchart of a multi-region computing power scheduling method for an inclusive computing power intelligent computing center provided by the present invention;

[0055] Figure 2 is a flowchart of a multi-region computing power scheduling method for an inclusive computing power intelligent computing center provided by the present invention;

[0056] Figure 3 is a structural block diagram of a multi-region computing power scheduling device for an inclusive computing power intelligent computing center provided by the present invention;

[0057] Figure 4 is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] First, the technical terms related to the present invention will be briefly described below.

[0060] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0061] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.

[0062] The "carrying capacity" (Network Power, NP) referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive indicator for measuring network transmission scheduling ability.

[0063] The "storage power" (Storage Power, SP) referred to in the present invention means: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0064] The "computing power infrastructure" referred to in the present invention means: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.

[0065] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0066] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0067] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0068] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.

[0069] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0070] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0071] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0072] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0073] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0074] The "supercomputing center" described in the present invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0075] The "computing power resources" described in the present invention refer to the technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0076] The "inclusive computing power" described in the present invention refers to providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost based on the requirements of equal opportunity and the principle of commercial sustainability.

[0077] Figure 1 There is shown a multi-region computing power scheduling method for an inclusive computing power intelligent computing center provided according to the present invention, as Figure 1 shown, the method includes:

[0078] Step S1: Determine the task to be migrated;

[0079] Among them, the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center;

[0080] Step S2: According to the task to be migrated, determine at least one optional intelligent computing center among each alternative intelligent computing center;

[0081] Among them, each alternative intelligent computing center is located in a different preset area;

[0082] Step S3: According to the task to be migrated, determine whether there is a target intelligent computing center among at least one optional intelligent computing center;

[0083] Step S4: If there is a target intelligent computing center, migrate the current intermediate running result of the task to be migrated to the target intelligent computing center;

[0084] Step S5: Based on the computing power of the target intelligent computing center and the intermediate running result, continue to run the task to be migrated.

[0085] In step S1, it is first necessary to determine the computing tasks to be migrated (such as model training tasks). These tasks usually rely on the powerful computing power of the intelligent computing center and are currently being executed on the source intelligent computing center. By clarifying the tasks to be migrated, it is possible to lay a foundation for subsequent steps, ensuring the effectiveness and necessity of the migration process, and thus avoiding unnecessary resource waste.

[0086] It should be noted that the user can submit computing power usage tasks (tasks to be migrated) through the Web (World Wide Web) interface and / or RESTful API (Representational State Transfer Application Programming Interface), and input the following parameters: computing tasks and requirement tags. The computing tasks include: training scripts, inference codes, and container images (i.e., the environment for code execution). The requirement tags include: GPU models (such as A100 and / or V100) and the size of the data volume. This task to be migrated is executed on the source intelligent computing center.

[0087] In step S2, at least one optional intelligent computing center needs to be selected from multiple alternative intelligent computing centers according to the characteristics of the task to be migrated. These alternative intelligent computing centers are located in different preset regions (such as the East China region and the North China region) to achieve load balancing and resource optimization between different geographical locations.

[0088] In a possible implementation, step S2 includes: step S21: determining the requirement tags of the task to be migrated, where the requirement tags include at least one of the following: GPU model, data volume to be migrated; step S22: determining the optional intelligent computing center according to the requirement tags and the estimated computing power of each alternative intelligent computing center within the second preset time period in the future. The estimated computing power includes at least one of the following: the GPU models available for use within the second preset time period in the future, the GPU utilization rate within the second preset time period in the future (it should be noted that the GPU utilization rate can be determined every 30 seconds).

[0089] It should be noted that, first, the characteristics of the task to be migrated need to be analyzed to extract its requirement tags. Requirement tags are key indicators that describe the computing resource requirements of the task, including at least one of the following: GPU model, that is, the task to be migrated requires a specific model of GPU to meet its computing needs, especially for tasks such as deep learning and graphics processing; the amount of data to be migrated, which affects the storage and bandwidth resources required for the task to be migrated, and thus affects the migration and execution efficiency of the task. After determining the requirement tags of the task to be migrated, the optional intelligent computing centers will be selected based on these tags and the estimated computing power of each alternative intelligent computing center within the second preset time period in the future. Thus, by combining the requirement tags and the estimated computing power, the optional intelligent computing centers can be more accurately and preliminarily selected, facilitating the subsequent determination of the target intelligent computing center from them.

[0090] Exemplarily, the GPU model required by the task to be migrated is A100, and the estimated computing power of an intelligent computing center in the East China region within the second preset time period (for example: within 2 hours) does not match it. For example, there is no GPU of model A100, so it cannot be used as an optional intelligent computing center; if the amount of data to be migrated for the task to be migrated is 500GB, and the estimated computing power of an intelligent computing center in the North China region within the second preset time period does not match it. For example, the utilization rate of the GPU is too high and the number of GPUs is small, so it cannot be used as an optional intelligent computing center.

[0091] In step S3, after confirming at least one optional intelligent computing center, it will be further checked whether there is a suitable target intelligent computing center among these intelligent computing centers. The target center needs to have sufficient computing power and resources to receive the task to be migrated and be able to continue to execute the task.

[0092] In step S4, if the target intelligent computing center is determined, the current intermediate running result of the task to be migrated needs to be migrated to this target intelligent computing center. By migrating the intermediate running result, the continuity of the task can be ensured, avoiding recalculation and data loss, thus saving time and computing power resources and improving the overall execution efficiency of the task.

[0093] In step S5, after the target intelligent computing center receives the intermediate running result, it will continue to execute the task to be migrated using the computing power of this target intelligent computing center. This process can be dynamically adjusted according to the needs of the task to ensure the best computing performance. Thus, by continuing to run the task to be migrated in the target intelligent computing center, not only can the computing power resources of the target intelligent computing center be fully utilized, but also the flexible scheduling and optimal configuration of computing resources can be realized, thereby reducing the economic cost and promoting the application of inclusive computing power.

[0094] In a possible implementation, step S5 includes: Step S51: Determine the task content of the task to be migrated, where the task content includes: training scripts, inference code, and container images; Step S52: On the target intelligent computing center, continue to run the task to be migrated based on the intermediate operation result and the task content.

[0095] It should be noted that the intermediate operation result can be saved in the area with high cost (the area where the source intelligent computing center is located), and the intermediate operation result is copied and migrated. The container computing power resources are started in the area with low cost (the area where the target intelligent computing center is located) to load the intermediate operation result and continue to run the task to be migrated.

[0096] In a possible implementation, as Figure 2 shown, step S3 includes:

[0097] Step S31: Calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center;

[0098] Among them, the migration cost is determined according to the data volume to be migrated of the task to be migrated and the traffic cost of the data volume to be migrated;

[0099] Step S32: Calculate the first continued execution cost of the task to be migrated on the source intelligent computing center;

[0100] Among them, the first continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price in the preset area where the source intelligent computing center is located;

[0101] Step S33: Calculate the second continued execution cost of the task to be migrated on each optional intelligent computing center respectively;

[0102] Among them, the second continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price in the corresponding preset area where the optional intelligent computing center is located;

[0103] Step S34: Determine the minimum second continued execution cost among the second continued execution costs;

[0104] Step S35: If the first continued execution cost is greater than the sum of the minimum second continued execution cost, the migration cost, and the preset threshold, it is determined that there is a target intelligent computing center;

[0105] Among them, the target intelligent computing center is the optional intelligent computing center corresponding to the minimum second continued execution cost;

[0106] Step S36: If the first continued execution cost is less than or equal to the sum of the minimum second continued execution cost, the migration cost, and the preset threshold, it is determined that there is no target intelligent computing center.

[0107] It should be noted that, first of all, it is necessary to calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center. The migration cost is mainly determined based on the amount of data to be migrated and the traffic cost of the task to be migrated. Next, calculate the first continued execution cost of the task to be migrated in the source intelligent computing center. This cost is determined by the remaining execution duration of the task to be migrated and the estimated electricity price in the region where the source intelligent computing center is located. This cost evaluates the economy of continuing to execute the task in the source intelligent computing center. Subsequently, the second continued execution cost of the task to be migrated on each optional intelligent computing center can be calculated respectively. The factors considered are also the remaining execution duration of the task to be migrated and the estimated electricity price in the region where the corresponding optional intelligent computing center is located. This process can compare the continued execution costs of different optional intelligent computing centers. After calculating the second continued execution costs of all optional intelligent computing centers, it is necessary to determine the minimum value among them, that is, determine the minimum second continued execution cost. This minimum second continued execution cost will be used as the basis for subsequent decisions. If the first continued execution cost is greater than the sum of the minimum second continued execution cost, the migration cost, and the preset threshold, it will be determined that there is a target intelligent computing center. If the first continued execution cost is less than or equal to the sum of the minimum second continued execution cost, the migration cost, and the preset threshold, it is judged that there is no suitable target intelligent computing center, and the migration will not be carried out temporarily, and the task to be migrated will still be executed on the source intelligent computing center.

[0108] Through this series of steps, it is possible to comprehensively evaluate the economy between the source intelligent computing center and the optional intelligent computing centers, so as to make a more reasonable migration decision. This method not only improves the utilization efficiency of computing power resources but also significantly reduces the economic cost of users using the computing power resources of the intelligent computing center, promotes the cross-regional sharing and collaboration of computing power resources, and realizes the wide application of inclusive computing power.

[0109] Figure 2 The method shown above can be expressed by the following formula:

[0110] Migration cost = Size of data to be migrated (G) × Traffic cost (G / yuan);

[0111] First continued execution cost = Remaining task execution duration × Estimated electricity price in the current region (hour / yuan);

[0112] Second continued execution cost = Remaining task execution duration × Estimated electricity price in other regions (hour / yuan);

[0113] The condition for triggering migration is: First continued execution cost > min(Second continued execution cost) + Migration cost + Migration cost anti-jitter threshold.

[0114] Among them, the estimated electricity price can be the average electricity price within the next two hours. Each optional intelligent computing center corresponds to a second continued execution cost. The migration cost prevention jitter threshold is equivalent to the above-mentioned preset threshold. By setting the migration cost prevention jitter threshold, frequent migrations caused by small cost fluctuations (such as electricity price fluctuations, etc.) can be avoided, thereby ensuring the stability of the task and the reliability of execution.

[0115] In a specific application scenario, assume that the type of task to be migrated is a deep learning model training task, the amount of data to be migrated is 50GB, the traffic cost is 0.1 yuan / GB, the remaining task execution duration is 20 hours, the electricity price in the region where the source intelligent computing center is located (East China) is 1.5 yuan / hour (estimated value for the next 2 hours), and the estimated electricity prices in the regions where the optional intelligent computing centers are located are as follows: North China region: electricity price 0.8 yuan / hour, South China region: electricity price 1.0 yuan / hour, and the migration cost prevention jitter threshold is 5 yuan.

[0116] Then the migration cost is: 50GB × 0.1 yuan / GB = 5 yuan, the first continued execution cost is: 20 hours × 1.5 yuan / hour = 30 yuan, the second continued execution cost in the North China region is: 20 hours × 0.8 yuan / hour = 16 yuan, the second continued execution cost in the South China region is: 20 hours × 1 yuan / hour = 20 yuan, and the minimum second continued execution cost is 16 yuan. The trigger condition = 30 yuan > (16 yuan + 5 yuan + 5 yuan) = 26 yuan, so the trigger condition is established, and the task to be migrated can be migrated from the source intelligent computing center (East China region) to the target intelligent computing center (North China region).

[0117] It should be noted that the migration cost not only includes the data traffic cost (such as the transmission cost per GB), but also includes the time cost. And the transmission rate directly affects the data migration time. A low transmission rate will lead to a long migration time, and thus an increase in the time cost. Therefore, the transmission rate can also be taken into account in the migration cost. If the transmission rate from the source intelligent computing center to a certain intelligent computing center is too low, then carefully consider whether to select this intelligent computing center as the target intelligent computing center to ensure that the scheduling decision is more comprehensive.

[0118] In a possible implementation manner, after step S3, the method further includes:

[0119] Step S6: If there is no target intelligent computing center, continue to run the task to be migrated on the source intelligent computing center, and repeat step S3 every first preset time period until the task to be migrated is completed on the source intelligent computing center, or until it is determined that there is a target intelligent computing center.

[0120] It should be noted that by continuing to run tasks in the source intelligent computing center, costly data migration can be avoided under unnecessary circumstances. Migration will only be carried out when the evaluation results show that migration is a more cost-effective option. Thus, the economic cost can be further effectively controlled. And the mechanism of regular evaluation can flexibly respond to changing environmental conditions. For example, if it is found in subsequent evaluations that the electricity price of an alternative intelligent computing center decreases or the network conditions improve, a migration decision can be made in a timely manner, thereby reducing the economic cost of users using computing power resources and realizing the wide application of inclusive computing power.

[0121] In the present invention, first, screening is carried out among alternative intelligent computing centers in different regions to determine optional intelligent computing centers. By quantifying the migration cost (determined according to the amount of data to be migrated and the traffic cost) and the subsequent execution cost of the optional intelligent computing center (determined according to the remaining duration of task execution and the estimated electricity price in the region), and comparing it with the continued execution cost of the source intelligent computing center. When the continued execution cost of the source intelligent computing center is higher than the sum of the minimum execution cost, migration cost and preset threshold of the target intelligent computing center, the intermediate state of the task is migrated to the target intelligent computing center with the optimal electricity price for continued operation.

[0122] Thus, through the intelligent task migration mechanism, the efficient scheduling and optimal allocation of computing power resources are achieved. On the premise of ensuring the smooth execution of tasks, the economic cost of users using the computing power resources of the intelligent computing center is significantly reduced. It promotes the cross-regional sharing and collaboration of computing power resources and realizes the wide application of inclusive computing power.

[0123] Figure 3 Fig. shows a multi-region computing power scheduling device for an inclusive computing power intelligent computing center according to the present invention, as Figure 3 shown. The device 30 includes:

[0124] A determination module 301, configured to determine a task to be migrated, where the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center;

[0125] An execution module 302, configured to determine at least one optional intelligent computing center among each alternative intelligent computing center according to the task to be migrated, where each alternative intelligent computing center is located in a different preset region;

[0126] According to the task to be migrated, determine whether there is a target intelligent computing center among at least one optional intelligent computing center;

[0127] If there is a target intelligent computing center, migrate the current intermediate operation result of the task to be migrated to the target intelligent computing center;

[0128] Based on the computing power and intermediate operation results of the target intelligent computing center, continue to run the task to be migrated;

[0129] Among them, the execution module 302 is also used to calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center, where the migration cost is determined according to the data volume to be migrated of the task to be migrated and the traffic cost of the data volume to be migrated;

[0130] Calculate the first continued execution cost of the task to be migrated on the source intelligent computing center, where the first continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price in the preset area where the source intelligent computing center is located;

[0131] Calculate the second continued execution cost of the task to be migrated on each optional intelligent computing center respectively, where the second continued execution cost is determined by the remaining task execution duration of the task to be migrated and the estimated electricity price in the corresponding preset area where the optional intelligent computing center is located;

[0132] Among the second continued execution costs, determine the minimum second continued execution cost;

[0133] If the first continued execution cost is greater than the sum of the minimum second continued execution cost, the migration cost and the preset threshold, it is determined that there is a target intelligent computing center; among them, the target intelligent computing center is the optional intelligent computing center corresponding to the minimum second continued execution cost;

[0134] If the first continued execution cost is less than or equal to the sum of the minimum second continued execution cost, the migration cost and the preset threshold, it is determined that there is no target intelligent computing center.

[0135] In a possible implementation manner, after the execution module 302 determines whether there is a target intelligent computing center in the optional intelligent computing centers according to the task to be migrated, if there is no target intelligent computing center, it continues to run the task to be migrated on the source intelligent computing center, and repeats the step of determining whether there is a target intelligent computing center in the optional intelligent computing centers according to the task to be migrated every first preset time period until the task to be migrated is completed on the source intelligent computing center, or until it is determined that there is a target intelligent computing center.

[0136] In a possible implementation manner, the execution module 302 is also used to determine the requirement label of the task to be migrated, where the requirement label includes at least one of the following: GPU model, data volume to be migrated;

[0137] According to the requirement tags and the estimated computing power of each alternative intelligent computing center in the second preset time period in the future, determine the optional intelligent computing centers, where the estimated computing power includes at least one of the following: the GPU models available for use in the second preset time period in the future, and the GPU utilization rate in the second preset time period in the future.

[0138] In a possible implementation manner, the execution module 302 is further configured to determine the task content of the task to be migrated, where the task content includes: training scripts, inference codes, and container images; and continue to run the task to be migrated on the target intelligent computing center based on the intermediate operation results and the task content.

[0139] In the present invention, first, screen among the alternative intelligent computing centers in different regions to determine the optional intelligent computing centers. By quantifying the migration cost (determined according to the data volume to be migrated and the traffic cost) and the subsequent execution cost of the optional intelligent computing centers (determined according to the remaining duration of task execution and the estimated electricity price in the region), and comparing it with the continued execution cost of the source intelligent computing center. When the continued execution cost of the source intelligent computing center is higher than the sum of the lowest execution cost, migration cost, and preset threshold of the target intelligent computing center, then migrate the intermediate state of the task to the target intelligent computing center with the optimal electricity price to continue running.

[0140] Thus, through the intelligent task migration mechanism, the efficient scheduling and optimal allocation of computing power resources are realized. On the premise of ensuring the smooth execution of tasks, the economic cost of users using the computing power resources of the intelligent computing center is significantly reduced. It promotes the cross-regional sharing and collaboration of computing power resources and realizes the wide application of inclusive computing power.

[0141] Please refer to Figure 4 , the present invention also provides an electronic device 40, including a processor 401, a memory 402, and a computer program stored on the memory 402 and executable on the processor 401. When the computer program is executed by the processor 401, it implements the steps of the above-mentioned multi-region computing power scheduling method for an inclusive computing power intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0142] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned multi-region computing power scheduling method for an inclusive computing power intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0143] The present invention also provides a computer program product, including computer instructions which, when executed by a processor, implement the steps of the above-mentioned multi-region computing power scheduling method for the inclusive computing power intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0144] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in the present invention.

[0146] The present invention has been described above in conjunction with the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.

Claims

1. A multi-region computing power scheduling method for inclusive computing power intelligent computing center, characterized in that: The method comprises: Step S1: determining a task to be migrated, wherein the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center; Step S2: determining at least one optional intelligent computing center from among the candidate intelligent computing centers according to the task to be migrated, wherein each candidate intelligent computing center is located in a different preset area; Step S3: determining whether there is a target intelligent computing center in the at least one optional intelligent computing center according to the task to be migrated; Wherein, the step S3 comprises: Step S31: Calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center, wherein the migration cost is determined according to the amount of data to be migrated of the task to be migrated and the traffic cost of the amount of data to be migrated; Step S32: Calculate a first continuing execution cost of the task to be migrated on the source intelligent computing center, wherein the first continuing execution cost is determined by the remaining task execution time of the task to be migrated and the estimated electricity price of the preset area where the source intelligent computing center is located; Step S33: calculating the second continuing execution cost of the task to be migrated on each optional intelligent computing center respectively, wherein the second continuing execution cost is determined by the remaining task execution time of the task to be migrated and the estimated electricity price of the preset area where the corresponding optional intelligent computing center is located; Step S34: determining a minimum second continuing execution cost among the second continuing execution costs; Step S35: If the first continuing execution cost is greater than the sum of the minimum second continuing execution cost, the migration cost and a preset threshold, it is determined that the target intelligent computing center exists; wherein the target intelligent computing center is an optional intelligent computing center corresponding to the minimum second continuing execution cost; Step S36: if the first continuing execution cost is less than or equal to the sum of the minimum second continuing execution cost, the migration cost and the preset threshold, determining that the target intelligent computing center does not exist; Step S4: if the target intelligent computing center exists, migrating the current intermediate running result of the task to be migrated to the target intelligent computing center; Step S5: Based on the computing power of the target intelligent computing center and the intermediate operation results, continue to run the task to be migrated.

2. The method according to claim 1, characterized in that After step S3, the method further includes: Step S6: If the target intelligent computing center does not exist, continue to run the task to be migrated on the source intelligent computing center, and repeat step S3 every first preset time period until the task to be migrated is completed on the source intelligent computing center, or until it is determined that the target intelligent computing center exists.

3. The method according to claim 1, characterized in that The step S2 comprises: Step S21: determining a requirement tag of the task to be migrated, wherein the requirement tag includes at least one of the following: GPU model, amount of data to be migrated; Step S22: Determine the optional intelligent computing center based on the demand label and the estimated computing power of each alternative intelligent computing center in the future second preset time period, wherein the estimated computing power includes at least one of the following: the GPU model available in the future second preset time period and the GPU utilization rate in the future second preset time period.

4. The method according to any one of claims 1 to 3, characterized in that The step S5 comprises: Step S51: determining the task content of the task to be migrated, wherein the task content includes: training script, reasoning code and container image; Step S52: On the target intelligent computing center, based on the intermediate running result and the task content, continue to run the task to be migrated.

5. A multi-region computing power scheduling device for inclusive computing power intelligent computing center, characterized in that: The device comprises: A determination module, used to determine a task to be migrated, wherein the task to be migrated needs to be executed based on the computing power of the intelligent computing center, and the task to be migrated is currently running on the source intelligent computing center; An execution module, configured to determine at least one optional intelligent computing center from among the candidate intelligent computing centers according to the task to be migrated, wherein each candidate intelligent computing center is located in a different preset area; According to the task to be migrated, determining whether there is a target intelligent computing center in the at least one optional intelligent computing center; If the target intelligent computing center exists, migrating the current intermediate running result of the task to be migrated to the target intelligent computing center; Based on the computing power of the target intelligent computing center and the intermediate operation result, continue to run the task to be migrated; The execution module is further used to calculate the migration cost of migrating the task to be migrated from the source intelligent computing center to each optional intelligent computing center, wherein the migration cost is determined according to the amount of data to be migrated of the task to be migrated and the traffic cost of the amount of data to be migrated; Calculating a first continuing execution cost of the task to be migrated on the source intelligent computing center, wherein the first continuing execution cost is determined by a remaining task execution time of the task to be migrated and an estimated electricity price of a preset area where the source intelligent computing center is located; Calculate the second continuing execution cost of the task to be migrated on each optional intelligent computing center respectively, wherein the second continuing execution cost is determined by the remaining task execution time of the task to be migrated and the estimated electricity price of the preset area where the corresponding optional intelligent computing center is located; Among the second continuing execution costs, determining a minimum second continuing execution cost; If the first continuing execution cost is greater than the sum of the minimum second continuing execution cost, the migration cost and a preset threshold, it is determined that the target intelligent computing center exists; wherein the target intelligent computing center is an optional intelligent computing center corresponding to the minimum second continuing execution cost; If the first continuing execution cost is less than or equal to the sum of the minimum second continuing execution cost, the migration cost and the preset threshold, it is determined that the target intelligent computing center does not exist.

6. The device according to claim 5, characterized in that The execution module is further configured to, after determining whether there is a target intelligent computing center in the optional intelligent computing centers according to the task to be migrated, If the target intelligent computing center does not exist, the task to be migrated continues to run on the source intelligent computing center, and the step of determining whether there is a target intelligent computing center in the optional intelligent computing center based on the task to be migrated is repeated every first preset time period until the task to be migrated is completed in the source intelligent computing center, or until it is determined that the target intelligent computing center exists.

7. The device according to claim 5, characterized in that The execution module is further used to determine a requirement tag of the task to be migrated, wherein the requirement tag includes at least one of the following: GPU model and amount of data to be migrated; The optional intelligent computing center is determined based on the demand label and the estimated computing power conditions of each alternative intelligent computing center in the second preset time period in the future, wherein the estimated computing power conditions include at least one of the following: the GPU model available in the second preset time period in the future and the GPU utilization rate in the second preset time period in the future.

8. A server, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a multi-region computing power scheduling method for a universal computing power intelligent computing center are implemented as described in any one of claims 1 to 4.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a multi-region computing power scheduling method for a universal computing power intelligent computing center as described in any one of claims 1 to 4.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a multi-region computing power scheduling method for a universal computing power intelligent computing center as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Automatic network and route selection method for achieving multi-objective achievement of computing power network

    CN115632939A

  • Calculation power resource scheduling method, device, system and equipment and storage medium

    CN116643873A