Intelligent computing center model training task allocation method and device capable of providing computing power across

By finding and assigning tasks to the second intelligent computing center to provide computing power for advanced users in the intelligent computing center in the intelligent computing center, the problem of inefficient training of advanced users is solved, and efficient and stable model training is achieved.

CN120371538APending Publication Date: 2025-07-25DATACANVAS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510859922.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the intelligent computing center, the model training tasks of advanced users are easily blocked by low-priority tasks, resulting in low training efficiency. How to improve the efficiency of model training tasks of advanced users has become an urgent problem.

Method used

By obtaining user identification information and model training tasks, we can judge the computing power resources of the local intelligent computing center. If it is insufficient, find a second intelligent computing center that provides computing power services for advanced users, and assign tasks to the center to perform to ensure that advanced users receive sufficient computing power support.

Benefits of technology

When the local computing power resources of advanced users are tight, the second intelligent computing center that provides computing power for advanced users performs tasks, which improves the stability and speed of model training and significantly improves the training efficiency of advanced users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371538A_ABST
    Figure CN120371538A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent computing center model training task allocation method and device for cross-providing computing power, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the steps: obtaining the identification information of a user and a model training task of the user; judging whether the computing power resource of the first intelligent computing center is greater than a preset threshold; under the condition that the computing power resource of the first intelligent computing center is smaller than a preset threshold value and the identification information of the user indicates that the user is an advanced user, a second intelligent computing center is searched, and the second intelligent computing center is specially used for providing computing power service for the advanced user; and under the condition that the computing power resource of the second intelligent computing center is greater than a preset threshold value, distributing the model training task of the user to the second intelligent computing center so as to execute the model training task of the user by utilizing the second intelligent computing center. According to the method, the execution stability and speed of the model training task of the senior user are improved, so that the model training efficiency of the senior user is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for allocating model training tasks of an intelligent computing center that provides computing power across different levels. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to achieve the output of the target result by processing information data. It is a new type of productive force that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] In an intelligent computing center, different users have different requirements for model training tasks. Ordinary users pay lower fees and usually purchase basic versions, free versions or low-cost services. They don't mind that the model training process takes a long time and are more concerned about costs. Advanced users pay higher fees and usually purchase advanced versions, professional versions or customized services. They are very concerned about the efficiency of model training and hope to obtain results in a shorter time. If advanced users share the same computing power resources with general users, high-priority model training tasks may be blocked by low-priority model training tasks, resulting in lower training efficiency of high-priority model training tasks. Therefore, since the emergence of intelligent computing centers, how to improve the training efficiency of advanced users' model training tasks has become an urgent technical problem to be solved. Summary of the Invention

[0008] The present invention provides a method and device for allocating model training tasks across intelligent computing centers that provide computing power, to solve the problem of how to improve the training efficiency of model training tasks for advanced users.

[0009] To solve the above problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for allocating model training tasks across intelligent computing centers that provide computing power, including: Step S1, obtaining the identification information of the user and the model training task of the user; Step S2, determining whether the computing power resources of the first intelligent computing center are greater than a preset threshold, where the first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for ordinary users and advanced users; Step S3, in the case where the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is an advanced user, searching for a second intelligent computing center, where the second intelligent computing center is used to specifically provide computing power services for the advanced users; Step S4, in the case where the computing power resources of the second intelligent computing center are greater than the preset threshold, allocating the model training task of the user to the second intelligent computing center to execute the model training task of the user by using the second intelligent computing center.

[0010] In one embodiment, the method further includes: Step S5, in the case where the computing power resources of the second intelligent computing center are less than the preset threshold, searching for a third intelligent computing center, where the third intelligent computing center is deployed near the user, and the third intelligent computing center is used to provide computing power services for the ordinary users and the advanced users; Step S6, in the case where the computing power resources of the third intelligent computing center are greater than the preset threshold, allocating the model training task of the user to the third intelligent computing center to execute the model training task of the user by using the third intelligent computing center.

[0011] In one embodiment, the method further includes: Step S7, in the case where the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is an ordinary user, searching for a third intelligent computing center, where the third intelligent computing center is deployed near the user, and the third intelligent computing center is used to provide computing power services for the ordinary users and the advanced users; Step S8: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the third intelligent computing center to execute the user's model training task by using the third intelligent computing center.

[0012] In one embodiment, the method further includes: Step S9: When the computing power resources of the third intelligent computing center are less than the preset threshold, compare the computing power resources of the first intelligent computing center and the third intelligent computing center; Step S10: Allocate the user's model training task to the target intelligent computing center to execute the user's model training task by using the target intelligent computing center, where the target intelligent computing center is the intelligent computing center with the most computing power resources among the first intelligent computing center and the third intelligent computing center.

[0013] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, and the user is the premium user. Step S6 includes: Step S61: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the idle second GPU graphics card of the third intelligent computing center to execute the user's model training task by using the idle second GPU graphics card.

[0014] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, and the user is the ordinary user. Step S8 includes: Step S81: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the target GPU graphics card of the third intelligent computing center to execute the user's model training task by using the target GPU graphics card. Wherein, when there is an idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle second GPU graphics card; when there is no idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle first GPU graphics card.

[0015] In one embodiment, when the user is a premium user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the priority order of each user among the N users. The priority order of each user is determined based on the payment amount of each user and the required completion time of the model training task, where N is an integer greater than 1.

[0016] In one embodiment, when the user is a regular user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the order of the initiation time of the model training tasks of each user among the N users, where N is an integer greater than 1.

[0017] In a second aspect, the present invention further provides a device for allocating model training tasks across intelligent computing centers that provide computing power, including: A first acquisition module, configured to acquire the identification information of the user and the model training task of the user; A first judgment module, configured to judge whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for regular users and premium users; A first search module, configured to search for a second intelligent computing center when the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user. The second intelligent computing center is an intelligent computing center dedicated to providing computing power services for premium users; A first allocation module, configured to allocate the model training task of the user to the second intelligent computing center when the computing power resources of the second intelligent computing center are greater than the preset threshold, so as to use the second intelligent computing center to execute the model training task of the user.

[0018] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for allocating model training tasks across intelligent computing centers that provide computing power as described in the first aspect above.

[0019] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for allocating model training tasks across intelligent computing centers that provide computing power as described in the first aspect above.

[0020] Fifth aspect, the present invention further provides a computer program product, including computer instructions, which when executed by a processor, implement the steps in the method for allocating model training tasks across intelligent computing centers providing computing power as described in the first aspect above.

[0021] In the present invention, obtain the identification information of the user and the model training task of the user; determine whether the computing power resources of the first intelligent computing center are greater than a preset threshold, the first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for ordinary users and premium users; in the case where the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user, search for a second intelligent computing center, the second intelligent computing center is used to specifically provide computing power services for the premium user; in the case where the computing power resources of the second intelligent computing center are greater than the preset threshold, allocate the model training task of the user to the second intelligent computing center, so as to use the second intelligent computing center to execute the model training task of the user.

[0022] In this method, when the computing power resources of the intelligent computing center local to the premium user are insufficient, the second intelligent computing center that specifically provides computing power services for the premium user executes the model training task of the premium user, so that the premium user can still obtain sufficient computing power support when the computing power resources are tense. At the same time, the specialization of the second intelligent computing center reduces the competition for computing power resources, improves the stability and speed of the execution of the model training task of the premium user, and thus significantly improves the model training efficiency of the premium user. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for the description of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 is a flowchart of a method for allocating model training tasks across intelligent computing centers providing computing power provided by the present invention; Figure 2 is one of the schematic diagrams of an intelligent computing center provided by the present invention; Figure 3 is another schematic diagram of an intelligent computing center provided by the present invention; Figure 4 is a schematic diagram of the scheduling of the model training task of an ordinary user provided by the present invention; Figure 5It is a structural diagram of an intelligent computing center model training task allocation device for providing computing power according to the present invention; Figure 6 It is a structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0025] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0026] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0027] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and output results, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .

[0028] The "carrying capacity" (Network Power, NP) described in the present invention refers to: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.

[0029] The "Storage Power" (SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and internal storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0030] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.

[0031] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0032] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0033] The "general computing power" described in the present invention refers to: the computing ability provided by servers based on central processing unit (CPU) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0034] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform based on the large-scale deployment of dedicated chips such as GPU (Graphics Processing Unit), Field Programmable Gate Array (FPGA), and Application Specific Integrated Circuit (ASIC), such as natural language processing, machine vision, and so on.

[0035] The "super computing power" described in the present invention mainly refers to: the computing ability provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system, and is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0036] The "Intelligent Computing Center" as described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0037] The "Intelligent Computing Center" as described in the present invention includes, but is not limited to, the "Intelligent Computing Center".

[0038] The "Intelligent Computing Center" as described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0039] The "Computing Power Center" as described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0040] The "Supercomputing Center" as described in the present invention refers to, that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0041] The "Computing Power Resources" as described in the present invention refers to technologies and facilities required for the development of the digital society and having information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0042] In the prior art, in an intelligent computing center, different users have different requirements for model training tasks. Ordinary users pay lower fees and usually purchase basic versions, free versions, or low-cost services. They don't mind that the model training process takes a long time and are more concerned about costs. Advanced users pay higher fees and usually purchase premium versions, professional versions, or customized services. They are very concerned about the efficiency of model training and hope to obtain results in a shorter time. If advanced users share the same computing power resources with ordinary users, high-priority model training tasks may be blocked by low-priority model training tasks, resulting in lower training efficiency for high-priority model training tasks. Therefore, since the emergence of intelligent computing centers, how to improve the training efficiency of advanced users' model training tasks has become a technical problem to be solved urgently. To improve the training efficiency of advanced users' model training tasks, in the present invention, the identification information of the user and the user's model training task are obtained; it is determined whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally to the user and is used to provide computing power services for ordinary users and advanced users; in the case where the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is an advanced user, a second intelligent computing center is searched for. The second intelligent computing center is used to provide computing power services specifically for the advanced users; in the case where the computing power resources of the second intelligent computing center are greater than the preset threshold, the user's model training task is assigned to the second intelligent computing center to execute the user's model training task by using the second intelligent computing center. This method, when the computing power resources of the intelligent computing center local to the advanced user are insufficient, has the second intelligent computing center that provides computing power services specifically for the advanced users execute the model training tasks of the advanced users, so that advanced users can still obtain sufficient computing power support when computing power resources are scarce. At the same time, the specialization of the second intelligent computing center reduces the competition for computing power resources and improves the stability and speed of the execution of the model training tasks of the advanced users, thus significantly improving the model training efficiency of the advanced users.

[0043] Specifically, please refer to Figure 1 , Figure 1 which is a flowchart of a method for allocating model training tasks across intelligent computing centers that provide computing power provided by the present invention. As Figure 1 shown, it includes the following steps: Step S1, obtain the identification information of the user and the user's model training task; In this step, the identification information of the user can be the user's ID, which is used to identify the user's identity. For example, the user can be an advanced user or an ordinary user. The user's model training task refers to the specific task that the user hopes to execute. Subsequently, resource allocation and task scheduling can be performed based on the identification information of the user and the user's model training task.

[0044] Step S2: Determine whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally at the user's location and is used to provide computing power services for ordinary users and premium users. In this step, refer to Figure 2 , since the local intelligent computing center can reduce network latency, save bandwidth, and improve task execution efficiency, it is preferred to search for the intelligent computing center deployed locally at the user's location (such as the first intelligent computing center mentioned above). The first intelligent computing center can provide computing power services for both ordinary users and premium users, that is, the model training tasks of premium users and ordinary users can be executed in the first intelligent computing center. Further, determine whether the current computing power resources of the first intelligent computing center are sufficient to meet the training requirements of the model training task. The above preset threshold can be determined according to the training requirements of the model training task.

[0045] Step S3: When the computing power resources of the first intelligent computing center are less than the preset threshold and the user's identification information indicates that the user is a premium user, search for a second intelligent computing center, which is used to provide computing power services specifically for the premium user. In this step, refer to Figure 2 , when the computing power resources of the first intelligent computing center are less than the preset threshold, that is, the computing power resources of the first intelligent computing center are not sufficient to support the user's model training task, and the user's identification information indicates that the user is a premium user, the system will search for a second intelligent computing center. It should be noted that the second intelligent computing center is an intelligent computing center that provides computing power services specifically for premium users, that is, the model training tasks of ordinary users cannot be executed in the second intelligent computing center. The difference between the second intelligent computing center and the aforementioned first intelligent computing center is that the second intelligent computing center has more powerful hardware resources, such as a GPU cluster with better performance, and can provide higher computing performance, more flexible resource scheduling, and exclusive services for premium users to support more complex and high-load model training tasks, so as to meet the requirements of premium users for training time.

[0046] Step S4: When the computing power resources of the second intelligent computing center are greater than the preset threshold, allocate the user's model training task to the second intelligent computing center to execute the user's model training task by using the second intelligent computing center.

[0047] In this step, if the computing power resources of the second intelligent computing center are greater than the preset threshold, that is, when the computing power resources of the second intelligent computing center are sufficient to meet the training requirements of the user's model training task, the system will allocate the user's model training task to the second intelligent computing center, that is, send the task parameters, training data, etc. of the user's model training task to the second intelligent computing center, so as to use the second intelligent computing center to execute the user's model training task.

[0048] In the above embodiment, the identification information of the user and the user's model training task are obtained; it is judged whether the computing power resources of the first intelligent computing center are greater than the preset threshold, the first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for ordinary users and premium users; in the case where the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user, a second intelligent computing center is searched for, and the second intelligent computing center is used to provide computing power services specifically for the premium user; in the case where the computing power resources of the second intelligent computing center are greater than the preset threshold, the user's model training task is allocated to the second intelligent computing center to use the second intelligent computing center to execute the user's model training task.

[0049] In this embodiment, when the computing power resources of the intelligent computing center local to the premium user are insufficient, the second intelligent computing center that provides computing power services specifically for the premium user executes the model training task of the premium user, so that the premium user can still obtain sufficient computing power support when the computing power resources are tight. At the same time, the specialization of the second intelligent computing center reduces the competition of computing power resources and improves the stability and speed of the execution of the model training task of the premium user, thus significantly improving the model training efficiency of the premium user.

[0050] In one embodiment, after step S4, the method further includes: Step S5, in the case where the computing power resources of the second intelligent computing center are less than the preset threshold, search for a third intelligent computing center, the third intelligent computing center is deployed near the user, and the third intelligent computing center is used to provide computing power services for the ordinary user and the premium user; Step S6, in the case where the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the third intelligent computing center to use the third intelligent computing center to execute the user's model training task.

[0051] In the above embodiment, continue to refer to Figure 2, when the user is a premium user and the computing power resources of the intelligent computing center that specifically provides computing power services for the user are insufficient, the system will further search for an intelligent computing center deployed near the user (such as the third intelligent computing center mentioned above). The third intelligent computing center can provide computing power services for both premium users and ordinary users. When the computing power resources of the third intelligent computing center are greater than the preset threshold, that is, the computing power resources of the third intelligent computing center are sufficient to meet the task requirements of the user's model training task, the system will allocate the user's model training task to the third intelligent computing center, that is, send the task parameters, training data, etc. of the user's model training task to the third intelligent computing center, so as to use the third intelligent computing center to execute the user's model training task.

[0052] In this embodiment, through multi-level computing power resource scheduling, the system ensures that premium users can still obtain available computing power when the computing power resources at multiple levels are insufficient, improving the continuity and reliability of the execution of the premium user's model training task. As an intelligent computing center close to the user, the third intelligent computing center has low latency. Compared with other intelligent computing centers farther away from the user, using the third intelligent computing center to execute the user's model training task can improve the efficiency of the user's model training task.

[0053] In one embodiment, the method further includes: Step S7, when the computing power resources of the first intelligent computing center are less than the preset threshold and the user identification information indicates that the user is an ordinary user, search for a third intelligent computing center. The third intelligent computing center is deployed near the user and is used to provide computing power services for the ordinary user and the premium user; Step S8, when the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the third intelligent computing center to use the third intelligent computing center to execute the user's model training task.

[0054] In the above embodiment, whether the user is a premium user or an ordinary user, when searching for an intelligent computing center to execute their model training task, they will first give priority to considering the intelligent computing center deployed locally (such as the first intelligent computing center mentioned above). However, for premium users, when the computing power resources of the local intelligent computing center are insufficient, the system will search for an intelligent computing center that specifically provides computing power services for premium users (such as the second intelligent computing center mentioned above); for ordinary users, when the computing power resources of the local intelligent computing center are insufficient, the system will further search for an intelligent computing center deployed near the user (such as the third intelligent computing center mentioned above).

[0055] Since high - level users have a higher willingness to pay and higher requirements for model training time, an intelligent computing center dedicated to providing computing power services for high - level users can be pre - created to ensure the training efficiency of high - level users' model training tasks. For ordinary users, since their willingness to pay is lower and their requirements for model training time are lower, it is sufficient to find an intelligent computing center to execute the model training task for ordinary users according to the general resource - level scheduling.

[0056] In this embodiment, through multi - level resource scheduling and priority management, both the exclusive resources of high - level users are guaranteed and a flexible alternative solution is provided for ordinary users, thereby improving the stability, efficiency, and user experience of each user's model training task.

[0057] In one embodiment, the method further includes: Step S9: When the computing power resources of the third intelligent computing center are less than the preset threshold, compare the computing power resources of the first intelligent computing center and the third intelligent computing center; Step S10: Allocate the user's model training task to the target intelligent computing center to use the target intelligent computing center to execute the user's model training task, where the target intelligent computing center is the intelligent computing center with the most computing power resources among the first intelligent computing center and the third intelligent computing center.

[0058] In the above - mentioned embodiment, regardless of whether the user is a high - level user or a low - level user, when the computing power resources of the third intelligent computing center are insufficient, the system will compare the computing power resources of the first intelligent computing center and the third intelligent computing center. The system will allocate the user's model training task to the intelligent computing center with the most computing power resources (i.e., the one with better resources among the first or the third center) to execute the model training task.

[0059] In this embodiment, through dynamic resource comparison and intelligent allocation, the dual optimization of resource utilization efficiency and task execution efficiency is achieved. First, in step S9, when the resources of the third center are insufficient, the computing power resources of the first and the third centers are compared to avoid blind allocation and ensure that the task selects the optimal path among the available resources. In step S10, according to the resource comparison result, the task is allocated to the center with the strongest computing power, maximizing the use of existing resources, reducing waiting time, and improving training efficiency.

[0060] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, the user is the high - level user, and step S6 includes: Step S61: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the idle second GPU graphics card of the third intelligent computing center to execute the user's model training task using the idle second GPU graphics card.

[0061] In the above embodiment, refer to Figure 3 , the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card. The first GPU graphics card can be an H800, and the second GPU graphics card can be an H100. The computing power of the second GPU graphics card is stronger than that of the first GPU graphics card. The GPU graphics cards in the black squares are occupied GPU graphics cards, and the GPU graphics cards in the white squares are idle GPU graphics cards. If the user is an advanced user, their model training task needs to be run on the second GPU graphics card because the second GPU graphics card has stronger computing power and can significantly improve the model training speed to meet the efficiency requirements of the advanced user for the model training task.

[0062] In this embodiment, allocating the model training task of the advanced user to a high-performance GPU can significantly improve the training efficiency of their model training task and meet their high demand for computing power.

[0063] It should be noted that continue to refer to Figure 3 , the first intelligent computing center also includes a first GPU graphics card and a second GPU graphics card. If the model training task of the advanced user is executed on the first intelligent computing center, then the model training task of the advanced user will be allocated to the second GPU graphics card of the first intelligent computing center; the second intelligent computing center only includes a second GPU graphics card and is dedicated to serving the model training tasks of advanced users.

[0064] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, the user is an ordinary user, and step S8 includes: Step S81: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the target GPU graphics card of the third intelligent computing center to execute the user's model training task using the target GPU graphics card. Wherein, when there is an idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle second GPU graphics card; when there is no idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle first GPU graphics card.

[0065] In the above embodiment, continue to refer to Figure 3, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card. The first GPU graphics card can be an H800, and the second GPU graphics card can be an H100. The computing power of the second GPU graphics card is stronger than that of the first GPU graphics card. If the user is an ordinary user, in principle, the model training tasks of ordinary users can only run on the idle first GPU graphics card. However, if the second GPU graphics card is idle, in order to make full use of the computing power resources of the second GPU graphics card, the model training tasks of ordinary users can also be run on the second GPU graphics card. However, it should be noted that if the computing power resources for the model training tasks of high-level users are insufficient at this time, ordinary users need to give up the resources of the second GPU graphics card to ensure that the model training tasks of high-level users can be executed smoothly. The model training tasks of ordinary users can be transferred to the idle first GPU graphics card for execution.

[0066] In this embodiment, ordinary users can share high-performance computing power when the second GPU graphics card is idle, improving the utilization rate of computing power resources. At the same time, high-level users can preferentially occupy the second GPU graphics card to ensure their high-efficiency requirements. When resource conflicts occur, ordinary users take the initiative to give way to ensure that key tasks are executed first, avoiding resource waste and maintaining system stability, taking into account both fairness and efficiency.

[0067] It should be noted that continue to refer to Figure 3 , the first intelligent computing center also includes a first GPU graphics card (H800) and a second GPU graphics card (H100). The execution rules of the model training tasks of ordinary users in the first intelligent computing center can refer to their execution rules in the third intelligent computing center, which will not be elaborated here.

[0068] In one embodiment, when the user is a high-level user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the priority order of each user among the N users. The priority order of each user is determined based on the payment amount of each user and the required completion time of the model training task, where N is an integer greater than 1.

[0069] In the above embodiment, when there are multiple high-level users and it is possible that the idle computing power resources of the third intelligent computing center are not sufficient to support the model training tasks of all high-level users, the system will sort the multiple high-level users according to the payment amount of each high-level user and the required completion time of the model training task. Exemplarily, the higher the fee paid by a high-level user, the higher the priority; if a high-level user hopes to complete the task as soon as possible, its priority will be higher.

[0070] The third intelligent computing center will schedule model training tasks according to the priority order of high - level users, ensuring that the model training tasks of high - priority users can obtain computing power resources first. For example, if high - level user A pays more or the task completion time is more urgent, their model training tasks will be executed first; while the model training tasks of low - priority users need to wait for resources to be idle.

[0071] In this embodiment, by combining the payment amount and the urgency of the task, the computing power resources of the third intelligent computing center are dynamically allocated to ensure that high - level users can obtain more efficient priority services in resource competition, and at the same time, the reasonable utilization of computing power resources can be improved.

[0072] It should be noted that when the model training tasks of multiple high - level users are executed in the first intelligent computing center or the second intelligent computing center, in the case where the computing power resources of the first intelligent computing center or the second intelligent computing center are not sufficient to meet the training requirements of all the model training tasks of high - level users, the model training tasks of multiple high - level users can be scheduled according to the priority order of multiple high - level users determined by the above rules, and the details will not be elaborated here.

[0073] In one embodiment, when the user is an ordinary user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the order of the start time of each user's model training task, where N is an integer greater than 1.

[0074] In the above - mentioned embodiment, refer to Figure 4 , when there are multiple ordinary users and it is possible that the idle computing power resources of the third intelligent computing center are not sufficient to support all the model training tasks of ordinary users, the third intelligent computing center will determine the execution order of their model training tasks according to the time sequence of each ordinary user submitting the model training task.

[0075] In this embodiment, since ordinary users have lower requirements for the urgency or priority of model training task execution, queuing in time sequence provides a fair and simple resource allocation method for ordinary users.

[0076] It should be noted that when the model training tasks of multiple ordinary users are executed in the first intelligent computing center, in the case where the computing power resources of the first intelligent computing center are not sufficient to meet the training requirements of all the model training tasks of ordinary users, the scheduling order of the model training tasks of multiple ordinary users determined by the above rules can be followed, and the details will not be elaborated here.

[0077] Please refer to Figure 5 , Figure 5 is the structural diagram of a device for allocating model training tasks across intelligent computing centers providing computing power according to the present invention. As Figure 5As shown, the intelligent computing center model training task allocation device 500 for providing computing power includes: A first acquisition module 501, configured to acquire the identification information of the user and the model training task of the user; A first judgment module 502, configured to judge whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for ordinary users and premium users; A first search module 503, configured to search for a second intelligent computing center when the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user. The second intelligent computing center is an intelligent computing center dedicated to providing computing power services for the premium user; A first allocation module 504, configured to allocate the model training task of the user to the second intelligent computing center when the computing power resources of the second intelligent computing center are greater than the preset threshold, so as to use the second intelligent computing center to execute the model training task of the user.

[0078] In one embodiment, the device further includes: A second search module, configured to search for a third intelligent computing center when the computing power resources of the second intelligent computing center are less than the preset threshold. The third intelligent computing center is deployed near the user, and the third intelligent computing center is used to provide computing power services for the ordinary user and the premium user; A second allocation module, configured to allocate the model training task of the user to the third intelligent computing center when the computing power resources of the third intelligent computing center are greater than the preset threshold, so as to use the third intelligent computing center to execute the model training task of the user.

[0079] In one embodiment, the device further includes: A third search module, configured to search for a third intelligent computing center when the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is an ordinary user. The third intelligent computing center is deployed near the user, and the third intelligent computing center is used to provide computing power services for the ordinary user and the premium user; A third allocation module, configured to allocate the model training task of the user to the third intelligent computing center when the computing power resources of the third intelligent computing center are greater than the preset threshold, so as to use the third intelligent computing center to execute the model training task of the user.

[0080] In one embodiment, the device further includes: A first comparison module, configured to compare the computing power resources of the first intelligent computing center and the third intelligent computing center when the computing power resources of the third intelligent computing center are less than the preset threshold. A fourth allocation module, configured to allocate the model training task of the user to a target intelligent computing center to execute the model training task of the user by using the target intelligent computing center, where the target intelligent computing center is the intelligent computing center with the most computing power resources among the first intelligent computing center and the third intelligent computing center.

[0081] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, the user is the premium user, and the second allocation module includes: A first allocation unit, configured to allocate the model training task of the user to the idle second GPU graphics card of the third intelligent computing center to execute the model training task of the user by using the idle second GPU graphics card when the computing power resources of the third intelligent computing center are greater than the preset threshold.

[0082] In one embodiment, the third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card, the user is the ordinary user, and the third allocation module includes: A second allocation unit, configured to allocate the model training task of the user to the target GPU graphics card of the third intelligent computing center to execute the model training task of the user by using the target GPU graphics card when the computing power resources of the third intelligent computing center are greater than the preset threshold, where, when there is an idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle second GPU graphics card; when there is no idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle first GPU graphics card.

[0083] In one embodiment, when the user is a premium user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the priority order of each user among the N users, and the priority order of each user is determined based on the payment amount of each user and the required completion time of the model training task, where N is an integer greater than 1.

[0084] In one embodiment, when the user is an ordinary user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the order of the start times of the model training tasks of each user among the N users, where N is an integer greater than 1.

[0085] The device for allocating model training tasks in an intelligent computing center across provided computing power according to the present invention can implement each process of the above-described method for allocating model training tasks in an intelligent computing center across provided computing power. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0086] It should be noted that the device for distributing computing power operation data of the intelligent computing center in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0087] The present invention also provides an electronic device. Refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 601, a processor 602, and a program or instruction running on the memory 601. When the program or instruction is executed by the processor 602, it can implement Figure 1 any step in the corresponding embodiment of the method for allocating model training tasks in an intelligent computing center across provided computing power and achieve the same beneficial effects, which will not be elaborated here.

[0088] Among them, the processor 602 can be a CPU, an ASIC, an FPGA, or a GPU.

[0089] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above-described embodiment of the method for allocating model training tasks in an intelligent computing center across provided computing power can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0090] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above Figure 1 any step in the corresponding embodiment of the method for allocating model training tasks in an intelligent computing center across provided computing power and can achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium can be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0091] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the above Figure 1The various processes of the corresponding method for allocating intelligent computing center model training tasks across different providers of computing power are not elaborated here to avoid repetition, as they can achieve the same technical effects.

[0092] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, as well as the cases where A and B exist together, B and C exist together, A and C exist together, and A, B, and C exist together, a total of 7 cases.

[0093] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to this process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0094] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or other second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0095] The above describes the embodiments of this application in conjunction with the accompanying drawings. However, this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of this application and without departing from the purpose of this application and the scope protected by the claims, can still make many forms, all of which fall within the protection scope of this application.

Claims

1. A method for allocating intelligent computing center model training tasks across computing power providers, characterized in that Including: Step S1: Obtain the identification information of the user and the model training task of the user; Step S2: Determine whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally for the user and is used to provide computing power services for ordinary users and premium users; Step S3: When the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user, search for a second intelligent computing center, which is used to provide computing power services specifically for the premium user; Step S4: When the computing power resources of the second intelligent computing center are greater than the preset threshold, allocate the model training task of the user to the second intelligent computing center to execute the model training task of the user by using the second intelligent computing center.

2. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 1, wherein After the step S4, the method further includes: Step S5: When the computing power resources of the second intelligent computing center are less than the preset threshold, search for a third intelligent computing center, which is deployed near the user and is used to provide computing power services for the ordinary user and the premium user; Step S6: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the model training task of the user to the third intelligent computing center to execute the model training task of the user by using the third intelligent computing center.

3. The method for allocating the intelligent computing center model training task across the provided computing power according to claim 1, wherein The method further includes: Step S7: When the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is an ordinary user, search for a third intelligent computing center, which is deployed near the user and is used to provide computing power services for the ordinary user and the premium user; Step S8: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the model training task of the user to the third intelligent computing center to execute the model training task of the user by using the third intelligent computing center.

4. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 2 or 3, characterized in that, The method further includes: Step S9: When the computing power resources of the third intelligent computing center are less than the preset threshold, compare the computing power resources of the first intelligent computing center and the third intelligent computing center; Step S10: Allocate the model training task of the user to the target intelligent computing center to execute the model training task of the user by using the target intelligent computing center. The target intelligent computing center is the intelligent computing center with the most computing power resources among the first intelligent computing center and the third intelligent computing center.

5. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 2, wherein The third intelligent computing center includes a first GPU graphics card and a second GPU graphics card, and the computing power of the second GPU graphics card is stronger than that of the first GPU graphics card. The user is the premium user, and the step S6 includes: Step S61: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the idle second GPU graphics card of the third intelligent computing center, so as to execute the user's model training task by using the idle second GPU graphics card.

6. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 3, wherein The third intelligent computing center includes a first GPU graphics card and a second GPU graphics card. The computing power of the second GPU graphics card is stronger than that of the first GPU graphics card. The user is an ordinary user. Step S8 includes: Step S81: When the computing power resources of the third intelligent computing center are greater than the preset threshold, allocate the user's model training task to the target GPU graphics card of the third intelligent computing center, so as to execute the user's model training task by using the target GPU graphics card. Wherein, when there is an idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle second GPU graphics card; when there is no idle second GPU graphics card in the third intelligent computing center, the target GPU graphics card is the idle first GPU graphics card.

7. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 2, wherein When the user is a premium user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the priority order of each user among the N users. The priority order of each user is determined based on the payment amount of each user and the required completion time of the model training task, where N is an integer greater than 1.

8. The method for allocating intelligent computing center model training tasks across provided computing power according to claim 3, wherein, When the user is an ordinary user and the number of users is N, the third intelligent computing center executes the model training tasks of the N users based on the order of the start time of the model training tasks of each user among the N users, where N is an integer greater than 1.

9. An intelligent computing center model training task allocation device for cross-providing computing power, characterized in that, Including: A first acquisition module, configured to acquire the identification information of the user and the user's model training task; A first judgment module, configured to judge whether the computing power resources of the first intelligent computing center are greater than a preset threshold. The first intelligent computing center is deployed locally to the user, and the first intelligent computing center is used to provide computing power services for ordinary users and premium users; A first search module, configured to search for a second intelligent computing center when the computing power resources of the first intelligent computing center are less than the preset threshold and the identification information of the user indicates that the user is a premium user. The second intelligent computing center is an intelligent computing center dedicated to providing computing power services for premium users; A first allocation module, configured to allocate the user's model training task to the second intelligent computing center when the computing power resources of the second intelligent computing center are greater than the preset threshold, so as to execute the user's model training task by using the second intelligent computing center.

10. An electronic device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the method for allocating model training tasks across intelligent computing centers providing computing power as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for allocating cross-providing computing power intelligent computing center model training tasks as described in any one of claims 1 to 8 are implemented.

12. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method for allocating cross-providing computing power intelligent computing center model training tasks as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Cluster resource scheduling method, device and equipment and computer readable storage medium

    CN109634748A

  • Computing power resource allocation method and device

    CN112988390A

  • Priority scheduling method and system for meteorological machine learning algorithm operation, electronic equipment and computer program product

    CN116643860A

  • Numerical simulation-oriented computing power scheduling method and system

    CN118567840A

  • Method and device for intelligent computing center to provide computing power resources through computing power package

    CN119739442A