Method and device for multi-agent scheduling computing power of intelligent computing center cloud platform

Through the multi-agent scheduling method of the intelligent computing center cloud platform, the computing power resource demand and the acceleration card status are dynamically matched, which solves the problems of low computing power resource utilization and high rental costs, and achieves more efficient resource utilization and cost reduction.

CN120256148BActive Publication Date: 2025-09-09DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510748073.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-09
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The computing power resource utilization rate of the intelligent computing center is low and the rental cost is high. The existing scheduler cannot flexibly adapt to different computing power running tasks and accelerator card status, resulting in excess computing power resources and excessively high rental costs.

Method used

Through the multi-agent scheduling method of the intelligent computing center cloud platform, computing power operation task management, resource monitoring and scheduling agents are utilized to dynamically match computing power resource requirements and accelerator card status to achieve flexible allocation and management of computing power resources.

Benefits of technology

It improves the utilization rate of computing resources, reduces leasing costs, and enables the widespread application of inclusive computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256148B_ABST
    Figure CN120256148B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for multi-agent scheduling of computing power on an intelligent computing center cloud platform, relating to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure. The method comprises: step S1, obtaining computing power resource demand information of a first computing power operation task in a computing power operation task scheduling queue through a computing power operation task management agent of the intelligent computing center cloud platform; step S2, obtaining computing power resource usage information of multiple accelerator cards through a computing power resource monitoring agent of the intelligent computing center cloud platform; step S3, allocating a first accelerator card to process the first computing power operation task based on the computing power resource demand information and computing power resource usage information through a computing power scheduling agent of the intelligent computing center cloud platform, and matching the computing power resource usage information of the first accelerator card with the computing power resource demand information of the first computing power operation task. The present invention can greatly improve computing power resource utilization, thereby realizing the universal application of computing power resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a method and device for multi-agent scheduling of computing power on an intelligent computing center cloud platform. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that utilizes large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, providing a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of the "Intelligent Computing Center", "Intelligent Computing Center" cloud platform and "Intelligent Computing Center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] In intelligent computing centers, computing power resources are usually provided by accelerator cards to execute computing power operation tasks. In the prior art, the computing power operation tasks to be scheduled and the status of the current accelerator card are obtained through a pre-written scheduler, and then the accelerator card is allocated to the computing power operation task through the scheduler. However, in the prior art, the pre-written scheduler can only simply allocate accelerator cards according to preset rules, and cannot flexibly adapt to different situations. There is a situation where different computing power operation tasks and the status of different accelerator cards cannot be taken into account, resulting in an excess of computing power resources of the accelerator cards after the accelerator cards are allocated, resulting in a very low utilization rate of the computing power resources of the intelligent computing center. At the same time, due to the excess computing power resources provided by the intelligent computing center, this part of the computing power resources also needs to be borne by users when leasing computing power services, resulting in high leasing costs, making it difficult to achieve widespread application of inclusive computing power.

[0008] It can be seen that the existing technology has the problems of low computing resource utilization of intelligent computing centers and high computing power leasing costs. Summary of the Invention

[0009] The embodiments of the present invention provide a method and device for multi-agent scheduling computing power of an intelligent computing center cloud platform to solve the problems in the prior art of low computing power resource utilization and high computing power leasing costs in intelligent computing centers.

[0010] To solve the above problems, the present invention is achieved as follows:

[0011] In a first aspect, an embodiment of the present invention provides a method for scheduling computing power of multiple agents on an intelligent computing center cloud platform, comprising:

[0012] Step S1: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management agent of the intelligent computing center cloud platform;

[0013] Step S2: Obtain computing resource usage information of multiple accelerator cards through the computing resource monitoring agent of the intelligent computing center cloud platform;

[0014] Step S3: Through the computing power scheduling intelligent body of the intelligent computing center cloud platform, a first accelerator card is allocated to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information. The first accelerator card is one of the multiple accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task.

[0015] In one embodiment, the computing power operation task management agent, the computing power resource monitoring agent, and the computing power scheduling agent are agents within a first virtual space, and the first virtual space is used to schedule the multiple accelerator cards to process the computing power operation tasks included in the computing power operation task scheduling queue;

[0016] The step S3 comprises:

[0017] Step S31: broadcasting the computing power resource demand information in the first virtual space through the computing power operation task management agent;

[0018] Step S32: broadcast computing resource usage information of the multiple accelerator cards in the first virtual space through the computing resource monitoring agent;

[0019] Step S33: Receive, in the first virtual space, the computing resource demand information and the computing resource usage information of the multiple accelerator cards through the computing resource scheduling agent;

[0020] Step S34: Allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0021] In one embodiment, step S3 includes:

[0022] Step S31′: broadcasting first information through the computing power operation task management agent, where the first information includes a first identifier and the computing power resource requirement information, where the first identifier is used to represent the computing power operation task scheduled to be processed by the multiple accelerator cards;

[0023] Step S32′: broadcasting second information through the computing resource monitoring agent, where the second information includes the first identifier and computing resource usage information of the multiple accelerator cards;

[0024] Step S33′: receiving the first information and the second information based on the first identifier through the computing power scheduling agent;

[0025] Step S34': Allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0026] In one embodiment, step S1 includes:

[0027] Step S11: Obtaining, through the computing power operation task management agent, the task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue;

[0028] Step S12: determining, by the computing power operation task management agent, the first computing power operation task from the at least one computing power operation task based on the task priority;

[0029] Step S13: Obtain computing power resource requirement information of the first computing power operation task through the computing power operation task management agent.

[0030] In one embodiment, step S11 includes:

[0031] Step S111: Obtaining, by the computing power running task management agent, the task priority of each computing power running task in at least one computing power running task included in the computing power running task scheduling queue based on a first model context protocol MCP interface;

[0032] The step S13 includes:

[0033] Step S131: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue based on the first MCP interface through the computing power running task management agent.

[0034] In one embodiment, step S2 includes:

[0035] Step S21: Obtain computing resource usage information of multiple accelerator cards from a computing resource collector based on a second MCP interface through the computing resource monitoring agent. The computing resource collector is a component deployed on the intelligent computing center cloud platform for collecting computing resource usage information of the multiple accelerator cards.

[0036] The step S3 comprises:

[0037] Step S31'': Determine the first accelerator card based on the computing resource demand information and the computing resource usage information by the computing resource scheduling agent;

[0038] Step S32'': through the computing power scheduling agent, allocate the first accelerator card based on the third MCP interface to process the first computing power operation task.

[0039] In a second aspect, an embodiment of the present invention further provides a device for scheduling computing power of multiple agents on an intelligent computing center cloud platform, comprising:

[0040] The first acquisition module is used to obtain computing power resource demand information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management agent of the intelligent computing center cloud platform;

[0041] A second acquisition module is used to obtain computing resource usage information of multiple accelerator cards through the computing resource monitoring agent of the intelligent computing center cloud platform;

[0042] A scheduling module is used to allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling intelligent body of the intelligent computing center cloud platform, wherein the first accelerator card is one of the multiple accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task.

[0043] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for multi-agent scheduling computing power of an intelligent computing center cloud platform as described in the first aspect above are implemented.

[0044] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for multi-agent scheduling computing power of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0045] In a fifth aspect, the present invention also provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps in the method for multi-agent scheduling computing power of an intelligent computing center cloud platform as described in the first aspect above.

[0046] In the present invention, step S1, through the computing power operation task management intelligent body of the intelligent computing center cloud platform, obtain the computing power resource demand information of the first computing power operation task in the computing power operation task scheduling queue; step S2, through the computing power resource monitoring intelligent body of the intelligent computing center cloud platform, obtain the computing power resource usage information of multiple acceleration cards; step S3, through the computing power scheduling intelligent body of the intelligent computing center cloud platform, allocate the first acceleration card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information, the first acceleration card is one of the multiple acceleration cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task. In this way, the computing power running task management agent obtains the computing power resource demand information of the first computing power running task in the computing power running task scheduling queue, and can realize the management of the computing power running tasks in the computing power running task scheduling queue; the computing power resource monitoring agent obtains the computing power resource usage information of multiple accelerator cards, and realizes the monitoring of the computing power resource usage information of different accelerator cards; the computing power scheduling agent realizes the allocation of the first accelerator card to execute the first computing power task, thereby taking into account the management of computing power running tasks, the monitoring of the computing power resource usage information of accelerator cards, and the allocation of accelerator cards through different agents. Compared with pre-written scheduling programs, it can adapt to different situations more flexibly, taking into account the status of different computing power running tasks and different accelerator cards, and greatly improving the computing power resource utilization rate of the intelligent computing center. At the same time, users no longer need to bear the excess computing power resources when renting computing power services, which greatly reduces the rental cost and realizes the widespread application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in describing the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0048] Figure 1This is a flow chart of a method for scheduling computing power among multiple agents on an intelligent computing center cloud platform provided by an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of interaction between intelligent agents provided by an embodiment of the present invention;

[0050] Figure 3 This is a structural diagram of a device for scheduling computing power of multiple agents on an intelligent computing center cloud platform provided by an embodiment of the present invention;

[0051] Figure 4 This is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.

[0053] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0054] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator for measuring the computing power of a data center, including general computing power, supercomputing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0055] The "Network Power" (NP) mentioned in this invention refers to: it is a manifestation of the data transmission capability of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0056] "Storage Power" (SP) as used in this document refers to the comprehensive capabilities of a data center in four areas: data storage capacity, performance, security and reliability, and environmental friendliness and low-carbon development. It serves as a comprehensive indicator of a data center's data storage capabilities, encompassing both external storage devices like storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second). Disaster recovery ratio is a key indicator of security and reliability.

[0057] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0058] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0059] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0060] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0061] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.

[0062] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0063] The term "intelligent computing center" as used in this document refers to a facility that utilizes large-scale heterogeneous computing resources, including general-purpose computing power (CPUs) and intelligent computing power (GPUs, FPGAs, ASICs, etc.), primarily to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). An intelligent computing center encompasses facilities, hardware, and software, providing a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0064] The "intelligent computing center cloud platform" mentioned in the present invention refers to: a cloud computing platform that provides comprehensive services based on the hardware resources and software resources of the intelligent computing center.

[0065] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0066] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0067] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0068] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0069] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water, and electricity.

[0070] The “model” mentioned in the present invention includes but is not limited to a “large language model” and a “large multimodal model”.

[0071] The "large language model" mentioned in this invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0072] The "Multimodal Large Models" mentioned in this invention refer to models that combine multimodal information such as text, images, video, and audio for training, including but not limited to multimodal large language models.

[0073] An "agent," as used in this article, is an agent capable of perceiving its environment and taking actions to achieve specific goals. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. An agent perceives changes in its environment (e.g., through sensors or data input), makes judgments and decisions based on learned knowledge and algorithms, and then executes actions to influence the environment or achieve a predetermined goal. Agents are widely used in the field of artificial intelligence, often found in automated systems, robots, virtual assistants, and game characters. Their core capability lies in their ability to autonomously learn and continuously evolve to better complete tasks and adapt to complex environments.

[0074] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.

[0075] The "inclusive computing power" mentioned in this invention refers to providing appropriate and effective computing power services at an affordable cost to all social classes and groups that have computing power service needs based on the requirements of equal opportunity and the principle of commercial sustainability.

[0076] See Figure 1 , Figure 1 This is a flowchart of a method for scheduling computing power of multiple intelligent agents on an intelligent computing center cloud platform provided by an embodiment of the present invention. Figure 1 As shown, the following steps are included:

[0077] Step S1: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management intelligent agent of the intelligent computing center cloud platform.

[0078] The aforementioned computing task management agent is created based on the computing resources of the intelligent computing center cloud platform and is used to obtain information about the computing task scheduling queue. This agent acts as an expert in task management, enabling better management of computing tasks in the task management queue.

[0079] It should be noted that there are usually multiple computing power running tasks in the computing power running task scheduling queue. The computing power running task management intelligent agent can manage multiple computing power running tasks based on the importance, priority, computing power resource demand information and other aspects of different computing power running tasks.

[0080] The above-mentioned computing power resource demand information is used to characterize the computing power resource situation required by the computing power running task. By obtaining the computing power resource demand information of the first computing power running task, the subsequent computing power scheduling intelligent body can allocate the first accelerator card to the first computing power running task.

[0081] The first accelerator card mentioned above is an accelerator card of the intelligent computing center cloud platform. The accelerator cards other than the first accelerator card of the intelligent computing center cloud platform can be called "other accelerator cards" or "second accelerator cards".

[0082] Step S2: Obtain computing power resource usage information of multiple acceleration cards through the computing power resource monitoring agent of the intelligent computing center cloud platform.

[0083] The computing resource monitoring agent described above is created based on the computing resources of the Intelligent Computing Center cloud platform. It is used to monitor the resources of different accelerator cards on the Intelligent Computing Center cloud platform. The computing resource monitoring agent acts as an expert in task monitoring, enabling better monitoring of computing resource usage across different accelerator cards.

[0084] The above-mentioned computing power resource usage information is used to characterize the computing power resource usage of the accelerator card. It should be noted that in the intelligent computing center cloud platform, computing power resources are usually provided by different accelerator cards (such as GPU accelerator cards). At the same time, the intelligent computing center cloud platform provides computing power resources for different computing power operation tasks to execute computing power operation tasks. In this case, the computing power resources of the accelerator card are in a changing state (for example, unused computing power resources, partially used computing power resources of the accelerator card, and all used computing power resources of the accelerator card). The computing power resource monitoring agent monitors the computing power resource usage information of different accelerator cards, so that the subsequent computing power scheduling agent can allocate the first accelerator card to the first computing power operation task.

[0085] Step S3: Through the computing power scheduling intelligent body of the intelligent computing center cloud platform, a first accelerator card is allocated to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information. The first accelerator card is one of the multiple accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task.

[0086] The aforementioned computing power scheduling agent is created based on the computing power resources of the intelligent computing center cloud platform. It is used to schedule computing power resources on the intelligent computing center cloud platform to allocate different accelerator cards to different computing power tasks. The computing power scheduling agent acts as an expert in task scheduling, assigning the most optimized accelerator card to each computing power task, ensuring that the computing power resources of each accelerator card are fully utilized, thereby improving computing power resource utilization.

[0087] In the present invention, step S1, through the computing power operation task management intelligent body of the intelligent computing center cloud platform, obtain the computing power resource demand information of the first computing power operation task in the computing power operation task scheduling queue; step S2, through the computing power resource monitoring intelligent body of the intelligent computing center cloud platform, obtain the computing power resource usage information of multiple acceleration cards; step S3, through the computing power scheduling intelligent body of the intelligent computing center cloud platform, allocate the first acceleration card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information, the first acceleration card is one of the multiple acceleration cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task. In this way, the computing power running task management agent obtains the computing power resource demand information of the first computing power running task in the computing power running task scheduling queue, and can realize the management of the computing power running tasks in the computing power running task scheduling queue; the computing power resource monitoring agent obtains the computing power resource usage information of multiple accelerator cards, and realizes the monitoring of the computing power resource usage information of different accelerator cards; the computing power scheduling agent realizes the allocation of the first accelerator card to execute the first computing power task, thereby taking into account the management of computing power running tasks, the monitoring of the computing power resource usage information of accelerator cards, and the allocation of accelerator cards through different agents. Compared with pre-written scheduling programs, it can adapt to different situations more flexibly, taking into account the status of different computing power running tasks and different accelerator cards, and greatly improving the computing power resource utilization rate of the intelligent computing center. At the same time, users no longer need to bear the excess computing power resources when renting computing power services, which greatly reduces the rental cost and realizes the widespread application of inclusive computing power.

[0088] Further, if Figure 2As shown, the computing power operation task management agent, the computing power resource monitoring agent and the computing power scheduling agent realize information interaction through broadcasting, so that after the computing power operation task management agent obtains the computing power resource demand information of the first computing power operation task in the computing power operation task scheduling queue, and the computing power resource monitoring agent obtains the computing power resource usage information of multiple acceleration cards, the computing power scheduling agent can obtain the computing power resource demand information of the first computing power operation task and the computing power resource usage information of multiple acceleration cards through broadcasting, and then realize the allocation of the first acceleration card to process the first computing power operation task.

[0089] It should be noted that information exchange via broadcasting can cause unrelated agents to also receive computing resource demand information and computing resource usage information for multiple accelerator cards. To enable accurate information exchange between the computing task management agent, computing resource monitoring agent, and computing scheduling agent, the present invention provides two methods. One is to place the computing task management agent, computing resource monitoring agent, and computing scheduling agent in a virtual space; the other is to implement information exchange between the computing task management agent, computing resource monitoring agent, and computing scheduling agent through identification.

[0090] Specifically, in one embodiment, the computing power operation task management agent, the computing power resource monitoring agent, and the computing power scheduling agent are agents in a first virtual space, and the first virtual space is used to schedule the multiple accelerator cards to process the computing power operation tasks included in the computing power operation task scheduling queue;

[0091] The step S3 comprises:

[0092] Step S31: broadcasting the computing power resource demand information in the first virtual space through the computing power operation task management agent;

[0093] Step S32: broadcast computing resource usage information of the multiple accelerator cards in the first virtual space through the computing resource monitoring agent;

[0094] Step S33: Receive, in the first virtual space, the computing resource demand information and the computing resource usage information of the multiple accelerator cards through the computing resource scheduling agent;

[0095] Step S34: Allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0096] The first virtual space can be Figure 2As shown, the computing power task management agent, computing power resource monitoring agent, and computing power scheduling agent are located in the same first virtual space. Data broadcast by agents in the first virtual space can only be received by other agents in the first virtual space, thus achieving stable information exchange between agents. Thus, the computing power task management agent, computing power resource monitoring agent, and computing power scheduling agent are configured in the first virtual space to obtain computing power resource demand information and computing power resource usage information of multiple accelerator cards, and to allocate the first accelerator card to process the first computing power task through the computing power scheduling agent.

[0097] Specifically, the computing power operation task management agent broadcasts the computing power resource demand information within the first virtual space; the computing power resource monitoring agent broadcasts the computing power resource usage information of the multiple accelerator cards within the first virtual space; and the computing power scheduling agent receives the computing power resource demand information and the computing power resource usage information of the multiple accelerator cards within the first virtual space. In this way, information exchange is achieved through the first virtual space.

[0098] In addition, in one embodiment, step S3 includes:

[0099] Step S31′: broadcasting first information through the computing power operation task management agent, where the first information includes a first identifier and the computing power resource requirement information, where the first identifier is used to represent the computing power operation task scheduled to be processed by the multiple accelerator cards;

[0100] Step S32′: broadcasting second information through the computing resource monitoring agent, where the second information includes the first identifier and computing resource usage information of the multiple accelerator cards;

[0101] Step S33′: receiving the first information and the second information based on the first identifier through the computing power scheduling agent;

[0102] Step S34': Allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0103] Different from the first virtual space, in the present invention, the information interaction between the computing power operation task management agent, the computing power resource monitoring agent and the computing power scheduling agent is realized through the first identifier. Specifically, the computing power operation task management agent broadcasts the first information, the first information includes the first identifier and computing power resource demand information, and the first identifier is used to represent the computing power operation task processed by multiple acceleration cards; the computing power resource monitoring agent broadcasts the second information, the second information includes the first identifier and the computing power resource usage information of multiple acceleration cards; the computing power scheduling agent receives the first information and the second information based on the first identifier. In this way, different computing power scheduling agents determine whether the received first information and second information are the required information through the first identifier, and then determine the computing power resource demand information and computing power resource usage information based on the first information and second information including the first identifier, thereby realizing the allocation of the first acceleration card to process the first computing power operation task through the computing power scheduling agent.

[0104] In one embodiment, step S1 includes:

[0105] Step S11: Obtaining, through the computing power operation task management agent, the task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue;

[0106] Step S12: determining, by the computing power operation task management agent, the first computing power operation task from the at least one computing power operation task based on the task priority;

[0107] Step S13: Obtain computing power resource requirement information of the first computing power operation task through the computing power operation task management agent.

[0108] It should be noted that when managing the computing power running tasks in the computing power running task scheduling queue, the importance or priority of the computing power running tasks is different. It is necessary to determine the computing power tasks to be scheduled for execution by the computing power resources in turn according to the importance or priority of each computing power running task.

[0109] Specifically, in the present invention, the task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue is obtained through the computing power operation task management agent; the first computing power operation task is determined from the at least one computing power operation task based on the task priority through the computing power operation task management agent; and the computing power resource requirement information of the first computing power operation task is obtained through the computing power operation task management agent. In this way, the task priority and computing power resource requirement information of the computing power operation task are obtained through the computing power operation task scheduling queue, and the order of scheduling computing power resources to process the computing power operation tasks is realized according to the task priority to determine the first computing power operation task.

[0110] In some implementations, the above step S1 may also be implemented by the following steps to obtain the computing resource requirement information of the first computing power running task. Specifically, step S1 includes:

[0111] Step S11 ′: obtaining, through the computing power operation task management agent, the task priority and computing power resource requirement information of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue.

[0112] Step S12': determine the first computing power operation task from the at least one computing power operation task based on the task priority and the computing power resource requirement information through the computing power operation task management agent.

[0113] In this way, the computing power running tasks that are scheduled first in the queue are determined by combining task priority and computing power resource demand information.

[0114] In some embodiments, step S1 may further include:

[0115] Step S11', through the computing power operation task management intelligent body, obtain the task priority, computing power resource requirement information and time information of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue, and the time information is used to represent the time for adding each computing power operation task to the task scheduling.

[0116] Step S12': determine the first computing power operation task from the at least one computing power operation task based on the task priority, the computing power resource requirement information and the time learning through the computing power operation task management agent.

[0117] In this way, the computing power running tasks with priority scheduling in the queue are determined by combining task priority, computing power resource demand information and time information.

[0118] In one embodiment, step S11 includes:

[0119] Step S111: Obtaining, by the computing power running task management agent, the task priority of each computing power running task in at least one computing power running task included in the computing power running task scheduling queue based on a first Model Context Protocol (MCP) interface;

[0120] The step S13 includes:

[0121] Step S131: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue based on the first MCP interface through the computing power running task management agent.

[0122] It should be noted that the protocols for interaction between the computing power operation task scheduling queue and the computing power operation task management intelligent agent may be different, and smooth data interaction cannot be achieved. Therefore, in the present invention, the interaction between the computing power operation task scheduling queue and the computing power operation task management intelligent agent is realized through the first MCP interface.

[0123] Among them, the MCP service can be used to encapsulate the application programming interface (API) of the computing power running task scheduling queue. The API of the computing power running task scheduling queue can access the data of the computing power running task scheduling queue. The API is encapsulated through the MCP service to obtain the first MCP interface, which realizes the standardization of the interface, so that the computing power running task management intelligent body obtains the computing power resource demand information of the first computing power running task in the computing power running task scheduling queue through the first MCP interface.

[0124] In one embodiment, step S2 includes:

[0125] Step S21: Through the computing power resource monitoring agent, the computing power resource usage information of multiple accelerator cards is obtained from the computing power resource collector based on the second model context protocol MCP interface. The computing power resource collector is a component deployed by the intelligent computing center cloud platform for collecting the computing power resource usage information of the multiple accelerator cards.

[0126] The step S3 comprises:

[0127] Step S31'': Determine the first accelerator card based on the computing resource demand information and the computing resource usage information by the computing resource scheduling agent;

[0128] Step S32'': through the computing power scheduling agent, allocate the first accelerator card to process the first computing power operation task based on the third model context protocol MCP interface.

[0129] It should be noted that, similar to the interaction between the computing power operation task scheduling queue and the computing power operation task management agent, the interaction between the computing power resource monitoring agent and the computing power resource collector also has the problem that different protocols cannot smoothly interact with data. Therefore, in the present invention, the interaction between the computing power resource monitoring agent and the computing power resource collector is realized through the second MCP interface.

[0130] Specifically, the API of the computing power resource collector is encapsulated through MCP, and the API of the computing power resource collector can access the data of the computing power resource collector. The API is encapsulated through the MCP service to obtain the second MCP interface, thereby realizing the standardization of the interface, so that the computing power resource monitoring intelligent body can obtain the computing power resource usage information in multiple acceleration cards of the computing power resource collector through the second MCP interface.

[0131] Similarly, the computing power scheduling agent allocates the first accelerator card to process the first computing power operation task through the third MCP interface. Specifically, the computing power scheduling agent allocates the first accelerator card to process the first computing power operation task through the third MCP interface. This is achieved through the computing power resource scheduler, which is a component deployed in the intelligent computing center cloud platform for allocating accelerator cards to computing power operation tasks. The computing power scheduling agent exchanges data with the computing power resource scheduler through the third MCP interface.

[0132] Specifically, when the first accelerator card needs to be assigned to process the first computing power operation task, the computing power scheduling agent sends a scheduling instruction to the computing power resource scheduler via the third MCP interface. After receiving the scheduling instruction, the computing power resource scheduler assigns the first accelerator card to process the first computing power operation task. The scheduling instruction includes an identifier of the first accelerator card and an identifier of the first computing power operation task.

[0133] See Figure 3 , Figure 3 This is a structural diagram of a multi-agent computing power scheduling device for an intelligent computing center cloud platform provided by an embodiment of the present invention. Figure 3 As shown, the multi-agent computing power scheduling device 300 of the intelligent computing center cloud platform includes:

[0134] The first acquisition module 301 is used to obtain computing power resource demand information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management agent of the intelligent computing center cloud platform;

[0135] A second acquisition module 302 is configured to acquire computing resource usage information of multiple accelerator cards through a computing resource monitoring agent of the intelligent computing center cloud platform;

[0136] The scheduling module 303 is used to allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling intelligent body of the intelligent computing center cloud platform. The first accelerator card is one of the multiple accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task.

[0137] In one embodiment, the computing power operation task management agent, the computing power resource monitoring agent, and the computing power scheduling agent are agents within a first virtual space, and the first virtual space is used to schedule the multiple accelerator cards to process the computing power operation tasks included in the computing power operation task scheduling queue;

[0138] The step scheduling module 303 includes:

[0139] A first broadcasting unit, configured to broadcast the computing power resource demand information in the first virtual space through the computing power operation task management agent;

[0140] a second broadcasting unit, configured to broadcast computing resource usage information of the plurality of accelerator cards in the first virtual space through the computing resource monitoring agent;

[0141] A first receiving unit is configured to receive, in the first virtual space, the computing resource demand information and the computing resource usage information of the plurality of accelerator cards through the computing resource scheduling agent;

[0142] The first scheduling unit is used to allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0143] In one embodiment, the scheduling module 303 includes:

[0144] a third broadcasting unit, configured to broadcast first information through the computing power operation task management agent, where the first information includes a first identifier and the computing power resource requirement information, where the first identifier is used to represent the computing power operation task scheduled for processing by the multiple accelerator cards;

[0145] a fourth broadcasting unit, configured to broadcast second information through the computing resource monitoring agent, where the second information includes the first identifier and computing resource usage information of the plurality of accelerator cards;

[0146] A second receiving unit is configured to receive the first information and the second information based on the first identifier through the computing power scheduling agent;

[0147] The second scheduling unit is used to allocate the first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

[0148] In one embodiment, the first acquisition module 301 includes:

[0149] A first acquisition unit is configured to acquire, through the computing power operation task management agent, a task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue;

[0150] A first determining unit is configured to determine, through the computing power operation task management agent, the first computing power operation task from the at least one computing power operation task based on the task priority;

[0151] The second acquisition unit is used to obtain the computing power resource requirement information of the first computing power running task through the computing power running task management intelligent agent.

[0152] In one embodiment, the first acquiring unit includes:

[0153] A first acquisition subunit is configured to acquire, through the computing power operation task management agent, a task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue based on a first model context protocol MCP interface;

[0154] The second acquiring unit includes:

[0155] The second acquisition subunit is used to obtain the computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue based on the first model context protocol MCP interface through the computing power running task management intelligent agent.

[0156] In one embodiment, the second acquisition module 302 includes:

[0157] The third acquisition unit is used to obtain the computing power resource usage information of multiple accelerator cards from the computing power resource collector based on the second model context protocol MCP interface through the computing power resource monitoring intelligent agent. The computing power resource collector is a component deployed on the intelligent computing center cloud platform for collecting the computing power resource usage information of the multiple accelerator cards.

[0158] The scheduling module 303 includes:

[0159] A second determining unit is configured to determine, through the computing power scheduling agent, the first accelerator card based on the computing power resource demand information and the computing power resource usage information;

[0160] The third scheduling unit is used to allocate the first accelerator card to process the first computing power operation task based on the third MCP interface through the computing power scheduling agent.

[0161] The device for multi-agent scheduling computing power of the intelligent computing center cloud platform provided in the embodiment of the present invention is capable of realizing the various processes of each embodiment of the method for multi-agent scheduling computing power of the above-mentioned intelligent computing center cloud platform. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0162] It should be noted that the device for scheduling computing power of multiple agents of the intelligent computing center cloud platform in the embodiment of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0163] The present invention also provides an electronic device, see Figure 4 , Figure 4 The electronic device includes a memory 401, a processor 402, and a program or instruction stored in the memory 401 and executed by the processor 402. Figure 1 Any steps in the corresponding embodiment of the method for multi-agent scheduling computing power of the intelligent computing center cloud platform and the same beneficial effects are achieved will not be repeated here.

[0164] The processor 402 may be a CPU, an ASIC, an FPGA, or a GPU.

[0165] Those skilled in the art will understand that all or part of the steps of the embodiment of the method for implementing multi-agent scheduling computing power of the above-mentioned intelligent computing center cloud platform can be completed through hardware related to program instructions, and the program can be stored in a readable medium.

[0166] The present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above Figure 1 Any steps in the corresponding embodiments of the method for scheduling computing power of a multi-agent on an intelligent computing center cloud platform can achieve the same technical effect and are not described here in detail to avoid repetition. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0167] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the embodiment of the method for multi-agent scheduling computing power of the corresponding intelligent computing center cloud platform can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0168] The terms "first", "second" and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. In addition, the terms "comprise" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. In addition, "and / or" is used in this application to represent at least one of the connected objects, for example A and / or B and / or C, which means comprising seven situations including single A, single B, single C, and both A and B exist, both B and C exist, both A and C exist, and both A, B and C exist.

[0169] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of each embodiment of this application.

[0171] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for scheduling computing power of multiple agents on an intelligent computing center cloud platform, characterized in that: include: Step S1: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management agent of the intelligent computing center cloud platform; Step S2: Obtain computing resource usage information of multiple accelerator cards through the computing resource monitoring agent of the intelligent computing center cloud platform; Step S3: Allocating, by the computing power scheduling agent of the intelligent computing center cloud platform, a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information, where the first accelerator card is one of the multiple accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task; The step S3 comprises: Step S31′: broadcasting first information through the computing power operation task management agent, where the first information includes a first identifier and the computing power resource requirement information, where the first identifier is used to represent the computing power operation task scheduled to be processed by the multiple accelerator cards; Step S32′: broadcasting second information through the computing resource monitoring agent, where the second information includes the first identifier and computing resource usage information of the multiple accelerator cards; Step S33′: receiving the first information and the second information based on the first identifier through the computing power scheduling agent; Step S34': Allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

2. The method according to claim 1, wherein The step S1 comprises: Step S11: Obtaining, through the computing power operation task management agent, the task priority of each computing power operation task in at least one computing power operation task included in the computing power operation task scheduling queue; Step S12: Determine, by the computing power operation task management agent, the first computing power operation task from the at least one computing power operation task based on the computing power operation task priority; Step S13: Obtain computing power resource requirement information of the first computing power operation task through the computing power operation task management agent.

3. The method according to claim 2, wherein The step S11 includes: Step S111: Obtaining, by the computing power running task management agent, the task priority of each computing power running task in at least one computing power running task included in the computing power running task scheduling queue based on a first model context protocol MCP interface; The step S13 includes: Step S131: Obtain computing power resource requirement information of the first computing power running task in the computing power running task scheduling queue based on the first model context protocol MCP interface through the computing power running task management agent.

4. The method according to claim 1, wherein The step S2 comprises: Step S21: Obtain computing power resource usage information of multiple accelerator cards from a computing power resource collector based on a second model context protocol (MCP) interface through the computing power resource monitoring agent. The computing power resource collector is a component deployed on the intelligent computing center cloud platform for collecting computing power resource usage information of the multiple accelerator cards. The step S3 comprises: Step S31″: Determine the first accelerator card based on the computing resource demand information and the computing resource usage information through the computing resource scheduling agent; Step S32″: through the computing power scheduling agent, allocate the first accelerator card to process the first computing power operation task based on the third model context protocol MCP interface.

5. A device for scheduling computing power of multiple agents on an intelligent computing center cloud platform, characterized in that: include: The first acquisition module is used to obtain computing power resource demand information of the first computing power running task in the computing power running task scheduling queue through the computing power running task management agent of the intelligent computing center cloud platform; A second acquisition module is used to obtain computing resource usage information of multiple accelerator cards through the computing resource monitoring agent of the intelligent computing center cloud platform; a scheduling module, configured to allocate, through a computing power scheduling agent of the intelligent computing center cloud platform, a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information, wherein the first accelerator card is one of the plurality of accelerator cards, and the computing power resource usage information of the first accelerator card matches the computing power resource demand information of the first computing power operation task; The scheduling module includes: a third broadcasting unit, configured to broadcast first information through the computing power operation task management agent, where the first information includes a first identifier and the computing power resource requirement information, where the first identifier is used to represent the computing power operation task scheduled for processing by the multiple accelerator cards; a fourth broadcasting unit, configured to broadcast second information through the computing resource monitoring agent, where the second information includes the first identifier and computing resource usage information of the plurality of accelerator cards; A second receiving unit, configured to receive the first information and the second information based on the first identifier through the computing power scheduling agent; The second scheduling unit is used to allocate a first accelerator card to process the first computing power operation task based on the computing power resource demand information and the computing power resource usage information through the computing power scheduling agent.

6. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for multi-agent scheduling computing power of an intelligent computing center cloud platform as described in any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for multi-agent scheduling computing power of an intelligent computing center cloud platform as described in any one of claims 1 to 4.

8. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for multi-agent scheduling computing power of an intelligent computing center cloud platform as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Distributed computing power scheduling method and device, electronic equipment and storage medium

    CN116069498A

  • Intelligent computing center computing power asymmetric collaborative scheduling method oriented to common computing power

    CN119938284A