AI computing power scheduling method, system and device based on artificial intelligence public service platform and medium

The AI ​​computing power scheduling method and system of the artificial intelligence public service platform has solved the problem of low computing power resource utilization, realized the efficient integration and intelligent scheduling of computing power resources, supported the large-scale implementation of large model technology in industrial applications, simplified the development process of intelligent agents, and improved resource utilization and scheduling efficiency.

CN121785730APending Publication Date: 2026-04-03SHANGHAI INSPUR CLOUD COMPUTING SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Currently, the utilization rate of computing resources in the construction of large-scale models is low, the supply and demand are not accurately matched, and it is difficult to achieve collaborative optimization of computing power, model and intelligent agent. In particular, the computing power demand fluctuates greatly and the task types are diverse during the training and inference of large-scale models, and the existing scheduling methods are difficult to meet the requirements of real-time performance and intelligence.

Method used

This paper presents an AI computing power scheduling method and system based on an artificial intelligence public service platform. Through computing power management, monitoring, and scheduling, it achieves efficient integration and intelligent scheduling of computing resources. Computing power management includes resource pool tag management, device information display, and monitoring; computing power monitoring achieves cluster monitoring through a comprehensive observation system; and computing power scheduling is elastically allocated based on resource operation status to meet demand and reduce resource consumption.

Benefits of technology

It improves the utilization rate of computing resources, realizes the efficient integration and intelligent scheduling of computing resources, supports the large-scale implementation of large model technology in industrial applications, simplifies the development process of intelligent agents, and improves resource utilization and scheduling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785730A_ABST
    Figure CN121785730A_ABST
Patent Text Reader

Abstract

The invention discloses an AI computing power scheduling method, system and device based on an artificial intelligence public service platform and a medium, belongs to the technical field of artificial intelligence and large models, and aims to solve the technical problem of how to realize efficient integration, intelligent scheduling and dynamic optimization of computing power resources. In order to support large-scale landing of a large model technology in industrial application, the technical scheme adopted by the invention is as follows: computing power management: integrally displaying the computing power condition of a computing power platform, and presenting the operation state, capacity condition and equipment condition information of the platform for a user; computing power monitoring: uniformly monitoring a computing power resource pool network, a server and a storage device, displaying various index curves, and realizing computing power resource pool monitoring; an omnibearing observation system is adopted, a complete monitoring chain is constructed from bottom hardware to upper application, and cluster monitoring is achieved; and computing power scheduling: flexibly scheduling the computing power resources according to the operating condition of the computing power resources, and reducing the resource consumption while meeting the computing power demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and large-scale model technology, specifically to an AI computing power scheduling method, system, device, and medium based on an artificial intelligence public service platform. Background Technology

[0002] With open-source large-scale models like DeepSeek reducing the training cost of models with hundreds of billions of parameters from hundreds of millions of yuan to millions of yuan through architectural innovation and engineering optimization, and lowering inference costs by more than 80%, AI capabilities that were previously only available to tech giants or ultra-large institutions can now be made accessible to government departments through a "computing power as a service" model. For example, a government large-scale model platform built by a certain organization using AI cloud services allows small and medium-sized enterprises to quickly access pre-trained models without having to build specialized teams, enabling "out-of-the-box" use in scenarios such as intelligent customer service and grassroots governance. This revolutionary shift in technology acquisition has directly spurred a transformation in government departments from "passive observation" to "proactive embrace."

[0003] Currently, the construction of large-scale models generally faces problems such as low utilization of computing resources and inaccurate supply-demand matching. Especially during the training and inference processes of large-scale models, computing power demand fluctuates greatly and task types are diverse, making it difficult for existing scheduling methods to achieve collaborative optimization of "computing power-model-agent". At the same time, as the application scenarios of large-scale models evolve from single text processing to multimodal and multi-agent collaborative processes, higher requirements are placed on the real-time, intelligent, and unified nature of computing power scheduling. Therefore, how to achieve efficient integration, intelligent scheduling, and dynamic optimization of computing resources to support the large-scale deployment of large-scale model technology in industrial applications is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] The technical objective of this invention is to provide an AI computing power scheduling method, system, device, and medium based on an artificial intelligence public service platform, in order to solve the problem of how to achieve efficient integration, intelligent scheduling, and dynamic optimization of computing power resources, and support the large-scale implementation of large-scale model technology in industrial applications.

[0005] The technical objective of this invention is achieved as follows: an AI computing power scheduling method based on an artificial intelligence public service platform, the method being as follows:

[0006] Computing Power Management: Provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. Through various customizable data query and reporting functions, it enables data analysis and computing power platform optimization. It also allows for filtering and sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, it allows users to modify device names or add notes to any record through page information editing functions. Modified information is only for display and does not affect the operation of the public service platform. It allows setting computing power types and displaying resource pool usage based on computing power types, thereby showing the data status of computing power nodes and computing power cards in different resource pools.

[0007] Computing power monitoring: Unified monitoring of the computing power resource pool network, servers, and storage devices, displaying various indicator curves to achieve computing power resource pool monitoring; and adopting a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application to achieve cluster monitoring; at the same time, monitoring of computing power nodes and computing power cards;

[0008] Computing power scheduling: Flexible scheduling of computing power resources based on their operational status to meet computing power demands while reducing resource consumption.

[0009] As a preferred option, the computing power type is set as follows: when registering computing power resources, the computing power resources are divided into shared resource pools, dedicated resource pools and special resource pools according to the requirements of the computing power service provider;

[0010] Among them, the shared resource pool carries various computing power tasks, including training, inference, big data, and general computing power;

[0011] The dedicated resource pool is only used for training, computing power, or a certain type of task specified by the service provider.

[0012] The dedicated resource pool provides computing power only for any one task or any one user unit. Computing power requests from other tasks or user units will not be scheduled to the public service platform.

[0013] As a preferred option, the specific computing power management is as follows:

[0014] Node Management: Display information about computing power nodes in units of computing power nodes, bind and configure the resource pools associated with the computing power nodes, and manage and maintain the purpose (inference, training), node name, node IP (testing connectivity), memory, number of computing power cards, computing power card model and description information of the computing power nodes to ensure that the computing power nodes are statistically quantifiable and analyzable.

[0015] Computing power card management: A unified display of high-performance computing power cards or various domestically produced computing power cards is achieved through a hardware abstraction layer, constructing a multi-dimensional resource profile. The multi-dimensional resource profile includes key indicators such as device computing power density, memory capacity and bandwidth, energy efficiency ratio, and topology connection relationship. The computing power cards are displayed and queried based on various dimensions, specifically: displaying the operating status of the computing power cards and marking them with different colors, and viewing the computing power card details; among which, the computing power card details include computing power information, model information, brand information, operating information, and driver version information.

[0016] Resource Pool Management: Based on resource pools, inference and training resource pools are created, and the types (shared, dedicated, and special-purpose), node information, and attribute information of computing resources are maintained. The usage of computing cards is monitored and statistically analyzed based on the resource pool status, enabling the display and management of CPU resource pools. At the same time, unified comprehensive analysis of different computing resource pools is supported, comparing the overall computing power scale, computing power performance, and operational economy to obtain analysis reports, and continuous optimization of the computing power platform is carried out based on the analysis reports.

[0017] As a preferred method, the network metrics monitored by the computing power resource pool include network link bandwidth, network latency, node network card information, and switch throughput; server metrics include the number of servers, the number of computing cards, server load (memory, disk), computing card load, and network throughput information; storage metrics include storage disk type, storage service type, storage load rate, storage network bandwidth, storage IOPS, and storage throughput information.

[0018] Cluster monitoring specifically involves: installing lightweight data collection programs on each computing node to obtain basic operational data such as processor load, memory usage, and storage space in real time; implementing a health check mechanism for core cluster management components to continuously track the response speed and stability of critical services; and establishing a traffic analysis model for network communication to monitor data transmission quality and connection status, and automatically drawing call relationship diagrams between services.

[0019] As a preferred option, computing node monitoring specifically involves real-time monitoring of computing nodes. The monitoring dimensions include performance metrics such as utilization rate, memory utilization rate, video memory usage, video memory utilization rate, computing card utilization rate, computing card power, network bandwidth, and network inbound and outbound packet volume, helping operation and maintenance units to achieve real-time monitoring and maintenance of computing power.

[0020] As a preferred option, the monitoring of the computing power card specifically involves real-time monitoring of the GPU card, with the monitoring dimensions being the amount and rate of GPU memory usage. This allows for real-time monitoring of the dynamic operating information of the GPU card, providing reliable data support for subsequent dynamic scheduling of model service resources based on the usage of the GPU card.

[0021] More optimally, the computing power scheduling is as follows:

[0022] Dynamic allocation: The computing resource scheduling service supports dynamic allocation of resources based on task priority, ensuring that high-priority tasks (such as large model pre-training and real-time inference) receive high-performance computing resources first. At the same time, it dynamically schedules tasks to low-load nodes based on real-time monitoring of computing resource utilization to optimize resource allocation. For differentiated scheduling of task types, training tasks are prioritized to match high-performance computing cards and support multi-card collaborative training, while inference tasks select low-latency nodes according to concurrency requirements, achieving a balance between cluster resource utilization and task efficiency, and ensuring the stable operation and timely response of critical tasks.

[0023] Multi-source heterogeneous computing power scheduling: Supports the identification and classification of different types of computing power from multiple cloud vendors and multiple sources in the cluster, and records the corresponding performance parameters (such as memory size, number of computing cores, computing power, etc.); Each computing power device is tagged with a corresponding label through a label management mechanism for differentiation during scheduling, and appropriate computing power devices are selected based on the characteristics and requirements of the task; matching tasks with specific computing power devices is achieved through rule configuration; multi-dimensional statistical analysis capabilities based on demand type are provided to track the utilization rate, task distribution, and performance of various types of computing power resources in real time, and generate visual analysis reports to provide data support for resource optimization and capacity planning;

[0024] Load balancing: Based on various computing power nodes and computing power resource pools, in the actual computing power allocation process, the load balancing scheduling of computing power is carried out according to two levels: computing power resource pool and computing power node.

[0025] An AI computing power scheduling system based on an artificial intelligence public service platform is disclosed. This system implements the AI ​​computing power scheduling method described above. The system includes:

[0026] The computing power management module provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. It enables data analysis and computing power platform optimization through various customizable data query and reporting functions. It also allows for filtering and sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, the module allows users to modify device names or add notes to any record using page information editing functions. These modifications are for display purposes only and do not affect the operation of the public service platform. The module also allows users to set computing power types and display resource pool usage based on these types, thereby showing the data status of computing power nodes and computing power cards in different resource pools.

[0027] The computing power monitoring module is used to uniformly monitor the computing power resource pool network, servers, and storage devices, display various indicator curves, and realize the monitoring of the computing power resource pool. It adopts a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application to realize cluster monitoring. At the same time, it monitors computing power nodes and computing power cards.

[0028] The computing power scheduling module is used to flexibly schedule computing power resources based on their operational status, thereby meeting computing power demands while reducing resource consumption.

[0029] An electronic device includes: a memory and at least one processor;

[0030] The memory stores computer-executed instructions;

[0031] The at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to execute the AI ​​computing power scheduling method based on the artificial intelligence public service platform described above.

[0032] A computer-readable storage medium storing computer-executable instructions, wherein when a processor executes the computer-executable instructions, the AI ​​computing power scheduling method based on an artificial intelligence public service platform described above is implemented.

[0033] The AI ​​computing power scheduling method, system, device, and medium based on an artificial intelligence public service platform of the present invention have the following advantages:

[0034] (i) This invention reduces the cost of accessing AI technology by providing computing power, algorithm library and data resources, assists in the application of AI in various scenarios, and enables the rapid application of AI technology without the need for development from scratch, thereby achieving resource sharing and technology accessibility.

[0035] (ii) This invention optimizes resource allocation through AI algorithms, improves resource utilization efficiency, optimizes resource allocation, and enhances efficiency;

[0036] (III) Based on the public service platform for artificial intelligence, this invention relies on the core functions of multi-dimensional computing power management of GPU resources, unified supply of model services, and monitoring and alarm mechanisms for computing power monitoring. With the core concept of "unified supply and unified supervision", it constructs an intelligent computing power scheduling system, which can integrate computing power resource infrastructure to achieve efficient, stable and scalable computing power scheduling services. Through basic (industry) large models, algorithm components and AI intelligent agent application services, it constructs a full-scenario intelligent agent application ecosystem to empower AI applications and business innovation.

[0037] (iv) Based on the core functions of GPU resource multi-dimensional computing power management, unified supply of model services, and monitoring and alarm mechanisms for computing power monitoring, this invention provides a full-stack artificial intelligence solution of "computing power-model-intelligent agent", provides unified computing power support for resource scheduling and management, simplifies the intelligent agent development process, achieves optimal matching of computing power supply and demand, and greatly improves the utilization rate of computing power resources.

[0038] (v) This invention achieves optimal matching of computing power supply and demand and improves resource utilization through strategies such as computing power perception, scheduling and load management;

[0039] (vi) This invention achieves unified management of diverse heterogeneous computing power based on a standardized interface protocol, providing unified computing power support for resource scheduling and management, and realizing the management of diverse computing power. Attached Figure Description

[0040] The invention will be further described below with reference to the accompanying drawings.

[0041] Appendix Figure 1 A schematic diagram for the management of multiple computing power sources;

[0042] Appendix Figure 2 This is a schematic diagram of intelligent computing power scheduling. Detailed Implementation

[0043] The AI ​​computing power scheduling method, system, device, and medium based on the artificial intelligence public service platform of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] Example 1:

[0045] Appendix Figure 1 and attached Figure 2 As shown in the figure, this embodiment provides an AI computing power scheduling method based on an artificial intelligence public service platform. The method is as follows:

[0046] S1. Computing Power Management: Provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. Through various customizable data query and reporting functions, it enables data analysis and computing power platform optimization. It also allows for filtering or sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, it allows users to modify device names or add notes to any record through page information editing functions. Modified information is only for display and does not affect the operation of the public service platform. It allows setting computing power types and displaying resource pool usage based on computing power types, thereby showing the data status of computing power nodes and computing power cards in different resource pools.

[0047] S2. Computing Power Monitoring: Provides unified monitoring of the computing power resource pool network, servers, and storage devices, displays various indicator curves, and realizes the monitoring of the computing power resource pool; it also adopts a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application, realizing cluster monitoring; and monitors computing power nodes and computing power cards.

[0048] S3. Computing power scheduling: Flexible scheduling of computing power resources based on the operation status of computing power resources to meet computing power demand while reducing resource consumption.

[0049] In this embodiment, the computing power type is specifically set as follows: when registering computing power resources, the computing power resources are divided into shared resource pool, dedicated resource pool and special resource pool according to the requirements of the computing power service provider;

[0050] Among them, the shared resource pool carries various computing power tasks, including training, inference, big data, and general computing power;

[0051] The dedicated resource pool is only used for training, computing power, or a certain type of task specified by the service provider.

[0052] The dedicated resource pool provides computing power only for any one task or any one user unit. Computing power requests from other tasks or user units will not be scheduled to the public service platform.

[0053] The computing power management in step S1 of this embodiment is as follows:

[0054] S101, Node Management: Display information about computing power nodes in units of computing power nodes, bind and configure the resource pools associated with computing power nodes, and manage and maintain the purpose (inference, training), node name, node IP (testing connectivity), memory, number of computing power cards, computing power card model and description information of computing power nodes to ensure that computing power nodes are statistically quantifiable and analyzable.

[0055] S102, Computing Power Card Management: A unified display of high-performance computing power cards or various domestically produced computing power cards is achieved through a hardware abstraction layer, constructing a multi-dimensional resource profile. This profile includes key indicators such as device computing power density, memory capacity and bandwidth, energy efficiency ratio, and topology connections. The system displays and queries computing power cards based on various dimensions, specifically: displaying the operating status of the computing power card and marking it with different colors, and viewing computing power card details. These details include computing power information, model information, brand information, operating information, and driver version information.

[0056] S103, Resource Pool Management: Based on resource pools, create inference and training resource pools, maintain the types (shared, dedicated, special), node information, and attribute information of computing resources, and monitor and statistically analyze the usage of computing cards based on the resource pool status, realizing the display and management of CPU resource pools; at the same time, it supports unified comprehensive analysis of different computing resource pools, compares them from the perspectives of overall computing power scale, computing power performance, and operational economy, obtains analysis reports, and continuously optimizes the computing power platform based on the analysis reports.

[0057] In this embodiment, the network metrics monitored in step S2 of the computing resource pool include network link bandwidth, network latency, node network card information, and switch throughput; server metrics include the number of servers, the number of computing cards, server load (memory, disk), computing card load, and network throughput information; storage metrics include storage disk type, storage service type, storage load rate, storage network bandwidth, storage IOPS, and storage throughput information.

[0058] In this embodiment, the cluster monitoring in step S2 specifically involves: installing a lightweight acquisition program on each computing node to obtain basic operational data such as processor load, memory usage, and storage space in real time; for the core components of cluster management, a health check mechanism is provided to continuously track the response speed and stability of key services; in terms of network communication, a traffic analysis model is established to monitor data transmission quality and connection status, and to automatically draw a call relationship diagram between services.

[0059] In this embodiment, the monitoring of computing nodes in step S2 specifically involves real-time monitoring of computing nodes. The monitoring dimensions include performance indicators such as utilization rate, memory utilization rate, video memory usage, video memory utilization rate, computing card utilization rate, computing card power, network bandwidth, and network inbound and outbound packet volume, which helps the operation and maintenance unit to achieve real-time monitoring and maintenance of computing power.

[0060] In this embodiment, the computing card monitoring in step S2 specifically involves real-time monitoring of the GPU card, with the monitoring dimensions being memory usage and memory utilization rate. This allows for real-time monitoring of the GPU card's dynamic operating information, providing reliable data support for subsequent dynamic scheduling of model service resources based on the GPU card's usage.

[0061] In this embodiment, the computing power scheduling in step S3 is as follows:

[0062] S301, Dynamic Allocation: The computing resource scheduling service supports dynamic allocation of resources based on task priority, ensuring that high-priority tasks (such as large model pre-training and real-time inference) receive high-performance computing resources first. At the same time, it dynamically schedules tasks to low-load nodes based on real-time monitoring of computing resource utilization to optimize resource allocation. For differentiated scheduling of task types, training tasks are prioritized to match high-performance computing cards and support multi-card collaborative training, while inference tasks select low-latency nodes according to concurrency requirements, achieving a balance between cluster resource utilization and task efficiency, and ensuring the stable operation and timely response of critical tasks.

[0063] S302, Multi-source Heterogeneous Computing Power Scheduling: Supports the identification and classification of different types of computing power from multiple cloud vendors and multiple heterogeneous sources in the cluster, recording corresponding performance parameters (such as memory size, number of computing cores, computing power, etc.); Each computing power device is tagged with a corresponding label through a label management mechanism for differentiation during scheduling; suitable computing power devices are selected based on task characteristics and requirements; and matching tasks with specific computing power devices is achieved through rule configuration; furthermore, it provides multi-dimensional statistical analysis capabilities based on demand types, real-time tracking of the utilization rate, task distribution, and performance of various types of computing power resources, and generates visual analysis reports to provide data support for resource optimization and capacity planning;

[0064] S303, Load Balancing: Based on various computing power nodes and computing power resource pools, in the actual computing power allocation process, the load balancing scheduling of computing power is carried out according to the two levels of computing power resource pools and computing power nodes.

[0065] Example 2:

[0066] This embodiment provides an AI computing power scheduling system based on an artificial intelligence public service platform. This system is used to implement the AI ​​computing power scheduling method based on an artificial intelligence public service platform as described in Embodiment 1. The system includes:

[0067] The computing power management module provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. It enables data analysis and computing power platform optimization through various customizable data query and reporting functions. It also allows for filtering and sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, the module allows users to modify device names or add notes to any record using page information editing functions. These modifications are for display purposes only and do not affect the operation of the public service platform. The module also allows users to set computing power types and display resource pool usage based on these types, thereby showing the data status of computing power nodes and computing power cards in different resource pools.

[0068] The computing power monitoring module is used to uniformly monitor the computing power resource pool network, servers, and storage devices, display various indicator curves, and realize the monitoring of the computing power resource pool. It adopts a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application to realize cluster monitoring. At the same time, it monitors computing power nodes and computing power cards.

[0069] The computing power scheduling module is used to flexibly schedule computing power resources based on their operational status, thereby meeting computing power demands while reducing resource consumption.

[0070] The computing power management module in this embodiment includes:

[0071] The node management submodule is used to display computing power node information on a per-node basis, bind and configure the resource pool associated with the computing power nodes, and manage and maintain node purpose (inference, training), node name, node IP (for testing connectivity), memory, number of computing power cards, card model, and description information to ensure that computing power nodes are statistically quantifiable and analyzable.

[0072] The computing power card management submodule is used to achieve a unified display of high-performance computing power cards or various domestically produced computing power cards through a hardware abstraction layer, constructing a multi-dimensional resource profile. This profile includes key indicators such as device computing power density, memory capacity and bandwidth, energy efficiency ratio, and topology connections. It also supports displaying and querying computing power cards based on various dimensions. It displays the operating status of computing power cards, supports color-coding, and allows viewing computing power card details, including: computing power information, model information, brand information, operating information, and driver version information.

[0073] The resource pool management submodule is used to create inference and training resource pools based on resource pools, and to maintain the attributes (shared, dedicated, special), node information, and attribute information of computing resources. It also needs to monitor and statistically analyze data such as computing card usage based on resource pool status, enabling the display and management of GPU resource pools. Furthermore, it supports unified and comprehensive analysis of different computing pools, comparing overall computing power scale, computing power performance, and operational economy, and outputting analysis reports. These reports can be used for continuous optimization of the entire computing platform.

[0074] The computing power monitoring module in this embodiment includes:

[0075] The resource pool monitoring submodule is used for unified monitoring of network, server, and storage devices within the computing resource pool, supporting the display of various metric curves. Network metrics include: network link bandwidth, network latency, node network interface card (NIC) information, and switch throughput. Server metrics include: number of servers, number of computing NICs, server load (memory, disk), computing NIC load, and network throughput. Storage metrics include: storage disk type, storage service type, storage load rate, storage network bandwidth, storage IOPS, and storage throughput.

[0076] The cluster monitoring submodule employs a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer applications. The system acquires basic operational data such as processor load, memory usage, and storage space in real time by installing lightweight data collection programs on each computing node. For core cluster management components, a health check mechanism is implemented to continuously track the response speed and stability of critical services. In terms of network communication, a traffic analysis model is established to monitor data transmission quality and connection status, and automatically draw call relationship graphs between services.

[0077] The node monitoring submodule is used to monitor computing nodes in real time. The monitoring dimensions are not limited to performance indicators such as utilization rate, memory utilization rate, video memory usage, video memory utilization rate, computing card utilization rate, computing card power, network bandwidth, and network inbound and outbound packet volume, which helps operation and maintenance units to achieve real-time monitoring and maintenance of computing power.

[0078] The GPU monitoring submodule is used for real-time monitoring of GPU cards, primarily focusing on memory usage and memory utilization rate. It provides real-time insights into the dynamic operation of GPU cards and offers reliable data support for subsequent dynamic scheduling of model service resources based on GPU card usage.

[0079] The computing power scheduling module includes:

[0080] The elastic scaling submodule provides elastic scaling functionality for computing resources. When the computing resources initially requested by the model are found to be insufficient to meet performance requirements during operation, the computing resources can be expanded online based on the elastic scaling function of the computing power scheduling system to meet short-term peak demands during model operation. When the model's concurrency requirements decrease, the computing resources can be reduced according to pre-set scheduling strategies to meet model requirements while reducing resource consumption.

[0081] The dynamic allocation submodule supports dynamic resource allocation based on task priority, ensuring that high-priority tasks (such as large model pre-training and real-time inference) receive high-performance computing resources first. At the same time, it dynamically schedules tasks to low-load nodes based on real-time monitoring of computing resource utilization to optimize resource allocation. For differentiated scheduling of task types, training tasks are prioritized to match high-performance computing cards and support multi-card collaborative training, while inference tasks select low-latency nodes according to concurrency requirements, achieving a balance between cluster resource utilization and task efficiency, and ensuring the stable operation and timely response of critical tasks.

[0082] The multi-source heterogeneous computing power scheduling submodule supports the identification and classification of different types of computing power from multiple cloud vendors and heterogeneous sources within the cluster, recording their performance parameters (such as memory size, number of computing cores, and computing power). Through a tag management mechanism, each computing power device is tagged accordingly for differentiation during scheduling. Simultaneously, appropriate computing power devices are selected based on task characteristics and requirements. Rule configuration enables matching tasks with specific computing power devices. Furthermore, it provides multi-dimensional statistical analysis capabilities based on demand types, tracking the utilization rate, task distribution, and performance of various types of computing power resources in real time, and generating visual analysis reports to provide data support for resource optimization and capacity planning.

[0083] The load balancing submodule is used to manage various computing nodes and computing resource pools based on the computing power management platform. In the actual computing power allocation process, the computing power scheduling platform will perform load balancing scheduling of computing power based on two levels: computing resource pool and computing node.

[0084] Example 3:

[0085] This embodiment also provides an electronic device, including: a memory and at least one processor;

[0086] The memory stores computer-executed instructions;

[0087] The at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to execute the AI ​​computing power scheduling method based on the artificial intelligence public service platform according to any one of the present invention.

[0088] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0089] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0090] Example 4:

[0091] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the AI ​​computing power scheduling method based on an artificial intelligence public service platform according to any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.

[0092] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0093] Examples of storage media used to provide program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0094] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0095] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI computing power scheduling method based on an artificial intelligence public service platform, characterized in that, The method is as follows: Computing Power Management: Provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. Through various customizable data query and reporting functions, it enables data analysis and computing power platform optimization. It also allows for filtering and sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, it allows users to modify device names or add notes to any record through page information editing functions; the modified information is only for display purposes. It allows setting computing power types and displaying resource pool usage based on computing power types, thereby showing the data status of computing power nodes and computing power cards in different resource pools. Computing power monitoring: Unified monitoring of the computing power resource pool network, servers, and storage devices, displaying various indicator curves to achieve computing power resource pool monitoring; and adopting a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application to achieve cluster monitoring; at the same time, monitoring of computing power nodes and computing power cards; Computing power scheduling: Flexible scheduling of computing power resources based on their operational status to meet computing power demands while reducing resource consumption.

2. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to claim 1, characterized in that, Specifically, when registering computing power resources, the computing power resources are divided into shared resource pools, dedicated resource pools, and special resource pools according to the requirements of the computing power service provider. Among them, the shared resource pool carries various computing power tasks, including training, inference, big data, and general computing power; The dedicated resource pool is only used for training, computing power, or a certain type of task specified by the service provider. The dedicated resource pool provides computing power only for a single task or a single user unit.

3. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to claim 1, characterized in that, The specifics of computing power management are as follows: Node Management: Display information about computing power nodes on a per-node basis, bind and configure the resource pools associated with the computing power nodes, and manage and maintain the purpose, node name, node IP, memory, number of computing power cards, computing power card model and description information of the computing power nodes to ensure that the computing power nodes are statistically quantifiable and analyzable. Computing power card management: A unified display of high-performance computing power cards or various domestically produced computing power cards is achieved through a hardware abstraction layer, constructing a multi-dimensional resource profile. The multi-dimensional resource profile includes key indicators such as device computing power density, memory capacity and bandwidth, energy efficiency ratio, and topology connection relationship. The computing power cards are displayed and queried based on various dimensions, specifically: displaying the operating status of the computing power cards and marking them with different colors, and viewing the computing power card details; among which, the computing power card details include computing power information, model information, brand information, operating information, and driver version information. Resource Pool Management: Based on resource pools, inference and training resource pools are created, and the types, node information, and attribute information of computing resources are maintained. The usage of computing cards is monitored and statistically analyzed based on the resource pool status, enabling the display and management of CPU resource pools. At the same time, it supports unified and comprehensive analysis of different computing resource pools, comparing them from the perspectives of overall computing power scale, computing power performance, and operational economy, obtaining analysis reports, and continuously optimizing the computing platform based on the analysis reports.

4. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to claim 1, characterized in that, The network metrics monitored by the computing resource pool include network link bandwidth, network latency, node network card information, and switch throughput; server metrics include the number of servers, the number of computing cards, server load, computing card load, and network throughput information; storage metrics include storage disk type, storage service type, storage load rate, storage network bandwidth, storage IOPS, and storage throughput information. Cluster monitoring specifically involves: installing lightweight data collection programs on each computing node to obtain basic operational data such as processor load, memory usage, and storage space in real time; and having a health check mechanism for core cluster management components to continuously track the response speed and stability of critical services. In terms of network communication, a traffic analysis model is established to monitor data transmission quality and connection status, and to automatically draw a call relationship diagram between services.

5. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to claim 1, characterized in that, The specific monitoring of computing nodes involves real-time monitoring of computing nodes, including performance metrics such as utilization rate, memory utilization rate, video memory usage, video memory utilization rate, computing card utilization rate, computing card power, network bandwidth, and network inbound and outbound packet volume. This helps operations and maintenance units achieve real-time monitoring and maintenance of computing power.

6. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to claim 1, characterized in that, The computing power card monitoring specifically involves real-time monitoring of GPU cards, focusing on memory usage and memory utilization rate. This allows for real-time monitoring of the dynamic operation information of GPU cards, providing reliable data support for subsequent dynamic scheduling of model service resources based on GPU card usage.

7. The AI ​​computing power scheduling method based on an artificial intelligence public service platform according to any one of claims 1 to 6, characterized in that, The specifics of computing power scheduling are as follows: Dynamic allocation: The computing resource scheduling service supports dynamic allocation of resources based on task priority, ensuring that high-priority tasks receive high-performance computing resources first. At the same time, it dynamically schedules tasks to low-load nodes based on real-time monitoring of computing resource utilization to optimize resource allocation. For differentiated scheduling of task types, training tasks are prioritized to match high-performance computing cards and support multi-card collaborative training, while inference tasks select low-latency nodes according to concurrency requirements, achieving a balance between cluster resource utilization and task efficiency, and ensuring the stable operation and timely response of critical tasks. Multi-source heterogeneous computing power scheduling: Supports the identification and classification of different types of computing power from multiple cloud vendors and multiple sources in the cluster, and records the corresponding performance parameters; a tag management mechanism is used to assign corresponding tags to each computing power device for differentiation during scheduling, and appropriate computing power devices are selected based on the characteristics and requirements of the tasks; and matching tasks with specific computing power devices is achieved through rule configuration; furthermore, it provides multi-dimensional statistical analysis capabilities based on demand type, tracks the utilization rate, task distribution and performance of various types of computing power resources in real time, and generates visual analysis reports to provide data support for resource optimization and capacity planning; Load balancing: Based on various computing power nodes and computing power resource pools, in the actual computing power allocation process, the load balancing scheduling of computing power is carried out according to two levels: computing power resource pool and computing power node.

8. An AI computing power scheduling system based on an artificial intelligence public service platform, characterized in that, This system is used to implement the AI ​​computing power scheduling method based on an artificial intelligence public service platform as described in any one of claims 1 to A; the system includes: The computing power management module provides an overall view of the computing power platform, presenting users with information on the platform's operating status, capacity, and equipment status. It enables data analysis and computing power platform optimization through various customizable data query and reporting functions. It also allows for filtering and sorting of various data based on custom tags, enabling resource pool tag management. Furthermore, the module allows users to modify device names or add notes to any record through page information editing functions; the modified information is only for display purposes. It allows setting computing power types and displaying resource pool usage based on these types, thereby showing the data status of computing power nodes and computing power cards in different resource pools. The computing power monitoring module is used to uniformly monitor the computing power resource pool network, servers, and storage devices, display various indicator curves, and realize the monitoring of the computing power resource pool. It adopts a comprehensive observation system to build a complete monitoring chain from the underlying hardware to the upper-layer application to realize cluster monitoring. At the same time, it monitors computing power nodes and computing power cards. The computing power scheduling module is used to flexibly schedule computing power resources based on their operational status, thereby meeting computing power demands while reducing resource consumption.

9. An electronic device, characterized in that, include: Memory and at least one processor; The memory stores computer-executed instructions; The at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to execute the AI ​​computing power scheduling method based on an artificial intelligence public service platform as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, it implements the AI ​​computing power scheduling method based on an artificial intelligence public service platform as described in any one of claims 1 to 7.