Intelligent Computing Center Model Development Method and Device for Inclusive Computing Power

By receiving scheduling instructions, making computing power predictions and obtaining virtual accelerator cards in the intelligent computing center, the problems of low computing power utilization and high model development costs are solved, and efficient computing power utilization and the application of universal computing power are realized.

CN119902904BActive Publication Date: 2025-06-20DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510396850.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The computing power resource utilization rate of existing intelligent computing centers is low, resulting in high cost of model development and it is difficult to achieve widespread application of universal computing power.

Method used

By receiving scheduling instructions, computing power prediction is performed based on the data volume of the model code, and the target virtual accelerator card is a virtual accelerator card created by some GPU computing resources of the GPU accelerator card, which is used to execute model code.

Benefits of technology

It improves the utilization rate of GPU computing power resources of the intelligent computing center, reduces the economic cost of model development, and realizes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902904B_ABST
    Figure CN119902904B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for developing an intelligent computing center model for inclusive computing power, relating to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure. The method includes: Step S1 of the present invention, receiving a scheduling instruction, where the scheduling instruction includes model code; Step S2, performing computing power prediction based on the data volume of the model code to obtain required computing power information; Step S3, obtaining a target virtual acceleration card based on the required computing power information, where the target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center; Step S4, creating a first container based on the target virtual acceleration card, where the first container is used to execute the model code. The present invention can greatly improve the utilization rate of computing power resources of the intelligent computing center and greatly reduce the economic cost of model development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for developing an intelligent computing center model for inclusive computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios for developing, training, and inferring artificial intelligence deep learning models). An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Currently, intelligent computing centers provide computing power resources for users through GPU (Graphics Processing Unit) acceleration cards, so that users can develop models based on the computing power resources of the intelligent computing center and reduce the cost of model development. However, in the prior art, multiple GPU acceleration cards are arranged in the intelligent computing center, and each GPU acceleration card can only process one model development task. In the process of model development, it often happens that not all the computing power resources of the GPU acceleration card are needed, resulting in low utilization rate of the computing power resources of the intelligent computing center. At the same time, since the computing power services provided by the intelligent computing center are limited by the number of GPU acceleration cards, when users lease computing power services, they usually need to lease the entire GPU acceleration card, resulting in a relatively high cost of model development and making it difficult to achieve the widespread application of inclusive computing power.

[0008] It can be seen that there are problems of very low utilization rate of computing power resources and very high cost of model development in the prior art. Summary of the Invention

[0009] The present invention provides a method and device for developing an intelligent computing center model for inclusive computing power, so as to solve the problems of low utilization rate of computing power resources and high cost of model development in the prior art.

[0010] To solve the above problems, the present invention is implemented as follows:

[0011] In a first aspect, the present invention provides a method for developing an intelligent computing center model for inclusive computing power, including:

[0012] Step S1, receiving a scheduling instruction, where the scheduling instruction includes model code;

[0013] Step S2, predicting the computing power based on the data volume of the model code to obtain required computing power information;

[0014] Step S3, obtaining a target virtual acceleration card based on the required computing power information, where the target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center;

[0015] Step S4, creating a first container based on the target virtual acceleration card, where the first container is used to execute the model code.

[0016] In one embodiment, step S3 includes:

[0017] Step S31, scheduling the GPU computing power resources of the target GPU acceleration card based on the required computing power information, where the target GPU acceleration card is an acceleration card with idle GPU computing power resources, and the idle GPU computing power resources in the target GPU acceleration card match the required computing power information;

[0018] Step S32, creating the target virtual acceleration card based on the GPU computing power resources of the target GPU acceleration card.

[0019] In one embodiment, after step S4, the method further includes:

[0020] Step S5, when the execution of the model code is completed in the first container, closing the first container, canceling the registration of the target virtual acceleration card, and releasing the GPU computing power resources of the target virtual acceleration card.

[0021] In one embodiment, step S3 includes:

[0022] Step S33: Obtain multiple pre-created initial virtual acceleration cards and the first computing power resource information corresponding to each initial virtual acceleration card, where the GPU computing power resources of the multiple initial virtual acceleration cards are in an idle state;

[0023] Step S34: Determine a target virtual acceleration card from the multiple initial virtual acceleration cards based on the required computing power information, where the target virtual acceleration card is an acceleration card whose first computing power resource information meets the required computing power information.

[0024] In one embodiment, the scheduling instruction includes the priority corresponding to the model code, and step S3 includes:

[0025] Step S35: Add the scheduling instruction to the execution queue based on the priority;

[0026] Step S36: Obtain a target virtual acceleration card based on the required computing power information in the order of the execution queue.

[0027] In one embodiment, after step S4, the method further includes:

[0028] Step S6: Monitor the idle time of the first container, where the idle time is the time when the first container stops occupying the GPU computing power resources without completing the execution of the model code;

[0029] Step S7: Set the GPU computing power resources occupied by the first container to an idle state when the idle time is greater than or equal to a set time threshold.

[0030] In one embodiment, after step S7, the method further includes:

[0031] Step S8: When the first container receives a running instruction, schedule the idle GPU computing power resources of the first intelligent computing center to process the model code.

[0032] In one embodiment, before step S3, the method further includes:

[0033] Step S9: Obtain the second computing power resource information of the first intelligent computing center;

[0034] The step S3 includes:

[0035] Step S36: When the second computing power resource information meets the required computing power information, obtain a target virtual acceleration card based on the required computing power information.

[0036] In one embodiment, the method further includes:

[0037] Step S10: When the second computing power resource information does not meet the required computing power information, send the scheduling instruction to the second intelligent computing center;

[0038] Step S11: Receive the processing result sent by the second intelligent computing center. The processing result is obtained by the second container executing the model code. The second container is a container created by the second intelligent computing center scheduling the GPU computing power resources in the idle state of the second intelligent computing center.

[0039] Second, the present invention also provides an intelligent computing center model development device for inclusive computing power, including:

[0040] The first receiving module is used to receive a scheduling instruction, and the scheduling instruction includes a model code;

[0041] The prediction module is used to perform computing power prediction based on the data volume of the model code to obtain required computing power information;

[0042] The first obtaining module is used to obtain a target virtual acceleration card based on the required computing power information. The target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center;

[0043] The creation module is used to create a first container based on the target virtual acceleration card. The first container is used to execute the model code.

[0044] Third, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for developing an intelligent computing center model for inclusive computing power as described in the first aspect above.

[0045] Fourth, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for developing an intelligent computing center model for inclusive computing power as described in the first aspect above.

[0046] Fifth, the present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps in the method for developing an intelligent computing center model for inclusive computing power as described in the first aspect above.

[0047] In the present invention, a scheduling instruction is received, and the scheduling instruction includes a model code; computing power prediction is performed based on the data volume of the model code to obtain required computing power information; a target virtual acceleration card is obtained based on the required computing power information, and the target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of a GPU acceleration card in a first intelligent computing center; a first container is created based on the target virtual acceleration card, and the first container is used to execute the model code. In this way, by creating a target virtual acceleration card with partial GPU computing power resources of the GPU acceleration card, the remaining partial GPU computing power resources can be used to create other virtual acceleration cards to execute other computing power tasks, greatly improving the utilization rate of the GPU computing power resources in the intelligent computing center. At the same time, when a user leases computing power services, there is no need to lease the entire GPU acceleration card, and only the computing power resources of the target virtual acceleration card need to be leased, thereby greatly reducing the economic cost in the model development process and realizing the wide application of inclusive computing power. Description of the Drawings

[0048] To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a flowchart of a method for developing a model of an intelligent computing center for inclusive computing power provided by the present invention;

[0050] Figure 2 It is a structural diagram of a device for developing a model of an intelligent computing center for inclusive computing power provided by the present invention;

[0051] Figure 3 It is a structural diagram of an electronic device provided by the present invention. Detailed Embodiments

[0052] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0053] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, which is the ability of computer hardware and software to cooperate to execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which mainly provides services to society through computing power infrastructure.

[0054] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and achieve result output, which is a comprehensive index to measure the computing ability of a data center and includes general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。

[0055] The "carrying capacity" (Network Power, NP) described in the present invention refers to: the performance of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission inside and between data centers, and is a comprehensive index to measure network transmission scheduling ability.

[0056] The "storage power" (Storage Power, SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive index to measure the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), and the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0057] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.

[0058] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology infrastructures such as artificial intelligence, blockchain, and quantum computing.

[0059] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0060] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0061] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.

[0062] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0063] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0064] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0065] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0066] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0067] The "supercomputing center" described in the present invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0068] The "computing power resources" described in the present invention refer to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0069] The "inclusive computing power" described in the present invention refers to providing appropriate and effective computing power services for all social strata and groups with computing power service needs at an affordable cost based on the requirements of equal opportunity and the principle of commercial sustainability.

[0070] The "models" described in the present invention include but are not limited to "large language models" and "multimodal large models".

[0071] The "large language model" described in the present invention refers to a large language model (LLM), which is a language model with a large number of parameters, aiming to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0072] The "multimodal large model" (Multimodal Large Models) described in the present invention refers to a model that jointly trains multi-modal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0073] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for developing an intelligent computing center model for inclusive computing power provided by the present invention. As Figure 1 shown, it includes the following steps:

[0074] Step S1: Receive a scheduling instruction, and the scheduling instruction includes model code.

[0075] The above scheduling instruction is a scheduling instruction sent by a terminal or a CPU container. It should be noted that the terminal or the CPU container is deployed with an editing environment for model code, and the user can adjust and update the model code within the terminal or the CPU container. After the model code in the terminal or the CPU container is adjusted, the user needs to run the model code to confirm the adjustment effect. At this time, the terminal or the CPU container sends a scheduling instruction, and the scheduling instruction includes the model code, so that after receiving the scheduling instruction, the intelligent computing center can run the model code included in the scheduling instruction to obtain a running result.

[0076] Among them, the CPU container is a container created based on the CPU resources of the intelligent computing center, and the terminal is a terminal that has established a communication connection with the intelligent computing center.

[0077] Step S2: Perform computing power prediction based on the data volume of the model code to obtain required computing power information.

[0078] The above required computing power information is used to represent the amount of computing power resources required to execute the model code. It should be noted that different model codes require different amounts of computing power resources. To increase the number of computing power services provided by the intelligent computing center, it is necessary to allocate appropriate computing power resources to each model code according to the required computing power information to improve the utilization rate of computing power resources, and then realize providing computing power services for more users.

[0079] In some embodiments, performing computing power prediction based on the data volume of the model code may be training a prediction model based on the data volume of historical model codes, and then analyzing through the prediction model based on the data volumes of different model codes to obtain the required computing power information corresponding to the model code.

[0080] In some embodiments, performing computing power prediction based on the data volume of the model code may obtain the number of tokens of the model code, and combine and analyze the token number and the data volume of the model code to obtain the required computing power information corresponding to the model code.

[0081] Step S3: Obtain a target virtual acceleration card based on the required computing power information. The target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center.

[0082] The above-mentioned virtual acceleration card is used to provide partial GPU computing power resources of the GPU acceleration card. It should be noted that the computing power resources of a GPU acceleration card can be divided into multiple parts. After division, the partial GPU computing power resources can create virtual acceleration cards, enabling the creation of multiple virtual acceleration cards based on the GPU computing power resources of one GPU acceleration card. Each virtual acceleration card can be used to execute a computing power task for running model code. In this way, one GPU acceleration card can execute multiple computing power tasks in the form of virtual acceleration cards, thereby providing computing power services for more users.

[0083] The above-mentioned target virtual acceleration card is a virtual acceleration card obtained based on demand computing power information, and the target virtual acceleration card meets the demand computing power information. Among them, the target virtual acceleration card is obtained through the demand computing power information. When the target virtual acceleration card meets the demand computing power information, it occupies as little GPU computing power resources of the first intelligent computing center as possible, improving the utilization rate of the GPU computing power resources of the first intelligent computing center.

[0084] In some embodiments, obtaining the target virtual acceleration card based on the demand computing power information may be temporarily scheduling the GPU computing power resources of the first intelligent computing center to create a target virtual acceleration card that meets the demand computing power information.

[0085] In some embodiments, obtaining the target virtual acceleration card based on the demand computing power information may be pre-creating multiple virtual acceleration cards with different computing power resource sizes, and then screening a suitable virtual acceleration card from the multiple virtual acceleration cards as the target virtual acceleration card based on the demand computing power information.

[0086] Step S4: Create a first container based on the target virtual acceleration card, and the first container is used to execute the model code.

[0087] In the present invention, a scheduling instruction is received, and the scheduling instruction includes model code; computing power prediction is performed based on the data volume of the model code to obtain demand computing power information; a target virtual acceleration card is obtained based on the demand computing power information, and the target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center; a first container is created based on the target virtual acceleration card, and the first container is used to execute the model code. In this way, the target virtual acceleration card is created through partial GPU computing power resources of the GPU acceleration card, and the remaining partial GPU computing power resources can be used to create other virtual acceleration cards to execute other computing power tasks, greatly improving the utilization rate of the GPU computing power resources of the intelligent computing center. At the same time, when users lease computing power services, they do not need to lease the entire GPU acceleration card, but only need to lease the computing power resources of the target virtual acceleration card part, thereby greatly reducing the economic cost in the model development process and realizing the wide application of inclusive computing power.

[0088] In one embodiment, step S3 includes:

[0089] Step S31: Schedule the GPU computing power resources of the target GPU acceleration card based on the required computing power information. The target GPU acceleration card is an acceleration card with idle GPU computing power resources, and the idle GPU computing power resources in the target GPU acceleration card match the required computing power information.

[0090] Step S32: Create the target virtual acceleration card based on the GPU computing power resources of the target GPU acceleration card.

[0091] In the present invention, the GPU computing power resources of the target GPU acceleration card are scheduled based on the required computing power information. The target GPU acceleration card is an acceleration card with idle GPU computing power resources, and the idle GPU computing power resources in the target GPU acceleration card match the required computing power information. The target virtual acceleration card is created based on the GPU computing power resources of the target GPU acceleration card. In this way, when it is necessary to execute the model code, the idle GPU computing power resources in the GPU acceleration card are temporarily scheduled to create the target virtual acceleration card, so that the GPU computing power resources of the target virtual acceleration card can just meet the requirements for executing the model code, realizing flexible scheduling of GPU computing power resources, reducing the waste of GPU computing power resources in the first intelligent computing center, and greatly improving the utilization rate of GPU computing power resources in the first intelligent computing center.

[0092] Specifically, when scheduling the GPU computing power resources of the first intelligent computing center, first obtain the GPU acceleration cards with idle GPU computing power resources in the first intelligent computing center, screen the target GPU acceleration cards, and then schedule the idle GPU computing power resources of the GPU acceleration cards to create the target virtual acceleration card.

[0093] Among them, obtaining the GPU acceleration cards with idle GPU computing power resources in the first intelligent computing center includes: obtaining multiple GPU acceleration cards with idle GPU computing power resources and the remaining computing power resource information of each GPU acceleration card.

[0094] Screening the target GPU acceleration cards includes: screening the target GPU acceleration cards from the multiple GPU acceleration cards based on the remaining computing power resource information. The target GPU acceleration card is an acceleration card whose remaining computing power resources meet the required computing power information.

[0095] By screening the target GPU acceleration cards in the above manner, the remaining computing power resources of the target GPU acceleration card meet the required computing power information, and then the target virtual acceleration card can be created by scheduling the idle GPU computing power resources of the GPU acceleration card.

[0096] In one embodiment, after the step S4, the method further includes:

[0097] Step S5, when the execution of the model code is completed in the first container, close the first container, cancel the registration of the target virtual acceleration card, and release the GPU computing power resources of the target virtual acceleration card.

[0098] It should be noted that the target virtual acceleration card is a temporarily created virtual acceleration card. If a first container is created based on the target virtual acceleration card to execute other computing power tasks, there will be a situation of insufficient or excessive GPU computing power resources, and it is impossible to meet the customer's model development needs while reducing the waste of computing power resources. Therefore, in the present invention, when the execution of the model code is completed in the first container, the first container is closed, the registration of the target virtual acceleration card is cancelled, and the GPU computing power resources of the target virtual acceleration card are released, so that idle GPU computing power resources can be obtained when other model codes need to be executed, greatly improving the utilization rate of the GPU computing power resources of the first intelligent computing center.

[0099] In one embodiment, the step S3 includes:

[0100] Step S33, obtain a plurality of pre-created initial virtual acceleration cards and the first computing power resource information corresponding to each initial virtual acceleration card, and the GPU computing power resources of the plurality of initial virtual acceleration cards are in an idle state;

[0101] Step S34, determine a target virtual acceleration card from the plurality of initial virtual acceleration cards based on the required computing power information, where the target virtual acceleration card is an acceleration card whose first computing power resource information meets the required computing power information.

[0102] The above first computing power resource information is used to represent the computing power resource size of the initial pre-acceleration card. The above plurality of initial virtual acceleration cards are pre-created initial virtual acceleration cards. By pre-creating a plurality of initial virtual acceleration cards, there is no need to create virtual acceleration cards again when the model code needs to be executed, which can greatly improve the efficiency of executing the model code.

[0103] In some embodiments, the plurality of initial virtual acceleration cards are initial virtual acceleration cards with different computing power resource sizes, and the target virtual acceleration card is the virtual acceleration card with the smallest computing power resource size that meets the required computing power information. By screening the target virtual acceleration card that meets the required computing power information and has the smallest computing power resource size from the plurality of initial virtual acceleration cards in this way, the first container can execute the model code while occupying as little GPU computing power resources of the first intelligent computing center as possible, thereby greatly improving the utilization rate of the GPU computing power resources of the first intelligent computing center.

[0104] Further, since the multiple initial virtual acceleration cards are initial virtual acceleration cards with different computing power resources, after the model code is executed in the first container, it is not necessary to cancel the target virtual acceleration card. It is only necessary to close the first container and set the GPU computing power resources of the target virtual acceleration card to the idle state, so that the target virtual acceleration card can be obtained again for execution when other model codes need to be executed, improving the execution efficiency.

[0105] In one embodiment, the scheduling instruction includes the priority corresponding to the model code, and the step S3 includes:

[0106] Step S35: Add the scheduling instruction to the execution queue based on the priority;

[0107] Step S36: Obtain the target virtual acceleration card based on the required computing power information in the order of the execution queue.

[0108] In the present invention, by setting priorities to execute different model codes, tasks with higher priorities will be preferentially scheduled for computing power resources for execution, effectively improving the user experience without changing the overall computing power resources of the first intelligent computing center.

[0109] In some embodiments, adding the scheduling instruction to the execution queue based on the priority includes:

[0110] Obtain the sending time of the model code;

[0111] Calculate the score of the model code based on the sending time and the priority;

[0112] Add the scheduling instruction to the execution queue based on the score.

[0113] It should be noted that the priorities of different model codes are different, but the sending times also vary. If the first intelligent computing center continuously receives model codes with high priorities after receiving a model code with a low priority, continuously scheduling the computing power resources to execute the model codes with high priorities will cause the execution time of the model code with a low priority to be too long, resulting in a poor user experience. Therefore, in this embodiment, the sending time and the priority of the model code are combined to obtain the score of the model code, so that the model code with a low priority but an early sending time can also be executed in a timely manner, further improving the user experience.

[0114] Among them, calculating the score of the model code based on the sending time and the priority can specifically be to weight the sending time and the priority to obtain the score of the model code.

[0115] In one embodiment, after the step S4, the method further includes:

[0116] Step S6, monitor the idle time of the first container, where the idle time is the time when the first container stops occupying the GPU computing power resource without completing the execution of the model code;

[0117] Step S7, when the idle time is greater than or equal to the set time threshold, set the GPU computing power resource occupied by the first container to the idle state.

[0118] It should be noted that during the process of the first container executing the model code, the user has a need to stop executing the model code for adjustment. At this time, the model code in the first container has not been completed, but the GPU computing power resource of the first container has not been used. If it is in the state of stopping the execution of the model code for a long time, it will cause waste of the GPU computing power resource.

[0119] Therefore, in the present invention, the idle time of the first container is monitored, where the idle time is the time when the first container stops occupying the GPU computing power resource without completing the execution of the model code; when the idle time is greater than or equal to the set time threshold, the GPU computing power resource occupied by the first container is set to the idle state. In this way, setting the GPU computing power resource occupied by the first container to the idle state enables other containers to occupy this GPU computing power resource to execute the model code, realizing preemptive scheduling of this part of the GPU computing power resource and greatly improving the utilization rate of the GPU computing power resource of the first intelligent computing center.

[0120] In one embodiment, after the step S7, the method further includes:

[0121] Step S8, when the first container receives a running instruction, schedule the idle GPU computing power resource of the first intelligent computing center to process the model code.

[0122] It should be noted that when the idle time of the first container is greater than or equal to the set time threshold, the GPU computing power resource occupied by the first container is set to the idle state. At this time, the first container actually does not occupy the GPU computing power resource. When receiving a running instruction, re-schedule the idle GPU computing power resource of the first intelligent computing center, thereby realizing the processing of the model code and ensuring the normal operation of the first container.

[0123] In one embodiment, before the step S3, the method further includes:

[0124] Step S9, obtain the second computing power resource information of the first intelligent computing center;

[0125] The step S3 includes:

[0126] Step S36: When the second computing power resource information meets the required computing power information, obtain a target virtual acceleration card based on the required computing power information.

[0127] In one embodiment, the method further includes:

[0128] Step S10: When the second computing power resource information does not meet the required computing power information, send the scheduling instruction to a second intelligent computing center;

[0129] Step S11: Receive a processing result sent by the second intelligent computing center, where the processing result is obtained by a second container executing the model code, and the second container is a container created by the second intelligent computing center scheduling the GPU computing power resources in an idle state of the second intelligent computing center.

[0130] It should be noted that the total GPU computing power resources of the first intelligent computing center are limited. In the actual process of providing computing power services, there may be a situation where some intelligent computing centers are overloaded while some other intelligent computing centers have a relatively low load. To achieve the balance between different intelligent computing centers and enable a single intelligent computing center to provide more computing power services, the GPU computing power resources between different intelligent computing centers can be scheduled remotely in the present invention.

[0131] Specifically, when the first intelligent computing center receives a scheduling instruction, it first determines the usage of computing power resources locally (i.e., in the first intelligent computing center), that is, obtains the second computing power resource information of the first intelligent computing center. When the second computing power resource information meets the required computing power information, the model code can be directly processed without scheduling the computing power resources of other intelligent computing centers; when the second computing power resource information does not meet the required computing power information, the model code cannot be processed locally. At this time, the computing power resources of the second intelligent computing center need to be scheduled, and the computing power resources of other intelligent computing centers are required to execute the model code. By this means, the utilization rate of computing power resources of different intelligent computing centers is improved, and the number of computing power services that the intelligent computing center can provide is increased.

[0132] Among them, the above-mentioned second computing power resource information is used to represent the usage of computing power in intelligent computing.

[0133] Please refer to Figure 2 , Figure 2 which is a structural diagram of an intelligent computing center model development device for inclusive computing power provided by the present invention. As Figure 2 shown, the intelligent computing center model development device 200 for inclusive computing power includes:

[0134] A first receiving module 201, configured to receive a scheduling instruction, where the scheduling instruction includes a model code;

[0135] The prediction module 202 is configured to perform computing power prediction based on the data volume of the model code to obtain required computing power information;

[0136] The first acquisition module 203 is configured to acquire a target virtual acceleration card based on the required computing power information, where the target virtual acceleration card is a virtual acceleration card created based on partial GPU computing power resources of the GPU acceleration card of the first intelligent computing center;

[0137] The creation module 204 is configured to create a first container based on the target virtual acceleration card, and the first container is used to execute the model code.

[0138] In one embodiment, the first acquisition module 203 includes:

[0139] The scheduling unit is configured to schedule the GPU computing power resources of the target GPU acceleration card based on the required computing power information. The target GPU acceleration card is an acceleration card with idle GPU computing power resources, and the idle GPU computing power resources in the target GPU acceleration card match the required computing power information;

[0140] The creation unit is configured to create the target virtual acceleration card based on the GPU computing power resources of the target GPU acceleration card.

[0141] In one embodiment, after the creation module 204, the intelligent computing center model development device 200 for inclusive computing power further includes:

[0142] The closing module is configured to close the first container, cancel the registration of the target virtual acceleration card, and release the GPU computing power resources of the target virtual acceleration card when the execution of the model code is completed in the first container.

[0143] In one embodiment, the first acquisition module 203 includes:

[0144] The first acquisition unit is configured to acquire a plurality of pre-created initial virtual acceleration cards and the corresponding first computing power resource information of each initial virtual acceleration card, and the GPU computing power resources of the plurality of initial virtual acceleration cards are in an idle state;

[0145] The determination module is configured to determine a target virtual acceleration card from the plurality of initial virtual acceleration cards based on the required computing power information, and the target virtual acceleration card is an acceleration card whose first computing power resource information meets the required computing power information.

[0146] In one embodiment, the scheduling instruction includes the priority corresponding to the model code, and the first acquisition module 203 includes:

[0147] An adding unit, configured to add the scheduling instruction to an execution queue based on the priority;

[0148] A second obtaining unit, configured to obtain a target virtual acceleration card according to the order of the execution queue based on the required computing power information.

[0149] In one embodiment, after the creating module 204, the intelligent computing center model development device 200 for inclusive computing power further includes:

[0150] A monitoring module, configured to monitor the idle time of the first container, where the idle time is the time when the first container stops occupying the GPU computing power resource without completing the execution of the model code;

[0151] A setting module, configured to set the GPU computing power resource occupied by the first container to an idle state when the idle time is greater than or equal to a set time threshold.

[0152] In one embodiment, after the setting module, the intelligent computing center model development device 200 for inclusive computing power further includes:

[0153] A scheduling module, configured to schedule the GPU computing power resource in an idle state of the first intelligent computing center to process the model code when the first container receives a running instruction.

[0154] In one embodiment, before the first obtaining module 203, the intelligent computing center model development device 200 for inclusive computing power further includes:

[0155] A second obtaining module, configured to obtain second computing power resource information of a first intelligent computing center;

[0156] The first obtaining module 203 includes:

[0157] A third obtaining unit, configured to obtain a target virtual acceleration card according to the required computing power information when the second computing power resource information meets the required computing power information.

[0158] In one embodiment, the intelligent computing center model development device 200 for inclusive computing power further includes:

[0159] A sending module, configured to send the scheduling instruction to a second intelligent computing center when the second computing power resource information does not meet the required computing power information.

[0160] A second receiving module, configured to receive the processing result sent by the second intelligent computing center, where the processing result is obtained by the second container executing the model code, and the second container is a container created by the second intelligent computing center scheduling the GPU computing power resources in the idle state of the second intelligent computing center.

[0161] The intelligent computing center model development device for inclusive computing power provided by the present invention can implement each process of the above-mentioned intelligent computing center model development method for inclusive computing power. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0162] It should be noted that the intelligent computing center model development device for inclusive computing power in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0163] The present invention also provides an electronic device. Refer to Figure 3 , Figure 3 which is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 301, a processor 302, and a program or instruction running on the memory 301. When the program or instruction is executed by the processor 302, it can implement Figure 1 any step in the corresponding embodiment of the intelligent computing center model development method for inclusive computing power and achieve the same beneficial effects. Details will not be repeated here.

[0164] Among them, the processor 302 can be a CPU, an ASIC, an FPGA, or a GPU.

[0165] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned embodiment of the intelligent computing center model development method for inclusive computing power can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0166] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1 any step in the corresponding embodiment of the intelligent computing center model development method for inclusive computing power and achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium can be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0167] The present invention also provides a computer program product, including computer instructions, which when executed by a processor, implement the above-mentioned Figure 1The processes of each embodiment of the method for developing an intelligent computing center model corresponding to inclusive computing power are not described in detail here to avoid repetition, and the same technical effects can be achieved.

[0168] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including seven cases: A alone, B alone, C alone, A and B both present, B and C both present, A and C both present, and A, B, and C all present.

[0169] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of this application.

[0171] The embodiments of this application are described above in conjunction with the accompanying drawings, but this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all belong to the protection scope of this application.

Claims

1. A method for developing an intelligent computing center model for inclusive computing power, characterized in that: include: Step S1, receiving a scheduling instruction, wherein the scheduling instruction includes a model code; Step S2: perform computing power prediction based on the data volume of the model code to obtain required computing power information; Step S3: acquiring a target virtual accelerator card based on the required computing power information, where the target virtual accelerator card is a virtual accelerator card created based on part of the GPU computing power resources of the GPU accelerator card of the first intelligent computing center; Step S4: creating a first container based on the target virtual accelerator card, where the first container is used to execute the model code; Step S5: When the model code is executed in the first container, the first container is closed, the target virtual accelerator card is deregistered, and the GPU computing power resources of the target virtual accelerator card are released; Step S6: monitor the idle time of the first container, where the idle time is the time when the first container stops occupying GPU computing resources without completing the execution of the model code; Step S7: When the idle time is greater than or equal to a set time threshold, setting the GPU computing power resources occupied by the first container to an idle state; Step S8: When the first container receives the run instruction, scheduling the idle GPU computing resources of the first intelligent computing center to process the model code; The method further comprises: Step S9: Obtain second computing power resource information of the first intelligent computing center; Step S10: When the second computing power resource information does not meet the required computing power information, sending the scheduling instruction to the second intelligent computing center; Step S11, receiving the processing result sent by the second intelligent computing center, where the processing result is obtained by executing the model code in the second container, and the second container is a container created by the second intelligent computing center scheduling the idle GPU computing resources of the second intelligent computing center.

2. The method according to claim 1, characterized in that The step S3 comprises: Step S31: scheduling the GPU computing power resources of the target GPU accelerator card based on the required computing power information, wherein the target GPU accelerator card is an accelerator card with idle GPU computing power resources, and the idle GPU computing power resources in the target GPU accelerator card match the required computing power information; Step S32: Create the target virtual accelerator card based on the GPU computing power resources of the target GPU accelerator card.

3. The method according to claim 1, characterized in that The step S3 comprises: Step S33: acquiring a plurality of pre-created initial virtual accelerator cards and first computing power resource information corresponding to each initial virtual accelerator card, wherein the GPU computing power resources of the plurality of initial virtual accelerator cards are in an idle state; Step S34: determining a target virtual accelerator card from the multiple initial virtual accelerator cards based on the required computing power information, wherein the target virtual accelerator card is an accelerator card whose first computing power resource information meets the required computing power information.

4. The method according to claim 1, characterized in that The scheduling instruction includes the priority corresponding to the model code, and the step S3 includes: Step S35: adding the scheduling instruction to the execution queue based on the priority; Step S36: Obtain a target virtual accelerator card based on the required computing power information in the order of the execution queue.

5. The method according to any one of claims 1 to 4, characterized in that The step S3 comprises: Step S36: When the second computing power resource information meets the required computing power information, obtain a target virtual accelerator card based on the required computing power information.

6. An intelligent computing center model development device for universal computing power, characterized in that: include: A first receiving module, configured to receive a scheduling instruction, wherein the scheduling instruction includes a model code; A prediction module, used to perform computing power prediction based on the data volume of the model code to obtain required computing power information; A first acquisition module is used to acquire a target virtual accelerator card based on the required computing power information, where the target virtual accelerator card is a virtual accelerator card created based on part of the GPU computing power resources of the GPU accelerator card of the first intelligent computing center; A creation module, configured to create a first container based on the target virtual accelerator card, wherein the first container is used to execute the model code; A closing module, configured to close the first container, deregister the target virtual accelerator card, and release the GPU computing power resources of the target virtual accelerator card when the execution of the model code is completed in the first container; A monitoring module, used to monitor the idle time of the first container, where the idle time is the time when the first container stops occupying GPU computing resources without completing the execution of the model code; A setting module, configured to set the GPU computing resources occupied by the first container to an idle state when the idle time is greater than or equal to a set time threshold; A scheduling module, configured to schedule idle GPU computing resources of the first intelligent computing center to process the model code when the first container receives a run instruction; The intelligent computing center model development device for inclusive computing power also includes: A second acquisition module is used to obtain second computing power resource information of the first intelligent computing center; A sending module, configured to send the scheduling instruction to a second intelligent computing center when the second computing power resource information does not meet the required computing power information; The second receiving module is used to receive the processing result sent by the second intelligent computing center, where the processing result is obtained by executing the model code in the second container, and the second container is a container created by the second intelligent computing center to schedule the GPU computing power resources in the idle state of the second intelligent computing center.

7. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for developing an intelligent computing center model for inclusive computing power as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for developing an intelligent computing center model for inclusive computing power as described in any one of claims 1 to 5.

9. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the intelligent computing center model development method for universal computing power as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • GPU virtualization deployment method and system, computer equipment and storage medium

    CN115617364A

  • GPU task queue management method, system and device

    CN119003149A