Computing Power Scheduling Method and Device for the Development Status of Inclusive Computing Power Intelligent Computing Center
By creating a virtual GPU container in the intelligent computing center to adjust the model code and using computing resources to debug when needed, the problems of low computing resources utilization and high economic costs are solved, and more efficient resource utilization and cost reduction are achieved.
Patent Information
- Application Number
- CN202510302247.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, the computing power resource utilization rate of intelligent computing centers is low and the economic cost of model development is high, resulting in limited application of universal computing power.
Release unused computing resources by creating a virtual graphics processor GPU container and adjusting the model code and debugging it using the computing resources of the Intelligent Computing Center when needed.
It improves the computing power resource utilization rate of the intelligent computing center, reduces the economic cost of model development, and realizes the widespread application of universal computing power.
Smart Images

Figure CN119806851B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power infrastructure, and particularly to a computing power scheduling method and device for the development status of an inclusive computing power intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly to provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.
[0007] Currently, users use the computing power resources provided by intelligent computing centers to achieve model development. During the model development process, the model code needs to be continuously modified and adjusted. In this process, the computing power resources of the intelligent computing center are not required during the stage of modifying the model code. However, in the prior art, since the model is deployed in the intelligent computing center, the computing power resources of the intelligent computing center still need to be occupied during the modified code model, resulting in the long-term occupation of the computing power resources of the intelligent computing center and very low utilization rate of the computing power resources. At the same time, users need to lease the computing power services of the intelligent computing center for a long time, resulting in a very high economic cost for model development and making it difficult to achieve the wide application of inclusive computing power.
[0008] It can be seen that there are problems in the prior art such as very low utilization rate of computing power resources and very high economic cost for model development. Summary of the Invention
[0009] The present invention provides a computing power scheduling method and device for the development status of an inclusive computing power intelligent computing center, so as to solve the problems of low utilization rate of computing power resources and high economic cost of model development in the prior art.
[0010] To solve the above problems, the present invention is implemented as follows:
[0011] In a first aspect, the present invention provides a computing power scheduling method for the development status of an inclusive computing power intelligent computing center, including:
[0012] Step S1, when receiving a development request sent by a user, create a first container based on virtual graphics processing unit (GPU) resources, where the first container is used to adjust model code;
[0013] Step S2, send a debugging instruction to a second container based on the first container, and run the debugging instruction based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions corresponding to the adjusted model code, and the second container is a container created based on the computing power resources of the intelligent computing center.
[0014] In one embodiment, step S2 includes:
[0015] Step S21, when there is no running second container in the intelligent computing center, call a container operation application programming interface (API) to create a second container based on the computing power resources of the intelligent computing center;
[0016] Step S22, establish a communication channel between the first container and the second container;
[0017] Step S23, send the debugging instruction to the second container through the communication channel based on the first container;
[0018] Step S24, run the debugging instruction based on the second container to obtain a debugging result.
[0019] In one embodiment, step S2 includes:
[0020] Step S25, when there is a running second container in the intelligent computing center, send the debugging instruction to the second container through the communication channel based on the first container;
[0021] Step S26, run the debugging instruction based on the second container to obtain a debugging result.
[0022] In one embodiment, after step S2, the method further includes:
[0023] Step S3: Send the debugging result to the first container through the communication channel based on the second container.
[0024] In one embodiment, after the step S3, the method further includes:
[0025] Step S4: Close the second container and release the computing power resources occupied by the second container.
[0026] In one embodiment, the step S4 includes at least one of the following:
[0027] Step S41: When the cumulative time reaches the preset time threshold and the second container does not receive a new debugging instruction, close the second container and release the computing power resources occupied by the second container. The cumulative time is the time after the second container last received a debugging instruction.
[0028] Step S42: When the cumulative time does not reach the preset time threshold and the second container receives a new debugging instruction, recalculate the cumulative time.
[0029] In a second aspect, the present invention further provides a computing power scheduling device for the development status of an inclusive computing power intelligent computing center, including:
[0030] A creation module, configured to create a first container based on virtual graphics processing unit (GPU) resources when receiving a development request sent by a user. The first container is used to adjust model code.
[0031] A debugging module, configured to send a debugging instruction to a second container based on the first container and run the debugging instruction based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions corresponding to running the adjusted model code. The second container is a container created based on the computing power resources of the intelligent computing center.
[0032] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for scheduling computing power in the development status of an inclusive computing power intelligent computing center as described in the first aspect above.
[0033] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for scheduling computing power in the development status of an inclusive computing power intelligent computing center as described in the first aspect above.
[0034] Fifth aspect, the present invention further provides a computer program product, including computer instructions, which when executed by a processor, implement the steps in the development status computing power scheduling method for the inclusive computing power intelligent computing center as described in the first aspect above.
[0035] In the present invention, in the case of receiving a development request sent by a user, a first container is created based on virtual graphics processing unit (GPU) resources, and the first container is used to adjust model code; a debugging instruction is sent to a second container based on the first container, and the debugging instruction is run based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions corresponding to the adjusted model code, and the second container is a container created based on the computing power resources of the intelligent computing center. In this way, by adjusting the model code through the first container, the computing power resources of the intelligent computing center are not occupied, and when debugging is required, the adjusted model code is run based on the computing power resources of the intelligent computing center through the second container, thereby achieving the occupation of the computing power resources of the intelligent computing center only when the adjusted model code needs to be run, and greatly improving the utilization rate of the computing power resources of the intelligent computing center. Further, since there is no need for the user to rent the computing power service of the intelligent computing center for a long time, the economic cost of the user for model development is effectively reduced significantly, and the wide application of inclusive computing power is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0037] Figure 1 is a flowchart of a development status computing power scheduling method for an inclusive computing power intelligent computing center provided by the present invention;
[0038] Figure 2 is an interaction schematic diagram of the first container and the second container provided by the present invention;
[0039] Figure 3 is a structural diagram of a development status computing power scheduling device for an inclusive computing power intelligent computing center provided by the present invention;
[0040] Figure 4 is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] The "computing power" as referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0043] The "computational power" (Computational Power, CP) as referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .
[0044] The "carrying capacity" (Network Power, NP) as referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive indicator for measuring network transmission scheduling ability.
[0045] The "storage power" (Storage Power, SP) as referred to in the present invention means: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0046] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.
[0047] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general-purpose technologies, the form of the new type of information infrastructure will be more diverse.
[0048] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0049] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0050] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0051] The "super computing power" described in the present invention mainly refers to: the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0052] The "intelligent computing center" described in the present invention refers to: a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0053] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0054] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0055] The "Computing Power Center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0056] The "Supercomputing Center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0057] The "Computing Power Resources" described in the present invention refer to: technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as support and guarantee resources such as wind, fire, water, and electricity.
[0058] The "Development Status" described in the present invention refers to the status during the process of writing model code.
[0059] The "Model" described in the present invention includes but is not limited to "Large Language Model" and "Multimodal Large Model".
[0060] The "Large Language Model" described in the present invention refers to the large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0061] The "Multimodal Large Model" described in the present invention (Multimodal Large Models) refers to: a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.
[0062] Please refer to Figure 1 , Figure 1 which is a flowchart of a computing power scheduling method for the development status of an intelligent computing center for inclusive computing power provided by the present invention. As Figure 1 shown, it includes the following steps:
[0063] Step S1: When a development request sent by a user is received, create a first container based on virtual Graphics Processing Unit (GPU) resources, where the first container is used to adjust model code.
[0064] The above virtual GPU resources are virtual resources that simulate GPUs. The first container created based on virtual GPU resources does not occupy the real GPU resources of the intelligent computing center, but can simulate the operating environment of real GPU resources. Users can adjust model code through the first container, and do not need to occupy the computing power resources of the intelligent computing center during the entire adjustment process, avoiding waste of computing power.
[0065] The above development request is a request sent when a user needs to perform model development or adjustment, and can be sent by the terminal used by the user. For example, the user selects in the used terminal to adjust the model in the development state. At this time, the terminal generates a development request and sends the development request to the intelligent computing center; after receiving the development request, the intelligent computing center creates a first container based on virtual GPU resources, enabling the user to adjust the code of the model through the first container.
[0066] Among them, the development request is used to represent the current user's need to develop and adjust model code.
[0067] In some alternative embodiments, the first container can be a container (Pod) that starts lightweight development based on Kubernetes (k8s), with an integrated development environment (IDE) and debugging tools built in, and can simulate GPU resources through a Resource Fake mechanism, thereby obtaining virtual GPU resources.
[0068] Specifically, by modifying the Device Plugin interface of Kubernetes, report virtual GPU resources (such as nvidia.com / fake - gpu:1) to the scheduler, notify the scheduler to allocate "virtual resources", and actually do not occupy real GPU resources. During the process of creating the first container, pre - install the same CUDA driver version and dependency libraries (such as cuDNN) as the real GPU container through the container image to ensure seamless compatibility between the development environment and the operating environment and maintain the consistency of the operating environment.
[0069] Furthermore, the user can select different modes according to their needs, and the terminal sends different requests according to different modes. For example, in the case where the user selects the development mode, the terminal sends a development request to the intelligent computing center; in the case where the user selects the running mode (the running mode indicates that the user needs to repeatedly run the model), the terminal sends a running request to the intelligent computing center. At this time, the intelligent computing center creates a container based on the actual computing power resources, and runs the model based on the computing power resources of the intelligent computing center through the container.
[0070] Step S2: Send a debugging instruction to the second container based on the first container, and run the debugging instruction based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions for running the adjusted model code. The second container is a container created based on the computing power resources of the intelligent computing center.
[0071] The above-mentioned second container is a container created based on the computing power resources of the intelligent computing center. Through the second container, it is possible to run the adjusted model code by using the computing power resources of the intelligent computing center, so as to confirm the effect after the model is adjusted. Among them, the second container occupies the computing power resources of the intelligent computing center when running, and does not occupy the computing power resources of the intelligent computing center when not running, thus avoiding occupying the computing power resources of the intelligent computing center without using them, and improving the utilization rate of the computing power resources of the intelligent computing center.
[0072] The above-mentioned debugging instruction is an instruction sent by the user when they need to debug the adjusted model code after completing the adjustment of the model code. Sending the debugging instruction from the first container to the second container enables the adjusted model code to be run through the second container when the first container does not occupy the computing power resources of the intelligent computing center, achieving a reduction in the occupation of the computing power resources of the intelligent computing center.
[0073] Among them, the debugging instruction includes GPU instructions corresponding to running the adjusted model code, which are specifically sent through the virtual GPU service of the first container. The second container executes the GPU instructions based on the computing power resources of the intelligent computing center, and thus can run the adjusted model code.
[0074] In some alternative embodiments, the debugging instruction may also include the modified model code. After receiving the debugging instruction, the second container can run the modified model code to obtain a debugging result.
[0075] Among them, the modified model code can be full-scale code or incremental code.
[0076] In the present invention, in the case of receiving a development request sent by a user, a first container is created based on virtual graphics processing unit (GPU) resources, and the first container is used to adjust model code; a debugging instruction is sent to a second container based on the first container, and the debugging instruction is run based on the second container to obtain a debugging result, where the debugging instruction includes GPU instructions corresponding to the adjusted model code, and the second container is a container created based on the computing power resources of an intelligent computing center. In this way, the model code is adjusted by the first container without occupying the computing power resources of the intelligent computing center, and when debugging is required, the adjusted model code is run based on the computing power resources of the intelligent computing center by the second container, thereby realizing occupying the computing power resources of the intelligent computing center only when the adjusted model code needs to be run, and improving the utilization rate of the computing power resources of the intelligent computing center. Further, since there is no need for the user to lease the computing power service of the intelligent computing center for a long time, the economic cost of the user for model development is effectively and significantly reduced, and the wide application of inclusive computing power is realized.
[0077] It should be noted that there may or may not be a running second container in the intelligent computing center. In the case where there is a running second container, the second container can directly execute receiving and running the debugging instruction; while in the case where there is no running second container, the second container needs to be started to realize receiving and running the debugging instruction through the second container.
[0078] Specifically, step S2 includes:
[0079] Step S21: In the case where there is no running second container in the intelligent computing center, a second container is created based on the computing power resources of the intelligent computing center by calling a container operation application programming interface (API).
[0080] Step S22: Establish a communication channel between the first container and the second container.
[0081] Step S23: Based on the first container, send the debugging instruction to the second container through the communication channel.
[0082] Step S24: Run the debugging instruction based on the second container to obtain a debugging result.
[0083] In this case, there is no running second container in the intelligent computing center, and at this time, it is necessary to use the computing power resources of the intelligent computing center to realize receiving and running the debugging instruction through the second container. Among them, in order to realize data transmission between the first container and the second container, it is necessary to establish a communication channel between the first container and the second container.
[0084] In some alternative embodiments, the second container includes a virtual Compute Unified Device Architecture (CUDA) interface, and the virtual CUDA interface is used to establish a communication channel between the first container and the second container.
[0085] It should be noted that the virtual CUDA interface is a functional component added when creating the second container. Through the virtual CUDA interface, interception of debug instructions can be achieved, and CUDA API calls in the code (such as cudaMalloc, cudaMemcpy, etc.) can be implemented.
[0086] Specifically, as Figure 2 shown, in the case where there is no running second container in the intelligent computing center, the container operation application programming interface (API) is called through the container operation instruction orchestration (Operator) of Kubernetes (k8s), and the second container is created based on the computing power resources of the intelligent computing center. A communication channel is established between the second container and the first container, so that the second container can receive debug instructions through the communication channel.
[0087] Among them, the second container includes a virtual CUDA (rCUDA) interface, and the receipt of debug instructions is realized through the rCUDA interface. Further, the second container further includes a CUDA layer and a driver layer to implement the invocation of CUDA functions and the invocation of the driver.
[0088] In some alternative embodiments, the debug instruction is specifically a Remote Procedure Call Instruction (RPC Instruction). By sending the remote instruction packet to the operation component of k8s, the creation of the second container can be realized.
[0089] Further, after the container operation instruction orchestration receives the debug instruction, the container operation instruction orchestration selects the optimal GPU resource to start the second container based on the priority queue and resource pool status of the GPU.
[0090] Specifically, after the first container sends a debug instruction, the GPU list of the intelligent computing center is obtained based on the container operation instruction orchestration. The GPU list includes multiple GPUs, and the multiple GPUs in the GPU list are arranged according to the priority of the debug model;
[0091] Obtain the resource usage ratio of each GPU in the multiple GPUs;
[0092] Based on the resource usage ratio and priority, calculate the evaluation score of each GPU.
[0093] Determine the target GPU based on the resource usage status, where the target GPU is the GPU with the highest evaluation score among the multiple GPUs.
[0094] In some embodiments, the evaluation score is obtained by weighted calculation of the resource usage ratio and the priority.
[0095] In this way, by selecting the GPU with the highest evaluation score to create the second container, the second container can execute the debugging instructions quickly.
[0096] In some embodiments, different second containers in the intelligent computing center are bound to different GPUs to avoid interference between different tasks.
[0097] In some alternative embodiments, the debugging instructions use binary encoding (such as Protocol Buffers) to compress the instruction data, reducing the network transmission delay.
[0098] In the present invention, in the case that there is no running second container in the intelligent computing center, based on the computing power resources of the intelligent computing center, call the container operation application programming interface API to create a second container; establish a communication channel between the first container and the second container; based on the first container, send the debugging instructions to the second container through the communication channel; based on the second container, run the debugging instructions to obtain a debugging result. In this way, in the case that there is no running second container in the intelligent computing center, by calling the container operation API interface, a second container is created, and a communication channel between the first container and the second container is created, so that the second container can receive and execute the debugging instructions through the communication channel.
[0099] In some alternative embodiments, calling the container operation application programming interface API to create a second container based on the computing power resources of the intelligent computing center includes:
[0100] Obtain historical load data, where the historical load data is used to characterize the load situation of the same type of model during debugging;
[0101] Create the second container based on the historical load data.
[0102] In this way, by creating the second container through the historical load data, the computing power resources of the intelligent computing center occupied by the second container are close to the computing power resources of the real demand, further reducing the consumption of the computing power resources of the intelligent computing center and improving the utilization rate of the computing power resources of the intelligent computing center.
[0103] In one embodiment, the step S2 includes:
[0104] Step S25: When the second container is running in the intelligent computing center, send the debugging instruction to the second container through the communication channel based on the first container;
[0105] Step S26: Run the debugging instruction based on the second container to obtain a debugging result.
[0106] In the present invention, when the second container is running in the intelligent computing center, the debugging instruction is sent to the second container through the communication channel based on the first container; the debugging instruction is run based on the second container to obtain a debugging result. In this way, when the second container is already running in the intelligent computing center, the debugging instruction is directly received and executed by the existing second container, thereby achieving low-latency model debugging.
[0107] In one embodiment, after the step S2, the method further includes:
[0108] Step S3: Send the debugging result to the first container through the communication channel based on the second container.
[0109] In the present invention, after the second container runs the debugging instruction to obtain a debugging result, the debugging result is sent to the first container through the communication channel, so that the user can determine the effect after the model code is adjusted through the first container, facilitating the user to further adjust the model code through the first container.
[0110] In one embodiment, after the step S3, the method further includes:
[0111] Step S4: Shut down the second container and release the computing power resources occupied by the second container.
[0112] In the present invention, the second container needs to occupy the computing power resources of the intelligent computing center. To reduce the occupation of the computing power resources of the intelligent computing center, the second container is shut down when the second container is not required to run the code, and the computing power resources occupied by the second container are released to further improve the utilization rate of the computing power resources, thereby realizing the wide application of inclusive computing power.
[0113] In some alternative embodiments, the CUDA API call chain is monitored through container operation instruction orchestration. When an end signal (such as cudaDeviceReset) is monitored, the second container is shut down and the GPU resources are released to reduce the occupation of the computing power resources of the intelligent computing center by the second container for a long time.
[0114] In some alternative embodiments, a timeout threshold is set. When the container operation instruction orchestration monitors that the second container has not processed data within the time of the timeout threshold, the second container is forcibly closed to prevent the abnormal situation from causing the second container to occupy the computing power resources of the intelligent computing center for a long time.
[0115] In some embodiments, step S4 includes at least one of the following:
[0116] Step S41: When the cumulative time reaches the preset time threshold and the second container has not received a new debugging instruction, close the second container and release the computing power resources occupied by the second container. The cumulative time is the time after the second container last received a debugging instruction.
[0117] Step S42: When the cumulative time has not reached the preset time threshold and the second container receives a new debugging instruction, recalculate the cumulative time.
[0118] It should be noted that creating or starting the second container requires a certain amount of time and computing power resources. When the user needs to frequently run and adjust the model code, repeatedly closing and starting the second container not only consumes computing power resources but also results in a relatively high latency for debugging the model code. To solve this problem, a preset time threshold is set in the present invention to optimize the startup and shutdown of the second container through the preset time threshold.
[0119] Specifically, when the cumulative time reaches the preset time threshold and the second container has not received a new debugging instruction, close the second container and release the computing power resources occupied by the second container. The cumulative time is the time after the second container last received a debugging instruction. When the cumulative time has not reached the preset time threshold and the second container receives a new debugging instruction, recalculate the cumulative time. In this way, when the sending time between two adjacent debugging instructions is less than or equal to the preset time threshold, the second container will not be closed and can directly process the debugging instruction. When the sending time between two adjacent debugging instructions is greater than the preset time threshold, close the second container to reduce the occupation of the computing power resources of the intelligent computing center by the second container.
[0120] Please refer to Figure 3 , Figure 3 which is a structural diagram of a computing power scheduling device for the development state of an intelligent computing center for inclusive computing power provided by the present invention. As Figure 3 shown, the computing power scheduling device 300 for the development state of an intelligent computing center for inclusive computing power includes:
[0121] A creation module 301, configured to create a first container based on virtual graphics processing unit (GPU) resources when receiving a development request sent by a user. The first container is used to adjust model code.
[0122] A debugging module 302, configured to send a debugging instruction to a second container based on the first container, and run the debugging instruction based on the second container to obtain a debugging result, where the debugging instruction includes a GPU instruction corresponding to the adjusted model code, and the second container is a container created based on the computing power resources of an intelligent computing center.
[0123] In one embodiment, the debugging module 302 includes:
[0124] A creation unit, configured to call a container operation application programming interface (API) to create a second container based on the computing power resources of the intelligent computing center when the second container is not running in the intelligent computing center;
[0125] An establishment unit, configured to establish a communication channel between the first container and the second container;
[0126] A first sending unit, configured to send the debugging instruction to the second container based on the first container through the communication channel;
[0127] A first debugging unit, configured to run the debugging instruction based on the second container to obtain a debugging result.
[0128] In one embodiment, the debugging module 302 includes:
[0129] A second sending unit, configured to send the debugging instruction to the second container based on the first container through the communication channel when the second container is running in the intelligent computing center;
[0130] A second debugging unit, configured to run the debugging instruction based on the second container to obtain a debugging result.
[0131] In one embodiment, after the debugging module 302, the state computing power scheduling device 300 for the development of the inclusive computing power intelligent computing center further includes:
[0132] A sending module, configured to send the debugging result to the first container based on the second container through the communication channel.
[0133] In one embodiment, after the sending module, the state computing power scheduling device 300 for the development of the inclusive computing power intelligent computing center further includes:
[0134] A closing module, configured to close the second container and release the computing power resources occupied by the second container.
[0135] In one embodiment, the closing module includes at least one of the following:
[0136] A first closing unit, configured to close the second container and release the computing power resources occupied by the second container when the accumulated time reaches a preset time threshold and the second container does not receive a new debugging instruction, where the accumulated time is the time after the second container last received a debugging instruction;
[0137] A second closing unit, configured to recalculate the accumulated time when the second container receives a new debugging instruction while the accumulated time does not reach the preset time threshold.
[0138] The computing power scheduling device for the development status of the inclusive computing power intelligent computing center provided by the present invention can implement each process of the above-mentioned computing power scheduling method for the development status of the inclusive computing power intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0139] It should be noted that the computing power scheduling device for the development status of the inclusive computing power intelligent computing center in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.
[0140] The present invention also provides an electronic device. Refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 401, a processor 402, and a program or instruction stored on the memory 401 and running. When the program or instruction is executed by the processor 402, it can implement Figure 1 any step in the corresponding embodiment of the computing power scheduling method for the development status of the inclusive computing power intelligent computing center and achieve the same beneficial effects, which will not be elaborated here.
[0141] Among them, the processor 402 can be a CPU, an ASIC, an FPGA, or a GPU.
[0142] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned embodiment of the computing power scheduling method for the development status of the inclusive computing power intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0143] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1Any step in the corresponding embodiment of the computing power scheduling method for the development status of the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. The storage medium, such as Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disc, etc.
[0144] The present invention also provides a computer program product, including computer instructions, which when executed by a processor implement the above Figure 1 Each process of the corresponding embodiment of the computing power scheduling method for the development status of the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0145] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, and A and B both exist, B and C both exist, A and C both exist, and A, B, and C all exist, a total of 7 situations.
[0146] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of the present application.
[0148] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A computing power scheduling method for the development status of an inclusive computing power intelligent computing center, characterized in that, Including: Step S1: When receiving a development request sent by a user, create a first container based on virtual graphics processing unit (GPU) resources, where the first container is used to adjust model code; Step S2: Send a debugging instruction to a second container based on the first container, and run the debugging instruction based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions for running the adjusted model code, and the second container is a container created based on the computing power resources of an intelligent computing center; The step S2 includes: Step S21: When there is no running second container in the intelligent computing center, call the container operation application programming interface (API) to create a second container based on the computing power resources of the intelligent computing center; Step S22: Establish a communication channel between the first container and the second container; Step S23: Send the debugging instruction to the second container through the communication channel based on the first container; Step S24: Run the debugging instruction based on the second container to obtain a debugging result; After the step S2, the method further includes: Step S3: Send the debugging result to the first container through the communication channel based on the second container; Step S4: Close the second container and release the computing power resources occupied by the second container.
2. The method according to claim 1, characterized in that, The step S2 includes: Step S25: When there is a running second container in the intelligent computing center, send the debugging instruction to the second container through the communication channel based on the first container; Step S26: Run the debugging instruction based on the second container to obtain a debugging result.
3. The method according to claim 1, characterized in that, The step S4 includes at least one of the following: Step S41: When the cumulative time reaches a preset time threshold and the second container does not receive a new debugging instruction, close the second container and release the computing power resources occupied by the second container. The cumulative time is the time after the second container last received a debugging instruction; Step S42: When the cumulative time does not reach the preset time threshold and the second container receives a new debugging instruction, recalculate the cumulative time.
4. A computing power scheduling device for the development status of an inclusive computing power intelligent computing center, characterized in that, Including: A creation module, configured to create a first container based on virtual graphics processing unit (GPU) resources when receiving a development request sent by a user, where the first container is used to adjust model code; A debugging module, configured to send a debugging instruction to a second container based on the first container, and run the debugging instruction based on the second container to obtain a debugging result. The debugging instruction includes GPU instructions for running the adjusted model code, and the second container is a container created based on the computing power resources of an intelligent computing center; The debugging module includes: A creation unit, configured to call the container operation application programming interface (API) to create a second container based on the computing power resources of the intelligent computing center when there is no running second container in the intelligent computing center; An establishment unit, configured to establish a communication channel between the first container and the second container; A first sending unit, configured to send the debugging instruction to the second container through the communication channel based on the first container; A first debugging unit, configured to run the debugging instruction based on the second container to obtain a debugging result; A sending module, configured to send the debugging result to the first container through the communication channel based on the second container; A closing module, configured to close the second container and release the computing power resources occupied by the second container.
5. An electronic device, characterized in that, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the development state computing power scheduling method for the inclusive computing power intelligent computing center as described in any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the development state computing power scheduling method for the inclusive computing power intelligent computing center as described in any one of claims 1 to 3 are implemented.
7. A computer program product, characterized in that, Comprising computer instructions, and when the computer instructions are executed by a processor, the steps of the development state computing power scheduling method for the inclusive computing power intelligent computing center as described in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Processing method and device, electronic equipment and readable storage medium
CN113296950A
Intelligent computing center model development method and device oriented to popularity
CN119336517A