Model Compilation Assistance Method, Device, Computer Equipment, Readable Storage Medium, and Program Product
By adding a compilation controller to the container orchestration platform, using the compilation controller to pre-run model compilation tasks on multiple GPU devices, forming a model image, solving the problem of low efficiency of large-scale model compilation and achieving efficient compilation and resource utilization.
Patent Information
- Application Number
- CN202411083596.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-08-08
AI Technical Summary
When facing large-scale deep learning models and diversified hardware architectures, the compilation and optimization process takes a long time and cannot fully utilize distributed computing resources, resulting in low compilation efficiency.
By adding a compilation controller to the container orchestration platform, the compilation controller pre-runs the target model compilation tasks in the task container on various types of GPU devices, obtain the compilation results and model weights, and save them to cloud storage to form a model image, and respond to the client's task pre-execution instructions for pre-compilation.
It significantly improves the compilation efficiency and resource utilization of the model, supports unified compilation and optimization of multiple hardware architectures, simplifies the deployment process of the model, and improves the user experience.
Smart Images

Figure CN118963725B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a model compilation assistance method, device, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of computer technology, deep learning models for implementing artificial intelligence tasks and applications have emerged. However, due to the increasing scale and complexity of deep learning models, the model compilation and optimization processes have become increasingly complex.
[0003] In traditional technologies, the model compilation and optimization processes are usually carried out on a single node. Common compilation tools include TVM and TensorRT. These tools can effectively improve the inference performance of the model on a single node. However, when faced with large-scale models and diverse hardware architectures, the compilation and optimization processes are often time-consuming and cannot fully utilize distributed computing resources, resulting in low compilation efficiency and inability to meet the requirements of efficient compilation. Summary of the Invention
[0004] Based on this, it is necessary to provide a model compilation assistance method, device, computer device, computer-readable storage medium, and computer program product for the above technical problems.
[0005] In a first aspect, this application provides a model compilation assistance method, including:
[0006] Obtain a model compilation task to be executed, and according to the task type of the model compilation task, obtain a target model compilation task that requires the use of GPU resources, and mount the target model compilation task into a task container;
[0007] Add a compilation controller in the container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices;
[0008] Save the compilation results and the model weights to cloud storage to form a corresponding model image;
[0009] In response to a task pre-execution instruction sent by a client, obtain the model image from the cloud storage, and pre-execute the target model compilation task according to the model image to obtain a pre-compiled model;
[0010] Send the pre-compiled model to the client.
[0011] In one of the embodiments, before obtaining the model compilation task to be executed, it further includes:
[0012] Deploy a compiler container in the container orchestration platform, install the compiler container on each type of the GPU devices; register the GPU devices including the compiler container into the container orchestration platform according to a preset device management mechanism.
[0013] In one embodiment, the obtaining the target model compilation task that needs to use GPU resources according to the task type of the model compilation task includes:
[0014] Judge whether the model compilation task needs to use the GPU resources according to the task type of the model compilation task, and obtain a corresponding judgment result; identify the target model compilation task that needs to use the GPU resources in the model compilation task according to the judgment result.
[0015] In one embodiment, before pre-executing the target model compilation task according to the model image to obtain a pre-compiled model, it further includes: identifying idle nodes in the container orchestration platform according to the node running conditions in the container orchestration platform;
[0016] The pre-executing the target model compilation task according to the model image to obtain a pre-compiled model includes: deploying the model image to the idle nodes in the container orchestration platform, and pre-executing the target model compilation task to obtain a task pre-execution result; obtaining the pre-compiled model according to the task pre-execution result.
[0017] In one embodiment, before sending the pre-compiled model to the client, it further includes: collecting the identity information and access permission information of the user, and verifying the identity information and the access permission information;
[0018] The sending the pre-compiled model to the client includes: sending the pre-compiled model to the client when both the identity information and the access permission information are verified.
[0019] In one embodiment, the saving the compilation result and the model weights to the cloud storage to form a corresponding model image includes:
[0020] Match the target information entry in the cloud storage according to the identification information of the target model compilation task; enter the compilation result and the model weights into the target information entry in the cloud storage, and form the model image.
[0021] In a second aspect, the present application further provides a model compilation assistance device, including:
[0022] A resource management module, configured to obtain a model compilation task to be executed, obtain a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container;
[0023] A compilation control module, configured to add a compilation controller in a container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain compilation results and model weights output by various types of the GPU devices;
[0024] An image generation module, configured to save the compilation results and the model weights to cloud storage to form a corresponding model image;
[0025] An image management module, configured to, in response to a task pre-execution instruction sent by a client, obtain the model image from the cloud storage, pre-execute the target model compilation task according to the model image to obtain a pre-compiled model;
[0026] A result sending module, configured to send the pre-compiled model to the client.
[0027] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0028] Obtain a model compilation task to be executed, obtain a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container; add a compilation controller in a container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain compilation results and model weights output by various types of the GPU devices; save the compilation results and the model weights to cloud storage to form a corresponding model image; in response to a task pre-execution instruction sent by a client, obtain the model image from the cloud storage, pre-execute the target model compilation task according to the model image to obtain a pre-compiled model; send the pre-compiled model to the client.
[0029] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0030] Obtain the model compilation task to be executed. According to the task type of the model compilation task, obtain the target model compilation task that requires the use of GPU resources, and mount the target model compilation task into the task container; Add a compilation controller in the container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices; Save the compilation results and the model weights to cloud storage to form a corresponding model image; In response to the task pre-execution instruction sent by the client, obtain the model image from the cloud storage, and pre-execute the target model compilation task according to the model image to obtain the pre-compiled model; Send the pre-compiled model to the client.
[0031] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0032] Obtain the model compilation task to be executed. According to the task type of the model compilation task, obtain the target model compilation task that requires the use of GPU resources, and mount the target model compilation task into the task container; Add a compilation controller in the container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices; Save the compilation results and the model weights to cloud storage to form a corresponding model image; In response to the task pre-execution instruction sent by the client, obtain the model image from the cloud storage, and pre-execute the target model compilation task according to the model image to obtain the pre-compiled model; Send the pre-compiled model to the client.
[0033] For the above model compilation assistance method, device, computer device, computer-readable storage medium, and computer program product, by adding a compilation controller in the container orchestration platform where the compiler container is deployed, using the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices; then saving the compilation results and model weights to cloud storage to form a corresponding model image; in response to the task pre-execution instruction sent by the client, obtaining the model image from the cloud storage and pre-executing the target model compilation task according to the model image to obtain the pre-compiled model. This solution makes full use of the resource management and scheduling capabilities of the container orchestration platform, can automatically allocate and manage and make full use of distributed computing resources, and supports unified compilation and optimization of multiple hardware architectures, thereby significantly improving the model compilation efficiency and resource utilization rate, and effectively meeting the requirements of efficient compilation. Brief Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0035] Figure 1 It is an application environment diagram of the model compilation assistance method in an embodiment;
[0036] Figure 2 It is a schematic flowchart of the model compilation assistance method in an embodiment;
[0037] Figure 3 It is a schematic flowchart of the compiler containerization deployment steps in an embodiment;
[0038] Figure 4 It is a schematic flowchart of the model compilation assistance method in a specific embodiment;
[0039] Figure 5 It is a structural block diagram of the model compilation assistance device in an embodiment;
[0040] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Embodiments
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] The model compilation assistance method provided by the embodiments of the present application can be applied to the application environment as Figure 1 shown. Among them, the client communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be integrated on the server, or placed in the cloud or other network servers.
[0043] Specifically, the model compilation assistance method provided by the embodiments of the present application can be executed by the server.
[0044] Exemplarily, the server obtains a model compilation task to be executed, obtains a target model compilation task that requires GPU resources according to the task type of the model compilation task, and mounts the target model compilation task into a task container; the server adds a compilation controller in the container orchestration platform, and uses the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of GPU devices; the server saves the compilation results and model weights to cloud storage to form a corresponding model image; in response to a task pre-execution instruction sent by the client, the server obtains the model image from cloud storage, pre-executes the target model compilation task according to the model image to obtain a pre-compiled model; the server sends the pre-compiled model to the client.
[0045] In an application environment as Figure 1 shown, the client can be but is not limited to various personal computers, laptop computers, smartphones, and tablet computers. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0046] In one embodiment, as Figure 2 shown, a model compilation assistance method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0047] Step S201, obtain a model compilation task to be executed, obtain a target model compilation task that requires GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container.
[0048] Among them, the model compilation task generally refers to converting a deep learning model into an executable and optimized form for rapid inference or execution on a target device.
[0049] Among them, the task type of the model compilation task can be classified according to whether it is a compute-intensive task. If the model compilation task requires a large amount of numerical calculations, matrix operations, and machine learning model training, etc., it can usually benefit from the parallel computing power of the GPU.
[0050] Specifically, the server obtains a model compilation task to be executed, identifies a target model compilation task that requires GPU resources in the model compilation task according to the task type of the model compilation task, mounts the target model compilation task and related file resources into a task container, and configures relevant environment variables.
[0051] Step S202: Add a compilation controller in the container orchestration platform. Use the compilation controller to pre-run the target model compilation tasks in the task containers on various types of GPU devices, and obtain the compilation results and model weights output by various types of GPU devices.
[0052] Among them, the container orchestration platform can be a Kubernetes (K8s) cluster, specifically an open-source container orchestration engine that can be used for automated deployment, scaling, and management of containerized applications.
[0053] Among them, the compilation controller can be a part of the control system, used to manage resources and tasks in the cluster to achieve specific goals or functions. In deep learning, it can refer to certain parts in the neural network model, such as the controller unit in the recurrent neural network (RNN), which is used to manage the processing of sequential data.
[0054] Specifically, the server adds a compilation controller in the container orchestration platform. Use the compilation controller to pre-run the target model compilation tasks in the task containers on various types of GPU devices, and obtain the compilation results and model weights output by various types of GPU devices in advance.
[0055] Step S203: Save the compilation results and model weights to cloud storage to form a corresponding model image.
[0056] Among them, the model image can be a lightweight and executable software package that contains all the content required for software operation, including code, runtime, libraries, and configuration files.
[0057] Specifically, the server saves the compilation results and model weights to cloud storage and makes them into a corresponding model image.
[0058] Step S204: In response to the task pre-execution instruction sent by the client, obtain the model image from cloud storage, and pre-execute the target model compilation task according to the model image to obtain a pre-compiled model.
[0059] Among them, cloud storage generally refers to storing data on remote servers connected through the Internet, rather than on local computers or traditional local storage devices. This storage method allows users to access their stored data through the network and enables access and management of this data from multiple devices.
[0060] Specifically, the server, in response to the task pre-execution instruction sent by the client, obtains the model image from cloud storage, deploys the model image to the idle nodes in the container orchestration platform, and pre-executes the target model compilation task to obtain the task pre-execution result; then, based on the task pre-execution result, obtains the pre-compiled model.
[0061] Step S205: Send the pre-compiled model to the client.
[0062] Specifically, the server sends the pre-compiled model to the client for the client to display the pre-compiled model.
[0063] In the above model compilation assistance method, by adding a compilation controller in the container orchestration platform deployed with compiler containers, and using the compilation controller to pre-run the target model compilation tasks in the task containers on various types of GPU devices, the compilation results and model weights output by various types of GPU devices are obtained; then the compilation results and model weights are saved to the cloud storage to form corresponding model images; in response to the task pre-execution instruction sent by the client, the model image is obtained from the cloud storage, and the target model compilation task is pre-executed according to the model image to obtain the pre-compiled model. This solution makes full use of the resource management and scheduling capabilities of the container orchestration platform, can automatically allocate and manage and make full use of distributed computing resources, and supports unified compilation and optimization of multiple hardware architectures, thus significantly improving the model compilation efficiency and resource utilization rate, and effectively meeting the requirements of efficient compilation.
[0064] In one embodiment, as Figure 3 shown, before obtaining the model compilation task to be executed, the method of the present application further includes the following steps:
[0065] Step S301: Deploy a compiler container in the container orchestration platform and install the compiler container on various types of GPU devices.
[0066] Step S302: Register the GPU device containing the compiler container to the container orchestration platform according to the preset device management mechanism.
[0067] Among them, the compiler container can be a software tool used to compile the trained AI model into high-performance backend code to improve the inference performance, and this container can be customized and optimized based on existing AI compiler software (such as TVM, TensorRT, etc.) to support multiple hardware architectures.
[0068] Among them, the device management mechanism can be the device plugin mechanism, which can be used to manage and allocate special hardware resources in the cluster, such as GPUs, FPGAs, and network devices, etc.
[0069] Specifically, the server deploys a compiler container in the K8s cluster and installs the compiler container on the nodes of various types of GPU devices to ensure that each node has the AI model compilation ability; then according to the deviceplugin mechanism in the K8s cluster, the GPU device containing the compiler container is registered to the K8s cluster.
[0070] In this embodiment, through the containerized deployment of the compiler and the registration of GPU devices, each node of the GPU device has the ability to compile AI models, and the K8s cluster can identify and manage these GPU devices.
[0071] In one embodiment, in the above step S201, according to the task type of the model compilation task, the target model compilation task that needs to use GPU resources is obtained, which specifically includes the following steps:
[0072] According to the task type of the model compilation task, it is judged whether the model compilation task needs to use GPU resources to obtain the corresponding judgment result; according to the judgment result, the target model compilation task that needs to use GPU resources in the model compilation task is identified.
[0073] Specifically, the server determines the task type of the model compilation task based on whether it is a compute-intensive task, and then judges whether the model compilation task needs to use GPU resources according to the task type of the model compilation task to obtain the corresponding judgment result; then according to the judgment result, the target model compilation task that needs to use GPU resources in the model compilation task is identified.
[0074] Among them, compute-intensive tasks can be large-scale data processing, scientific computing, and numerical calculations of complex models, etc., and usually benefit from the parallel computing ability of GPUs.
[0075] In this embodiment, by determining the task type of the model compilation task, the target model compilation task that needs to use GPU resources in the model compilation task is accurately identified, which is beneficial to improving the resource utilization rate of distributed computing resources.
[0076] In one embodiment, before obtaining the pre-compiled model by pre-executing the target model compilation task according to the model image, the method of the present application further includes the following steps:
[0077] According to the node running status in the container orchestration platform, the idle nodes in the container orchestration platform are identified;
[0078] In the above step S204, obtaining the pre-compiled model by pre-executing the target model compilation task according to the model image specifically includes the following steps:
[0079] Deploy the model image to the idle nodes of the container orchestration platform and pre-execute the target model compilation task to obtain the task pre-execution result; according to the task pre-execution result, obtain the pre-compiled model.
[0080] Specifically, the server identifies the idle nodes in the container orchestration platform according to the running status of the nodes in the container orchestration platform; then deploys the model image to the idle nodes in the container orchestration platform and pre-executes the target model compilation task to obtain the task pre-execution result; and then obtains the pre-compiled model according to the task pre-execution result.
[0081] In this embodiment, when the user selects to deploy the large model, the server pulls the model image file from the cloud storage and deploys it to the idle nodes scheduled by the K8s cluster; in the form of an image, the user can quickly pull the pre-compiled and optimized model, reducing the deployment time and complexity, which is beneficial to improving the model compilation efficiency.
[0082] In one of the embodiments, before sending the pre-compiled model to the client, the method of the present application further includes the following steps:
[0083] Collect the user's identity information and access permission information, and verify the identity information and access permission information.
[0084] In the above step S205, sending the pre-compiled model to the client specifically includes the following steps:
[0085] When both the identity information and the access permission information are verified to be passed, send the pre-compiled model to the client.
[0086] Among them, the user's identity information can be the user's ID number, mobile phone number, and user's name, etc.
[0087] Specifically, the server collects the user's identity information and access permission information, and verifies the identity information and access permission information; then when it is recognized that both the identity information and the access permission information are verified to be passed, send the pre-compiled model to the client.
[0088] In this embodiment, by verifying the user's identity information and access permission information and sending the pre-compiled model to the client when both the identity information and the access permission information are verified to be passed, the security and reliability of user use are improved.
[0089] In one of the embodiments, in the above step S203, saving the compilation result and the model weight to the cloud storage to form a corresponding model image specifically includes the following steps:
[0090] Match the target information entry in the cloud storage according to the identification information of the target model compilation task; enter the compilation result and the model weight into the target information entry in the cloud storage and form a model image.
[0091] Among them, an information entry generally refers to a short information unit arranged according to certain rules in a certain book, document, database, encyclopedia or other materials, similar to a title or a sub-section.
[0092] Specifically, the server matches the target information entry in the cloud storage according to the identification information of the target model compilation task; then enters the compilation result and the model weight into the target information entry in the cloud storage, and forms a model image.
[0093] In this embodiment, by matching the target information entry in the cloud storage according to the identification information of the target model compilation task, and entering the compilation result and the model weight into the target information entry in the cloud storage, the compilation result and the model weight are entered into the target information entry in the cloud storage in an orderly manner with the identification information as the division criterion, which is beneficial to the efficient storage and utilization of the model image.
[0094] In one embodiment, as Figure 4 shown, a model compilation assistance method in a specific embodiment is provided, which specifically includes the following steps:
[0095] Step S401, deploy a compiler container in the container orchestration platform, install the compiler container on various types of GPU devices; register the GPU devices containing the compiler container into the container orchestration platform according to the preset device management mechanism.
[0096] Step S402, obtain the model compilation task to be executed, judge whether the model compilation task needs to use GPU resources according to the task type of the model compilation task, and obtain the corresponding judgment result; according to the judgment result, identify the target model compilation task that needs to use GPU resources in the model compilation task, and mount the target model compilation task into the task container.
[0097] Step S403, add a compilation controller in the container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation result and the model weight output by various types of GPU devices.
[0098] Step S404, match the target information entry in the cloud storage according to the identification information of the target model compilation task; enter the compilation result and the model weight into the target information entry in the cloud storage, and form a model image.
[0099] Step S405: In response to the task pre-execution instruction sent by the client, obtain the model image from the cloud storage; identify the idle nodes in the container orchestration platform according to the node running status in the container orchestration platform; deploy the model image to the idle nodes in the container orchestration platform, and pre-execute the target model compilation task to obtain the task pre-execution result; obtain the pre-compiled model according to the task pre-execution result.
[0100] Step S406: Collect the user's identity information and access permission information, and verify the identity information and access permission information; in the case where both the identity information and the access permission information are verified, send the pre-compiled model to the client.
[0101] In the above embodiments, by introducing a compiler container and adding a compilation controller in the K8s cluster, the problem that the existing compiler is difficult to efficiently compile and optimize under multiple hardware types is solved. The specific effects include:
[0102] 1) Improve compilation efficiency: Through distributed compilation, significantly reduce the compilation time of large-scale models and improve the development and optimization efficiency.
[0103] 2) Improve resource utilization: Make full use of the resource management and scheduling capabilities of the K8s cluster to achieve efficient management and use of computing resources.
[0104] 3) Improve hardware compatibility: Support unified compilation and optimization of multiple hardware architectures, and reduce the complexity and errors in the compilation process.
[0105] 4) Simplify the deployment process: Through image management and automatic deployment of server tasks, simplify the model deployment and usage process and improve the user experience.
[0106] It should be noted that in actual applications, some parts of this technical solution can be replaced and adjusted according to specific requirements. For example:
[0107] 1) Implementation of the compiler container: It can be customized and optimized according to different compiler software (such as TVM, TensorRT, etc.) to adapt to more hardware architectures and model types.
[0108] 2) Implementation of the compilation controller: It can be dynamically adjusted according to different GPU models and user requirements to improve the flexibility and efficiency of compilation and optimization.
[0109] 3) Image management method: The generation and management method of the image can be adjusted according to specific deployment requirements to adapt to different application scenarios.
[0110] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0111] Based on the same inventive concept, an embodiment of the present application also provides a model compilation assistance device for implementing the above-mentioned model compilation assistance method. The implementation solution provided by this device for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model compilation assistance device provided below can refer to the limitations on the model compilation assistance method in the above text, and will not be repeated here.
[0112] In an exemplary embodiment, as Figure 5 shown, a model compilation assistance device is provided, including:
[0113] A resource management module 501, configured to obtain a model compilation task to be executed, obtain a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container;
[0114] A compilation control module 502, configured to add a compilation controller in a container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of GPU devices;
[0115] An image generation module 503, configured to save the compilation results and model weights to cloud storage to form a corresponding model image;
[0116] An image management module 504, configured to, in response to a task pre-execution instruction sent by a client, obtain a model image from cloud storage, pre-execute the target model compilation task according to the model image, and obtain a pre-compiled model;
[0117] A result sending module 505, configured to send the pre-compiled model to the client.
[0118] In one embodiment, the model compilation assistance device further includes a container registration module, which is used to deploy a compiler container in a container orchestration platform, install the compiler container on various types of GPU devices, and register the GPU devices containing the compiler container into the container orchestration platform according to a preset device management mechanism.
[0119] In one embodiment, the resource management module 501 is further configured to determine whether a model compilation task requires GPU resources according to the task type of the model compilation task, obtain a corresponding determination result, and identify a target model compilation task that requires GPU resources in the model compilation task according to the determination result.
[0120] In one embodiment, the model compilation assistance device further includes a node identification module, which is used to identify idle nodes in the container orchestration platform according to the node running conditions in the container orchestration platform. The image management module 504 is further configured to deploy a model image to the idle nodes in the container orchestration platform, pre-execute a target model compilation task, obtain a task pre-execution result, and obtain a pre-compiled model according to the task pre-execution result.
[0121] In one embodiment, the model compilation assistance device further includes an information verification module, which is used to collect the user's identity information and access permission information and verify the identity information and access permission information. The result sending module 505 is further configured to send the pre-compiled model to the client when both the identity information and the access permission information are verified.
[0122] In one embodiment, the image generation module 503 is further configured to match a target information entry in cloud storage according to the identification information of the target model compilation task, enter the compilation result and the model weights into the target information entry in cloud storage, and form a model image.
[0123] Each module in the above model compilation assistance device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0124] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store compilation results and model weight data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a model compilation assistance method.
[0125] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0126] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0128] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0129] It should be noted that the client information (including but not limited to client device information, client personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the client or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0130] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0132] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for assisting model compilation, characterized in that The method includes: Obtain a model compilation task to be executed, obtain a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container; Add a compilation controller in the container orchestration platform, and use the compilation controller to pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices; Save the compilation results and the model weights to cloud storage to form a corresponding model image; In response to a task pre-execution instruction sent by a client, obtain the model image from the cloud storage, identify idle nodes in the container orchestration platform according to the node running status in the container orchestration platform; deploy the model image to the idle nodes in the container orchestration platform, and pre-execute the target model compilation task to obtain a task pre-execution result; obtain the pre-compiled model according to the task pre-execution result; Send the pre-compiled model to the client for the client to display the pre-compiled model.
2. The method according to claim 1, wherein Before obtaining the model compilation task to be executed, it further includes: Deploy a compiler container in the container orchestration platform and install the compiler container on various types of the GPU devices; Register the GPU devices including the compiler container to the container orchestration platform according to a preset device management mechanism.
3. The method according to claim 1, characterized in that, The obtaining a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task includes: Judge whether the model compilation task requires the use of the GPU resources according to the task type of the model compilation task to obtain a corresponding judgment result; Identify the target model compilation task that requires the use of the GPU resources in the model compilation task according to the judgment result.
4. The method according to claim 1, wherein Before sending the pre-compiled model to the client, it further includes: Collect the identity information and access permission information of the user, and verify the identity information and the access permission information; The sending the pre-compiled model to the client includes: Send the pre-compiled model to the client when both the identity information and the access permission information are verified.
5. The method according to any one of claims 1 to 4, characterized in that, The saving the compilation results and the model weights to cloud storage to form a corresponding model image includes: Match a target information entry in the cloud storage according to the identification information of the target model compilation task; Enter the compilation results and the model weights into the target information entry of the cloud storage and form the model image.
6. A model compilation assistance device, characterized in that, The device includes: A resource management module, configured to obtain a model compilation task to be executed, obtain a target model compilation task that requires the use of GPU resources according to the task type of the model compilation task, and mount the target model compilation task into a task container; A compilation control module, configured to add a compilation controller in a container orchestration platform, and by using the compilation controller, pre-run the target model compilation task in the task container on various types of GPU devices to obtain the compilation results and model weights output by various types of the GPU devices; An image generation module, configured to save the compilation results and the model weights to cloud storage to form a corresponding model image; An image management module, configured to, in response to a task pre-execution instruction sent by a client, obtain the model image from the cloud storage, identify idle nodes in the container orchestration platform according to the node running conditions in the container orchestration platform; deploy the model image to the idle nodes in the container orchestration platform, and pre-execute the target model compilation task to obtain a task pre-execution result; and obtain the pre-compiled model according to the task pre-execution result; A result sending module, configured to send the pre-compiled model to the client for the client to display the pre-compiled model.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Mirror image construction method, system and device and medium
CN112363803A
Deep learning distributed compiler for cloud edge computing and construction method
CN113127203A