A method, device, computer device, and storage medium for deploying an inference service
By separating the inference framework and model, removing the running parameters and customizing the startup parameters, the problem of low mirror reuse rate of different model files under the same inference framework is solved, and more efficient deployment and reuse rate is achieved, improving the deployment efficiency of inference services.
Patent Information
- Application Number
- CN202210602553.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-30
AI Technical Summary
In actual production, when different model files under the same inference framework, the mirror multiplexing rate is low and the transmission and deployment are slow.
Remove the running parameters of each inference framework and create a mirror file separately; create a model resource library; determine the target model and the target inference framework based on the inference service to be deployed; customize the running parameters for the target inference framework; load the image file and start the inference framework; pull and run the target model file from the model resource library under the startup framework.
The image size of the deployment of inference services is reduced, the reuse rate of inference frameworks and model files is improved, the inference frameworks have strong flexibility, the difficulty of combining different inference frameworks and different models is reduced, the deployment methods of inference services are enriched, and the deployment efficiency is significantly improved.
Smart Images

Figure CN114819160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technologies, and particularly to an inference service deployment method, an inference service deployment device, a computer device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence, more and more training frameworks and inference frameworks have come into people's sight, such as the familiar Caffe framework, TensorFlow framework, pytorch framework, etc. Sometimes in specific scenarios, in order to achieve extreme performance, users will optimize based on the original framework. In these cases, more and more different frameworks will appear.
[0003] Currently, the main way to deploy an inference service is to combine the model with the framework, that is, to package the two into an image package containing the framework (including framework running parameters) and the model. Although this method can provide a unified deployment method for different frameworks, the disadvantages of this method are also obvious. Specifically, there are the following defects in the traditional deployment of inference services: First, packaging the model and the framework into one image, firstly, the reuse rate of this method is low. When using the same framework with different models, the same framework cannot be reused, and separate images need to be made for each case. In actual production, the size of the inference service image can reach several gigabytes or even more than ten gigabytes. It can be seen that using this method to make, transfer, and deploy images is very time-consuming and space-consuming. Second, separating the inference framework from the model, and having the framework service pull different models during actual operation, this method solves the problem of low reuse rate of some images, but brings the problem that the startup parameters and configuration information of the framework service cannot be changed. Summary of the Invention
[0004] In view of this, in order to solve the situation of low image reuse rate, slow transmission, and slow deployment when there are different model files under the same inference framework in actual production, the present invention provides an inference service deployment method, an inference service deployment device, a computer device, and a computer-readable storage medium.
[0005] According to a first aspect of the present invention, there is provided an inference service deployment method, the inference service deployment method including:
[0006] Removing the running parameters of each inference framework, and respectively making corresponding image files for each inference framework after removing the running parameters;
[0007] Creating a model resource library including a number of model files;
[0008] Based on the inference service to be deployed, determining a target model and a target inference framework;
[0009] Customize the running parameters for the target inference framework;
[0010] Load the image file corresponding to the target inference framework and start the target inference framework based on the customized running parameters;
[0011] Under the started target inference framework, pull and run the model file corresponding to the target model from the model resource library.
[0012] In some embodiments, the inference service deployment method further includes:
[0013] Customize the pull protocol for the target model; and
[0014] When performing the step of pulling and running the model file corresponding to the target model from the model resource library under the started target inference framework, use the customized pull protocol to pull the model file.
[0015] In some embodiments, the customized pull protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
[0016] In some embodiments, each inference framework includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
[0017] According to the second aspect of the present invention, there is provided an inference service deployment device, and the inference service deployment device includes:
[0018] An image making module configured to remove the running parameters of each inference framework and make corresponding image files for each inference framework after removing the running parameters;
[0019] A creation module configured to create a model resource library including a number of model files;
[0020] A determination module configured to determine a target model and a target inference framework based on the inference service to be deployed;
[0021] A first customization module configured to customize the running parameters for the target inference framework;
[0022] A framework start module configured to load the image file corresponding to the target inference framework and start the target inference framework based on the customized running parameters;
[0023] A model pulling module, configured to pull and run a model file corresponding to the target model from the model resource library under the target inference framework after startup.
[0024] In some embodiments, the inference service deployment device further includes:
[0025] A second custom module, configured to customize a pulling protocol for the target model; and
[0026] The model pulling module pulls the model file using the custom pulling protocol.
[0027] In some embodiments, the custom pulling protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
[0028] In some embodiments, each of the inference frameworks includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
[0029] According to a third aspect of the present invention, there is also provided a computer device, which includes:
[0030] At least one processor; and
[0031] A memory, storing a computer program that can run on the processor, and the processor executes the program to perform the foregoing inference service deployment method when executing the program.
[0032] According to a fourth aspect of the present invention, there is also provided a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, it performs the foregoing inference service deployment method.
[0033] For the foregoing inference service deployment method, the inference framework and the model required for deploying the inference service are separated. At the same time, the running parameters of the inference framework are removed from the original framework, and an image is created only for the inference framework after removing the running parameters. The inference framework is started by using a separate custom framework to run the parameters, and the model file is pulled from the model resource library using the started framework. This not only reduces the size of the image for deploying the inference service, but also improves the reuse rate of the inference framework and the model file, making the inference framework highly flexible, reducing the difficulty of combining different inference frameworks and different models, enriching the deployment methods of the inference service, and significantly improving the deployment efficiency of the inference service.
[0034] In addition, the present invention also provides an inference service deployment device, a computer device, and a computer-readable storage medium, which can also achieve the above technical effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of a method for deploying an inference service provided by an embodiment of the present invention;
[0037] Figure 2 It is a schematic architecture diagram of an inference service provided by another embodiment of the present invention;
[0038] Figure 3 It is a schematic flowchart of deploying an inference service in Kubernetes provided by another embodiment of the present invention
[0039] Figure 4 It is a schematic structural diagram of an inference service deployment device provided by another embodiment of the present invention;
[0040] Figure 5 The internal structure diagram of a computer device in another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings.
[0042] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two entities or parameters with the same name but different, and it can be seen that "first" and "second" are only for the convenience of expression and should not be construed as a limitation to the embodiments of the present invention. This will not be elaborated one by one in the subsequent embodiments.
[0043] In one embodiment, please refer to Figure 1 As shown, the present invention provides a method for deploying an inference service. Specifically, the method for deploying an inference service includes the following steps:
[0044] Step 101: Remove the running parameters of each inference framework, and respectively create corresponding image files for each inference framework after removing the running parameters; wherein, the inference framework refers to the processing flow for solving a certain problem;
[0045] Step 102: Create a model resource library including several model files; wherein, the model refers to a trained neural network, and this model can be applied to actual tasks.
[0046] Step 103, determine the target model and the target inference framework based on the inference service to be deployed;
[0047] Step 104, customize the running parameters for the target inference framework;
[0048] Step 105, load the mirror file corresponding to the target inference framework and start the target inference framework based on the customized running parameters;
[0049] Step 106, pull and run the model file corresponding to the target model from the model repository under the started target inference framework.
[0050] The above method for deploying an inference service separates the inference framework and the model required for deploying the inference service. At the same time, the running parameters of the inference framework are removed from the original framework, and only the inference framework after removing the running parameters is used to create a mirror. The inference framework is started by the method of separately customizing the framework running parameters. When using the started framework to pull the model file from the model repository, it not only reduces the size of the mirror for deploying the inference service, but also improves the reuse rate of the inference framework and the model file, making the inference framework more flexible, reducing the difficulty of combining different inference frameworks and different models, enriching the deployment methods of the inference service, and significantly improving the deployment efficiency of the inference service.
[0051] In some instances, the method for deploying the inference service further includes:
[0052] Customize the pull protocol for the target model; and
[0053] When performing the step of pulling and running the model file corresponding to the target model from the model repository under the started target inference framework, use the customized pull protocol to pull the model file.
[0054] In some embodiments, the customized pull protocol is any one of the Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Hadoop Distributed FileSystem Protocol (HDFS), and PersistentVolumeClaim Protocol (PVC).
[0055] In some embodiments, each inference framework includes at least one of the Caffe framework, the TensorFlow framework, and the PyTorch framework.
[0056] Caffe, short for Convolutional Architecture for Fast Feature Embedding, is a deep learning framework that combines expressiveness, speed, and modularity of thinking. Caffe supports multiple types of deep learning architectures, is oriented towards image classification and image segmentation, and also supports the design of CNN, RCNN, LSTM, and fully connected neural networks. Caffe supports acceleration computing kernel libraries based on GPUs and CPUs, such as NVIDIA cuDNN and Intel MKL.
[0057] TensorFlow is one of the deep learning frameworks developed by the Google team. It is an open-source software designed entirely in the Python language. The original intention of TensorFlow is to implement the concepts of machine learning and deep learning in the simplest way. It combines optimization techniques of computational algebra, enabling it to compute many mathematical expressions. TensorFlow can train and run deep neural networks and can be applied in many scenarios, such as image recognition, handwritten digit classification, recurrent neural networks, word embedding, natural language processing, video detection, etc. TensorFlow can run on multiple CPUs or GPUs and can also run on mobile operating systems (such as Android, iOS, etc.). Its architecture is flexible and has good scalability, capable of supporting various network models (such as the seven-layer OSI and the four-layer TCP / IP).
[0058] The predecessor of PyTorch is Torch. Its underlying layer is the same as the Torch framework, but a lot of content has been rewritten in Python, making it not only more flexible, supporting dynamic graphs, but also providing Python interfaces. It is developed by the Torch7 team and is a Python-first deep learning framework that can not only achieve powerful GPU acceleration but also support dynamic neural networks.
[0059] In another embodiment, to facilitate the understanding of the technical solution of the present invention, the following takes the deployment of an inference service in Kubernetes, a container orchestration engine open-sourced by Google, as an example for illustration. This embodiment provides a method for deploying an inference service. Please refer to Figure 2 as shown, and the specific implementation is as follows:
[0060] This embodiment defines the inference service in the following three parts, namely: (1) Inference framework: The processing flow for solving a certain problem; When the inference framework runs: The startup parameters and configuration information required when the inference service is started, where the runtime refers to the environmental basis required when the trained model runs in the actual production environment; (3) Model file: The AI model obtained through training, which, when combined with the inference framework, forms an inference service that can be used for inference. In the specific implementation process, the inference service can refer to the following code:
[0061]
[0062] The following describes the process of deploying the inference service in combination with the above three defined parts: When creating the inference service, first obtain the runtime configuration required for the inference service. This runtime configuration defines the framework type, startup parameters, required resources, configuration information, etc., and is used to start the inference service framework; After pulling and starting the inference service framework during the runtime of the inference service, the required model file is pulled into the container through the init container mechanism of Kubernetes POD, and then the service is started to respond to real-time inference requests. In this embodiment, AIService is taken as an example. In AIService, the model file and the corresponding runtime required for the inference service are defined. The model file includes the model file location, pull protocol, and save location. The runtime of the inference service includes information such as the inference framework image, inference framework type, parameter list, and startup command. The AIService example can refer to the following code:
[0063]
[0064] Furthermore, the code for the runtime in AIService is as follows for reference:
[0065]
[0066]
[0067] For example, suppose there are two models, Model 1 and Model 2. If both Model 1 and Model 2 need to use the TensorFlow framework, but their parameters for starting the TensorFlow framework are different, then in the traditional way, two different configurations of the TensorFlow framework need to be set up, and two mirror packages for the TensorFlow framework and model packaging need to be made. However, when deploying the inference service based on the above two models using the method of the present invention, only one mirror needs to be made for the TensorFlow framework, and then the runtime and model pulling location can be customized for each model's corresponding framework respectively. For example, when deploying the inference service using Model 1, first obtain the TensorFlow framework according to the location of the TensorFlow framework mirror, then start the TensorFlow framework in combination with the custom runtime of Model 1, and finally pull Model 1 through the predefined model location to complete the deployment of the inference service. The deployment method for the inference service of Model 2 is the same. The difference from the previous deployment of the inference service using Model 1 is that the runtime of the TensorFlow framework and the model pulling location need to be re-customized.
[0068] A method for deploying an inference service according to this embodiment, on the one hand, solves the situation of low reuse rate, slow transmission and deployment when there are different model files under the same inference framework in actual production, and improves production efficiency; on the other hand, it also solves the problem that in actual production, for the same service framework, it is impossible to flexibly modify and configure information such as the startup parameters and resource requirements of the service.
[0069] In yet another embodiment, please refer to Figure 4 As shown, this embodiment provides an inference service deployment device 200. Specifically, the inference service deployment device 200 includes:
[0070] An image making module 201, configured to remove the running parameters of each inference framework and make corresponding image files for each inference framework after removing the running parameters;
[0071] A creation module 202, configured to create a model resource library including several model files;
[0072] A determination module 203, configured to determine a target model and a target inference framework based on the inference service to be deployed;
[0073] A first customization module 204, configured to customize the running parameters for the target inference framework;
[0074] A framework startup module 205, where the framework startup module 205 is configured to load the image file corresponding to the target inference framework and start the target inference framework based on the custom running parameters;
[0075] A model pulling module 206, where the model pulling module 206 is configured to pull and run the model file corresponding to the target model from the model resource library under the started target inference framework.
[0076] The above-mentioned inference service deployment device separates the inference framework and the model required for deploying the inference service. At the same time, it also removes the running parameters of the inference framework from the original framework and creates an image only for the inference framework after removing the running parameters. It starts the inference framework by using the method of separately customizing the framework running parameters, and pulls the model file from the model resource library by using the started framework. This not only reduces the size of the image for deploying the inference service, but also improves the reuse rate of the inference framework and the model file, makes the inference framework have strong flexibility, reduces the difficulty of combining different inference frameworks and different models, enriches the deployment methods of the inference service, and significantly improves the deployment efficiency of the inference service.
[0077] In some embodiments, the inference service deployment device further includes:
[0078] A second custom module, where the second custom module is configured to customize the pulling protocol for the target model; and
[0079] The model pulling module pulls the model file by using the custom pulling protocol.
[0080] In some embodiments, the custom pulling protocol is any one of the Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Hadoop Distributed FileSystem (HDFS), and PersistentVolumeClaim (PVC).
[0081] In some embodiments, each inference framework includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
[0082] It should be noted that for the specific limitations of the inference service deployment device, reference can be made to the limitations of the inference service deployment method in the above text, which will not be elaborated here. Each module in the above inference service deployment device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0083] According to another aspect of the present invention, a computer device is provided. The computer device can be a server, and its internal structure diagram is shown in Figure 5 as follows. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above-mentioned inference service deployment method is implemented. Specifically, the inference service deployment method may include the following steps:
[0084] Remove the running parameters of each inference framework, and respectively create corresponding image files for each inference framework after removing the running parameters;
[0085] Create a model resource library including several model files;
[0086] Determine the target model and the target inference framework based on the inference service to be deployed;
[0087] Customize the running parameters for the target inference framework;
[0088] Load the image file corresponding to the target inference framework and start the target inference framework based on the customized running parameters;
[0089] Under the started target inference framework, pull and run the model file corresponding to the target model from the model resource library.
[0090] The above computer device separates the inference framework and model required for deploying the inference service by running the program on the memory through the processor. At the same time, it also removes the running parameters of the inference framework from the original framework and creates an image only for the inference framework after removing the running parameters. It starts the inference framework by using the method of separately customizing the running parameters of the framework, and pulls the model file from the model repository by using the started framework. This not only reduces the size of the image for deploying the inference service, but also improves the reuse rate of the inference framework and the model file, making the inference framework have strong flexibility, reducing the difficulty of combining different inference frameworks and different models, enriching the deployment methods of the inference service, and significantly improving the deployment efficiency of the inference service.
[0091] In some embodiments, the inference service deployment method further includes:
[0092] Customizing a pull protocol for the target model; and
[0093] When performing the step of pulling and running the model file corresponding to the target model from the model repository under the started target inference framework, the custom pull protocol is used to pull the model file.
[0094] In some embodiments, the custom pull protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
[0095] In some embodiments, each inference framework includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
[0096] According to another aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned inference service deployment method. Specifically, it includes performing the following steps:
[0097] Removing the running parameters of each inference framework and respectively making corresponding image files for each inference framework after removing the running parameters;
[0098] Creating a model repository including several model files;
[0099] Determining a target model and a target inference framework based on the inference service to be deployed;
[0100] Customizing running parameters for the target inference framework;
[0101] Loading the image file corresponding to the target inference framework and starting the target inference framework based on the custom running parameters;
[0102] Pull and run the model file corresponding to the target model from the model repository under the target inference framework after startup.
[0103] When the above computer-readable storage medium is executed on a processor, it can separate the inference framework and the model required for deploying the inference service. At the same time, it also removes the running parameters of the inference framework from the original framework and only creates an image for the inference framework after removing the running parameters. It starts the inference framework by using the method of separately customizing the framework running parameters. When pulling the model file from the model repository by using the started framework, it not only reduces the image size of deploying the inference service, but also improves the reuse rate of the inference framework and the model file, making the inference framework have strong flexibility, reducing the difficulty of combining different inference frameworks and different models, enriching the deployment methods of the inference service, and significantly improving the deployment efficiency of the inference service.
[0104] In some embodiments, the inference service deployment method further includes:
[0105] Customize a pull protocol for the target model; and
[0106] When performing the step of pulling and running the model file corresponding to the target model from the model repository under the target inference framework after startup, use the custom pull protocol to pull the model file.
[0107] In some embodiments, the custom pull protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
[0108] In some embodiments, each of the inference frameworks includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
[0109] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0110] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0111] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for deploying an inference service, characterized in that, the method for deploying the inference service includes: Removing the running parameters of each inference framework, and respectively creating corresponding image files for each inference framework after removing the running parameters; Creating a model resource library including several model files; Determining a target model and a target inference framework based on the inference service to be deployed; Customizing the running parameters for the target inference framework; Loading the image file corresponding to the target inference framework and starting the target inference framework based on the customized running parameters; Pulling and running the model file corresponding to the target model from the model resource library under the started target inference framework; Customizing a pulling protocol for the target model; and When performing the step of pulling and running the model file corresponding to the target model from the model resource library under the started target inference framework, pulling the model file using the customized pulling protocol; The customized pulling protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
2. The method for deploying an inference service according to claim 1, characterized in that, each of the inference frameworks includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
3. An apparatus for deploying an inference service, characterized in that, the apparatus for deploying the inference service includes: An image making module configured to remove the running parameters of each inference framework and respectively create corresponding image files for each inference framework after removing the running parameters; A creating module configured to create a model resource library including several model files; A determining module configured to determine a target model and a target inference framework based on the inference service to be deployed; A first customizing module configured to customize the running parameters for the target inference framework; A framework starting module configured to load the image file corresponding to the target inference framework and start the target inference framework based on the customized running parameters; A model pulling module configured to pull and run the model file corresponding to the target model from the model resource library under the started target inference framework; A second customizing module configured to customize a pulling protocol for the target model; and The model pulling module pulls the model file using the customized pulling protocol; The customized pulling protocol is any one of the Hypertext Transfer Protocol, File Transfer Protocol, Distributed File System Protocol, and Persistent Volume Claim Protocol.
4. The apparatus for deploying an inference service according to claim 3, characterized in that, each of the inference frameworks includes at least one of the Caffe framework, TensorFlow framework, and PyTorch framework.
5. A computer device, characterized in that, including: At least one processor; And A memory that stores a computer program executable in the processor, and the processor, when executing the program, performs the inference service deployment method according to any one of claims 1-2.
6. A computer-readable storage medium that stores a computer program, characterized in that when the computer program is executed by a processor, it performs the inference service deployment method according to any one of claims 1-2.
Citation Information
Patent Citations
Deep learning model online reasoning method and device, electronic equipment and storage medium
CN111461332A
Framework deployment method and device, computer equipment and storage medium
CN113190238A