Deployment method and device of deep learning model, equipment and storage medium
By obtaining the structural data and weight transformation instructions of the deep learning model, the weight data in the model file is directly transformed, which solves the problem of increasing compilation time in the existing technology, and achieves efficient model deployment and compatibility improvement.
Patent Information
- Application Number
- CN202510448617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art requires compiling the model file when deploying deep learning models on heterogeneous processors, resulting in increased compilation time costs and reduced deployment efficiency.
By obtaining the structure data and weight transformation instructions of the deep learning model, the weight data in the model file is directly transformed, the target weight data is generated, and the structure data and target weight data are deployed, avoiding the repeated compilation process of the model file.
It improves the deployment efficiency of deep learning models, reduces compilation time, reduces system storage requirements, and improves system ease of use and compatibility.
Smart Images

Figure CN120562501A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model deployment technology, and in particular to a deployment method, apparatus, device, and storage medium for a deep learning model. Background Art
[0002] A deep learning model is mainly composed of two parts: one is the network structure, and the other is the weight data. When exporting, you can save only the weight data in a checkpoint format, or you can export a model file that includes both the weight data and the network structure. In a model file, the weight data occupies most of the volume. If you want to deploy a deep learning model on a heterogeneous processor, you generally need to compile the model file obtained by model training through a specific compiler, and then optimize and transform the weight data in the model file through the compiler. The generated deployment file includes all the weight data after optimization and transformation, which increases the compilation time cost. Loading the deployment file onto a heterogeneous processor reduces the deployment efficiency of the deep learning model. Summary of the Invention
[0003] The main purpose of this application is to provide a deployment method, device, equipment and storage medium for a deep learning model, which does not require compiling the model file to be deployed, saving the time of compiling the model file before deploying the target deep learning model, thereby improving the deployment efficiency of the deep learning model.
[0004] In a first aspect, the present application provides a method for deploying a deep learning model, comprising:
[0005] Obtaining structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, where the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model;
[0006] Obtain a model file of a second deep learning model, where the model file includes second weight data in the second deep learning model, where the second deep learning model has the same structure as the first deep learning model;
[0007] performing a weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data;
[0008] Deploy the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data.
[0009] In a second aspect, the present application also provides a deployment device for a deep learning model, comprising:
[0010] A first acquisition module is configured to acquire structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, wherein the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model;
[0011] A second acquisition module is configured to acquire a model file of a second deep learning model, wherein the model file includes second weight data in the second deep learning model, and the second deep learning model has the same structure as the first deep learning model;
[0012] a weight transformation module, configured to perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data;
[0013] A model deployment module is used to deploy the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data.
[0014] In a third aspect, the present application further provides a computer device, comprising a memory and a processor;
[0015] The memory is used to store computer programs;
[0016] The processor is used to execute the computer program and implement the deployment method of the deep learning model as described above when executing the computer program.
[0017] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the deep learning model deployment method as described above are implemented.
[0018] The present application provides a method, apparatus, device and storage medium for deploying a deep learning model, wherein the method comprises: obtaining structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, the first deployment file comprising a weight transformation instruction corresponding to the first weight data in the first deep learning model; obtaining a model file of a second deep learning model, the model file comprising second weight data in the second deep learning model, the second deep learning model having the same structure as the first deep learning model; performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data; deploying the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data. The first deployment file in the present application comprises a weight transformation instruction corresponding to the first weight data in the first deep learning model, and performing weight transformation on the second weight data in the model file according to the weight transformation instruction to obtain the target weight data; after compiling the first deep learning model once to generate the first deployment file, the target deep learning model corresponding to the target weight data and the structural data of the first deep learning model can be deployed without subsequently compiling the model file to be deployed that has the same structural data as the first deep learning model, thereby saving the time for compiling the model file before deploying the target deep learning model, thereby improving the deployment efficiency of the deep learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flowchart of a method for deploying a deep learning model provided in an embodiment of the present application;
[0021] Figure 2 A schematic diagram of the data structure of the first deployment file provided in an embodiment of the present application;
[0022] Figure 3 A schematic diagram of the data structure of an existing deployment file provided in an embodiment of the present application;
[0023] Figure 4 A schematic diagram of the data structure of the model file provided in the embodiment of the present application;
[0024] Figure 5 A schematic diagram of the framework of the deployment process of the deep learning model provided in the embodiment of the present application;
[0025] Figure 6A schematic block diagram of a deep learning model deployment device provided in an embodiment of the present application;
[0026] Figure 7 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0028] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0029] The embodiments of the present application provide a method, apparatus, device, and storage medium for deploying a deep learning model. The method for deploying a deep learning model can be applied to a terminal device, which can be a mobile phone, tablet computer, laptop computer, desktop computer, or other device. The method can also be applied to a server, which can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0030] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0031] See also Figure 1 , Figure 1 A flowchart of a method for deploying a deep learning model provided in an embodiment of the present application. It should be noted that the method for deploying a deep learning model provided in an embodiment of the present application can be used in terminal devices and of course can also be used in servers.
[0032] like Figure 1 As shown, the deployment method of the deep learning model includes steps S101 to S104.
[0033] Step S101: Obtain structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, where the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model.
[0034] Specifically, the first deep learning model includes structural data and first weight data. The structural data of the first deep learning model may include a hierarchical structure, neurons, and activation functions of the model. For example, the hierarchical structure may include an input layer, a hidden layer, and an output layer, and may also include the layer type (e.g., convolutional layer, fully connected layer, pooling layer, etc.) of each layer, parameters (e.g., convolution kernel size, stride, padding method, etc.), and output dimension.
[0035] The first deployment file can be compiled and generated based on the structural data of the first deep learning model. Figure 2 As shown, the first deployment file in the embodiment of the present application may include an index identifier corresponding to the first weight data in the first deep learning model, a weight transformation instruction corresponding to the first weight data, metadata, and heterogeneous processor instructions, etc.
[0036] For example, the weight transformation instruction corresponding to the first weight data refers to the step of transforming the first weight data during the compilation process. Generally, a model file can be obtained after each training of the deep learning model. By optimizing the network structure in the model file and transforming the weight data, the transformed weight data can be obtained. Specifically, it may include: 1. Several continuous operators are equivalently transformed, and the related operations are fused into one operator, and the corresponding weights are fused at the same time; 2. The weight data of a single operator is usually type-converted, calculated, and arranged according to the calculation requirements. The embodiment of the present application does not directly transform the first weight data, but uses the steps that require the transformation of the first weight data as weight transformation instructions.
[0037] For example, during the compilation process, the original weight data is not directly operated, but the weight data index identifier is recorded and the operation transformation process of the original weight data is recorded in the form of weight transformation instructions, and finally saved in the first deployment file.
[0038] Exemplarily, the weight transformation instruction may include multiple different instructions, for example, instruction 1, instruction 2, ... instruction N. It will be understood that the transformation processing steps corresponding to different first weight data (i.e., weight transformation instructions) may be the same or different. For example, the first weight data 1 corresponds to instruction 1, the first weight data 2 corresponds to instruction 2, the first weight data 3 and the first weight data 4 correspond to instruction 3, etc. The first deployment file in the embodiment of the present application records the weight transformation instruction corresponding to the transformation step that the first weight data needs to be transformed, rather than the weight data after the transformation.
[0039] like Figure 3 As shown, the data structure of the existing deployment file includes meta information, heterogeneous processor instructions and transformed weight data. Obviously, the existing deployment file includes the transformed weight data, which is equivalent to the prior art needing to maintain two copies of weight data in the deployment file and the model file; while the first deployment file in the embodiment of the present application records the weight transformation instructions corresponding to the first weight data, and does not include a large amount of transformed weight data. Compared with the existing deployment file, the first deployment file requires less data to be maintained and is more convenient, thereby reducing the capacity of the system storage file and reducing the hard disk storage requirements, and only the first deployment file and the deep learning model that do not include the weight data need to be maintained.
[0040] Step S102: Obtain a model file of a second deep learning model, where the model file includes second weight data in the second deep learning model. The second deep learning model has the same structure as the first deep learning model.
[0041] like Figure 4 As shown, it can be understood that the deep learning model is mainly composed of two parts: one is the network structure and the other is the weight data. After each training of the deep learning model, a model file including the weight data and the network structure data can be exported.
[0042] In the embodiments of the present application, the network structure of the second deep learning model is the same as that of the first deep learning model. For example, the algorithm used to train the second deep learning model is the same as that used to train the first deep learning model. The dataset used to train the second deep learning model can be the same as or different from the dataset used to train the first deep learning model.
[0043] like Figure 5 As shown, after the embodiment of the present application obtains the model file of the second deep learning model, the model file can be stored in the storage medium of the device, so that the device can obtain multiple second weight data from the model file.
[0044] For example, when the second deep learning model is retrained to obtain a new model file, there is no need to compile the new model file into a deployment file first. It is only necessary to replace the old model file in the storage medium with the new model file. For deep learning models that have been trained many times, reducing the steps of compiling the model file corresponding to the model after each training can reduce the time for deploying the model file, thereby improving the efficiency of model deployment. Obviously, in the application scenario where the model structure remains unchanged and multiple training optimizations are performed, the embodiment of the present application can directly replace the model file to be deployed without recompilation. On the one hand, it saves compilation and deployment time, and on the other hand, it can improve the usability of the system.
[0045] Step S103: Perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data.
[0046] The target weight data in the embodiment of the present application refers to the target weight data that can be deployed on a device with a heterogeneous processor after weight transformation.
[0047] Exemplarily, after obtaining the first deployment file and the model file, the second weight data in the model file can be weight-transformed according to the weight transformation instruction in the first deployment file to obtain the target weight data. Specifically, the first deployment file and the model file can be first stored on a storage medium of a device where the model needs to be deployed. Then, in the process of loading the first deployment file and the model file from the storage medium, the second weight data can be weight-transformed according to the first weight data index identifier and the corresponding weight transformation instruction recorded in the first deployment file to obtain the target weight data.
[0048] Step S104: deploy the target deep learning model corresponding to the structural data of the first deep learning model and the target weight data.
[0049] It can be understood that after the target weight data obtained according to the weight transformation instruction is loaded into the heterogeneous processor of the device, the heterogeneous processor can calculate the target weight data and rerun it according to the structural data of the first deep learning model to deploy the target deep learning model.
[0050] The deployment method of the deep learning model provided in the above embodiment includes: obtaining the structural data of the first deep learning model and the first deployment file corresponding to the first deep learning model, the first deployment file including the weight transformation instruction corresponding to the first weight data in the first deep learning model; obtaining the model file of the second deep learning model, the model file including the second weight data in the second deep learning model, and the second deep learning model has the same structure as the first deep learning model; performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data; deploying the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data. The first deployment file in the embodiment of the present application includes the weight transformation instruction corresponding to the first weight data in the first deep learning model, and performing weight transformation on the second weight data in the model file according to the weight transformation instruction to obtain the target weight data; after compiling the first deep learning model once to generate the first deployment file, there is no need to compile the model file to be deployed and which has the same structural data as the first deep learning model, and the target deep learning model corresponding to the target weight data and the structural data of the first deep learning model can be deployed, saving the time of compiling the model file before deploying the target deep learning model, thereby improving the deployment efficiency of the deep learning model.
[0051] It can be understood that in the process of compiling the first deployment file, the embodiment of the present application mainly relies on the model structure of the first deep learning model, and does not rely on the transformed weight data, thereby reducing the requirements for the model.
[0052] In an exemplary embodiment, the second deep learning model can be trained by the first deep learning model.
[0053] For example, the second deep learning model in the embodiment of the present application can be obtained by training the first deep learning model. Specifically, the algorithm used for training the second deep learning model is the same as the algorithm used for training the first deep learning model. The data set used for training the second deep learning model may be the same as or different from the data set used for training the first deep learning model. Furthermore, the second deep learning model may be obtained by training the first deep learning model once, or by training the first deep learning model more than twice. Deep learning models with the same network structure undergo different training. When deploying the newly trained deep learning model, it is only necessary to directly replace the old model file obtained by the previous training in the storage medium of the device, thereby improving the deployment efficiency of the deep learning model.
[0054] In an exemplary embodiment, obtaining a first deployment file corresponding to a first deep learning model may specifically include: when compiling the first deep learning model based on a preset compilation tool, obtaining a weight transformation instruction for the preset compilation tool to perform a weight transformation on the first weight data in the first deep learning model. For example, an embodiment of the present application first obtains a model file obtained after initial training or multiple training of the first deep learning model, and then compiles the model file corresponding to the first deep learning model based on the preset compilation tool. It can be understood that the preset compilation tool is used to convert the model defined in the deep learning framework into code that can be efficiently executed on different hardware. The preset compilation tool may include Adlik model compiler, TVM, Glow, MLIR, etc. Among them, Adlik is an end-to-end deep learning model inference acceleration tool chain, and its model compiler supports multiple deep learning frameworks and can compile the model into the model format required by multiple inference service engines. TVM is an open source deep learning compiler framework that supports multiple deep learning frameworks and hardware backends, can generate efficient code, and achieve high-performance inference on different hardware. Glow focuses on enabling efficient inference on mobile and edge devices, supporting a variety of deep learning models and hardware backends, and improving inference performance through graph optimization and code generation techniques. MLIR can provide a flexible and extensible intermediate representation framework. For example, it can be a compiler framework that supports heterogeneous device instruction conversion based on multi-layer intermediate representations to support the compilation and optimization of machine learning and deep learning models.
[0055] Specifically, the first weight data index identifier corresponding to the first weight data in the model file can be recorded first, and then the weight transformation instruction of the first weight data can be determined during the optimization of the first deep learning model.
[0056] like Figure 2 As shown, the first deployment file in the embodiment of the present application records the weight transformation instructions of the first weight data, rather than the weight data after weight transformation. The volume of the first deployment file is reduced. Compared with the existing solution of maintaining two copies of weight data in the model file and the deployment file, it can avoid the generation of additional storage costs and bandwidth pressure during frequent storage and transmission. Moreover, the hard disk storage requirements are reduced, and only the first deployment file without weight data and the deep learning model need to be maintained.
[0057] For example, the embodiment of the present application may further record the data type of the first weight data to facilitate loading and processing. At the same time, the size of the space required for the first weight data after transformation may be recorded to facilitate allocation of storage space.
[0058] In an exemplary embodiment, the weight transformation instruction includes at least one of the following: a spatial instruction, an arithmetic instruction, and an arrangement instruction. Step S103 may include steps S1031 to S1033.
[0059] Step S1031: Mark the storage location of the second weight data after weight transformation according to the spatial instruction.
[0060] Step S1032: Perform arithmetic processing on the second weight data according to an arithmetic instruction to obtain target weight data.
[0061] Step S1033: According to the sorting instructions, the second weight data is transposed, concatenated and / or sliced to obtain target weight data.
[0062] For example, in the weight transformation instruction, the first weight data index identifier corresponding to the first weight data can be used to represent it. It can be understood that the first weight data index identifier can be an integer. Generally, the index identifier less than a preset threshold (such as 10000) can be used as the index identifier corresponding to the first weight data before the weight transformation, and the index identifier greater than or equal to the preset threshold can be used as the index identifier corresponding to the first weight data after the weight transformation.
[0063] Specifically, a spatial instruction can be a storage location for a marked transformation or intermediate calculation result. For example, a spatial instruction is mem(10000,offset), which means that the transformed data of the first weight data with index identifier 10000 is placed at a position offset by offset in the transformation data space.
[0064] Arithmetic instructions may include operations such as addition, subtraction, multiplication, division, quantization, and dequantization. For example, an arithmetic instruction is 10000=add(1,2), which means adding the first weight data with index identifier 1 and the first weight data with index identifier 2 to generate the first weight data with index identifier 10000.
[0065] Sorting instructions may include operations such as transposition, splicing or slicing according to specific dimensions. For example, a sorting instruction is 10000=transpose(1,(0,2,3,1)), which means that the first weight data of the four dimensions with index identified as 1 is transformed according to order=(0,2,3,1).
[0066] The embodiment of the present application records the steps required for weight transformation of weight data according to different weight transformation instructions, which can further reduce the size of the first deployment file.
[0067] In an exemplary embodiment, step S103 may specifically include: loading the weight transformation instruction and the second weight data into the heterogeneous processor, and performing weight transformation on the second weight data according to the weight transformation instruction based on the heterogeneous processor to obtain target weight data.
[0068] For example, in the embodiment of the present application, the second weight data in the model file can be transformed during the process of loading the first deployment file and the model file online. Figure 5 As shown, in the process of loading the first deployment file and the model file from the storage medium of the device to the heterogeneous processor, the second weight data in the model file can be weight-transformed according to the weight transformation instruction based on the heterogeneous processor to obtain the target weight data. The embodiment of the present application performs weight transformation processing when loading data without changing the external interface, but only changes the internal loading process, which is imperceptible to the user. And the use of heterogeneous processors for on-site processing has little effect on the loading speed, which can meet the requirements of subsequent running calculation processes.
[0069] In an exemplary embodiment, step S103 may specifically include: inputting the weight transformation instruction and the second weight data into a preset conversion tool, and based on the preset conversion tool, performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data.
[0070] In some embodiments, a preset conversion tool (such as an application) may also be provided. It is understandable that the preset conversion tool is different from the device with a heterogeneous processor and may be a program separated from the online operating environment. In an embodiment of the present application, the weight transformation instruction and the second weight data may be input into a preset conversion tool to perform a weight transformation on the second weight data according to the weight transformation instruction based on the preset conversion tool to obtain target weight data. By providing a preset conversion tool, the second weight data is weight-transformed in advance according to the weight transformation instruction, and the weight data after weight transformation can be obtained without compilation, and then the subsequent loader is entered, which is compatible with the loader in the original deployment scheme, thereby improving the applicability and compatibility of the device.
[0071] In an exemplary embodiment, step S100 may be included before step S104.
[0072] Step S100: Store the target weight data into a first deployment file.
[0073] Exemplarily, after the second weight data is weight-transformed according to the weight transformation instruction based on the preset conversion tool to obtain the target weight data, the target weight data can be stored in the first deployment file to obtain a deployment file with the weight data after weight transformation. Specifically, the weight data after weight transformation can replace the data content of the relevant index identifier and weight transformation instruction in the first deployment file to obtain a second deployment file, and the data format of the second deployment file is the same as the data format of the existing deployment file. Then, the second deployment file is stored in the storage medium of the device. The embodiment of the present application does not need to go through the compilation step again, and the weight data after weight transformation can also be obtained to carry out the subsequent deployment process.
[0074] In actual applications, deployment files of different formats can be distinguished based on the version number recorded in the metadata of the deployment file or based on the data structure. For example, if the deployment file contains an index identifier of the weight data, it means that the format of the deployment file is the first deployment file.
[0075] See also Figure 6 , Figure 6 A schematic block diagram of a deep learning model deployment device provided in an embodiment of the present application. The deep learning model deployment device can be configured in a server or terminal device to execute the aforementioned deep learning model deployment method.
[0076] like Figure 6 As shown, the deployment device of the deep learning model includes: a first acquisition module 110, a second acquisition module 120, a weight transformation module 130 and a model deployment module 140.
[0077] The first acquisition module 110 is used to obtain structural data of the first deep learning model and a first deployment file corresponding to the first deep learning model. The first deployment file includes a weight transformation instruction corresponding to the first weight data in the first deep learning model.
[0078] The second acquisition module 120 is used to obtain a model file of a second deep learning model, where the model file includes second weight data in the second deep learning model. The second deep learning model has the same structure as the first deep learning model.
[0079] The weight transformation module 130 is used to perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data.
[0080] The model deployment module 140 is used to deploy the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data.
[0081] In an exemplary embodiment, the second deep learning model is trained by the first deep learning model.
[0082] In an exemplary embodiment, the first acquisition module 110 can be specifically used to obtain a weight transformation instruction for the preset compilation tool to perform weight transformation on the first weight data in the first deep learning model when the first deep learning model is compiled based on the preset compilation tool.
[0083] In an exemplary embodiment, the weight transformation instruction includes at least one of the following: a spatial instruction, an arithmetic instruction, and an arrangement instruction. The weight transformation module 130 may include a marking submodule, an operation submodule, and an arrangement submodule.
[0084] The marking submodule is used to mark the storage location of the second weight data after weight transformation according to the spatial class instruction.
[0085] The operator module is used to perform operation processing on the second weight data according to arithmetic instructions to obtain target weight data.
[0086] The arranging submodule is used to perform transposition, splicing and / or slicing processing on the second weight data according to the arranging instructions to obtain target weight data.
[0087] In an exemplary embodiment, the weight transformation module 130 can be specifically used to load the weight transformation instruction and the second weight data into the heterogeneous processor, and based on the heterogeneous processor, perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data.
[0088] In an exemplary embodiment, the weight transformation module 130 can be specifically used to input the weight transformation instruction and the second weight data into a preset conversion tool, and based on the preset conversion tool, perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data.
[0089] In an exemplary embodiment, the device further includes a storage module.
[0090] The storage module is used to store the target weight data in the first deployment file.
[0091] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0092] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0093] Illustratively, the above-mentioned method and apparatus may be implemented in the form of a computer program, which may be run on a computer device.
[0094] See also Figure 7 , Figure 7 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be a server or a terminal device.
[0095] like Figure 7 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a storage medium and an internal memory.
[0096] The storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause a processor to perform the steps of any one of the deep learning model deployment methods.
[0097] The processor is used to provide control and lightweight computing capabilities, supporting the operation of the entire computer device.
[0098] The processor may include a heterogeneous computing processor, which is used to provide high computing power and high bandwidth capabilities, and is mainly responsible for heavyweight computing tasks such as deep learning model inference tasks.
[0099] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can perform the steps of any deep learning model deployment method.
[0100] This network interface is used for network communication, such as sending assigned tasks.
[0101] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0102] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0103] In one embodiment, the processor is configured to execute a computer program and implement the following steps when executing the computer program:
[0104] Obtaining structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, where the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model;
[0105] Obtain a model file of a second deep learning model, where the model file includes second weight data in the second deep learning model, and the second deep learning model has the same structure as the first deep learning model;
[0106] Performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data;
[0107] Deploy the target deep learning model corresponding to the structural data of the first deep learning model and the target weight data.
[0108] In an exemplary embodiment, the second deep learning model is trained by the first deep learning model.
[0109] In an exemplary embodiment, obtaining a first deployment file corresponding to the first deep learning model includes:
[0110] When the first deep learning model is compiled based on the preset compilation tool, a weight transformation instruction for the preset compilation tool to perform weight transformation on the first weight data in the first deep learning model is obtained.
[0111] In an exemplary embodiment, the weight transformation instruction includes at least one of the following: a spatial instruction, an arithmetic instruction, and a sorting instruction; according to the weight transformation instruction, performing weight transformation on the second weight data to obtain target weight data includes:
[0112] Marking the storage location of the second weight data after weight transformation according to the spatial instruction;
[0113] According to the arithmetic instruction, the second weight data is processed to obtain the target weight data;
[0114] According to the sorting instructions, the second weight data is transposed, concatenated and / or sliced to obtain target weight data.
[0115] In an exemplary embodiment, performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data includes:
[0116] The weight transformation instruction and the second weight data are loaded into the heterogeneous processor, and based on the heterogeneous processor, the second weight data is weight transformed according to the weight transformation instruction to obtain the target weight data.
[0117] In an exemplary embodiment, performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data includes:
[0118] The weight transformation instruction and the second weight data are input into a preset transformation tool. Based on the preset transformation tool, the second weight data is weight transformed according to the weight transformation instruction to obtain the target weight data.
[0119] In an exemplary embodiment, before deploying the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data, the process includes:
[0120] Store the target weight data into the first deployment file.
[0121] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working process of the deployment of the deep learning model described above can refer to the corresponding process in the embodiment of the aforementioned deep learning model deployment method, and will not be repeated here.
[0122] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The method implemented when the computer program is executed by a processor can refer to the various embodiments of the deployment method of the deep learning model of the present application.
[0123] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.
[0124] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0125] It should also be understood that the term "and / or" used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.
[0126] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments. The above description is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for deploying a deep learning model, characterized in that: include: Obtaining structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, where the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model; Obtain a model file of a second deep learning model, where the model file includes second weight data in the second deep learning model, where the second deep learning model has the same structure as the first deep learning model; performing a weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data; Deploy the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data.
2. The method for deploying a deep learning model according to claim 1, wherein: The second deep learning model is trained by the first deep learning model.
3. The method for deploying a deep learning model according to claim 1, wherein: The obtaining a first deployment file corresponding to the first deep learning model includes: When the first deep learning model is compiled based on a preset compilation tool, a weight transformation instruction for the preset compilation tool to perform weight transformation on the first weight data in the first deep learning model is obtained to obtain a first deployment file.
4. The method for deploying a deep learning model according to claim 1, wherein: The weight transformation instruction includes at least one of the following: a spatial instruction, an arithmetic instruction, and a sorting instruction; performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data includes: Marking, according to the spatial instruction, a storage location of the second weight data after weight transformation; performing arithmetic processing on the second weight data according to the arithmetic instruction to obtain target weight data; According to the sorting instructions, the second weight data is transposed, concatenated and / or sliced to obtain target weight data.
5. The method for deploying a deep learning model according to claim 1, wherein: The step of performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data includes: The weight transformation instruction and the second weight data are loaded into a heterogeneous processor, and based on the heterogeneous processor, the second weight data is weight transformed according to the weight transformation instruction to obtain target weight data.
6. The method for deploying a deep learning model according to claim 1, wherein: The step of performing weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data includes: The weight transformation instruction and the second weight data are input into a preset conversion tool, and based on the preset conversion tool, the second weight data is weight transformed according to the weight transformation instruction to obtain target weight data.
7. The method for deploying a deep learning model according to claim 6, wherein: Before deploying the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data, the method includes: The target weight data is stored in the first deployment file.
8. A deployment device for a deep learning model, characterized in that: include: A first acquisition module is configured to acquire structural data of a first deep learning model and a first deployment file corresponding to the first deep learning model, wherein the first deployment file includes a weight transformation instruction corresponding to first weight data in the first deep learning model; A second acquisition module is configured to acquire a model file of a second deep learning model, wherein the model file includes second weight data in the second deep learning model, and the second deep learning model has the same structure as the first deep learning model; a weight transformation module, configured to perform weight transformation on the second weight data according to the weight transformation instruction to obtain target weight data; A model deployment module is used to deploy the structural data of the first deep learning model and the target deep learning model corresponding to the target weight data.
9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the deployment method of the deep learning model as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning model deployment method as described in any one of claims 1 to 7 are implemented.