Model Optimization Method, Device, Equipment, Storage Medium and Program Product
By automatically identifying the framework and structure types of artificial intelligence model components and automatically matching the optimization strategy based on the optimization strategy mapping table, a fully automated process of model optimization is realized, solving the problems of low optimization efficiency and insufficient flexibility in the existing technology.
Patent Information
- Application Number
- CN202411137679.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-08-19
AI Technical Summary
The existing artificial intelligence model framework needs to manually adapt different optimization engines to each component when optimizing the model, resulting in low optimization efficiency and lack of flexibility.
By identifying the framework types and structure types of each component in the target model to be optimized, and calling matching optimization strategies for each component according to the preset optimization strategy mapping table, thereby achieving inference-accelerated optimization of the target model.
The fully automated process of model optimization is realized, the optimization efficiency is improved, the accuracy and pertinence of optimization strategies are ensured, and the problems of low optimization efficiency and insufficient flexibility in the inference acceleration optimization of multi-frame models are solved.
Smart Images

Figure CN119150941B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a model optimization method, apparatus, device, storage medium, and program product. Background Art
[0002] With the diversification of task requirements, the structure of artificial intelligence models has become more complex, involving multiple components working together to form a complex pipeline structure, and these different components correspond to different types of model structures.
[0003] Existing artificial intelligence model frameworks, such as PyTorch, TensorFlow, and ONNX Runtime, describe the model structure and its running logic through different file formats. When a model includes multiple framework types, although the multi-framework model has advantages in model representation and execution, in terms of inference acceleration optimization of the model, when performing model optimization, it is necessary to manually adapt different optimization engines for each component, resulting in low optimization efficiency and lack of flexibility.
[0004] In summary, how to improve the efficiency and flexibility of model optimization has become an urgent technical problem in this field. Summary of the Invention
[0005] The main purpose of this application is to provide a model optimization method, apparatus, device, storage medium, and program product, aiming to improve the efficiency and flexibility of model optimization.
[0006] To achieve the above object, this application proposes a model optimization method, which includes:
[0007] Identify the framework type and structure type of each component in the target model to be optimized;
[0008] According to a preset optimization strategy mapping table, call a matching optimization strategy for each of the components, where the optimization strategy is mapped and stored in the optimization strategy mapping table with the framework type and structure type of the component;
[0009] Optimize each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0010] In one embodiment, the step of identifying the framework type and structure type of each component in the target model to be optimized includes:
[0011] For a target component in each component of the target model to be optimized, perform framework identification on the target component through a preset component framework classification identifier to determine the framework type of the target component;
[0012] Perform structure recognition on the target component through a preset model structure classifier to determine the structure type of the target component.
[0013] In one embodiment, before the step of performing structure recognition on the target component through a preset model structure classifier to determine the structure type of the target component, it further includes:
[0014] When there are multiple frame types of the target component, split the target component into structural parts according to different frame types to obtain each part of the target component;
[0015] The step of performing structure recognition on the target component through a preset model structure classifier to determine the structure type of the target component includes:
[0016] Perform structure recognition on each part of the target component through a preset model structure classifier to determine the structure type of each part of the target component.
[0017] In one embodiment, before the step of calling the matching optimization strategy for each component according to the preset optimization strategy mapping table, it further includes:
[0018] Based on a preset user description component optimization rule interface, receive the input optimization rules, where the optimization rules include the corresponding relationship between the frame type and the structure type and the optimization strategy;
[0019] Create an optimization strategy mapping table according to the optimization rules.
[0020] In one embodiment, the step of calling the matching optimization strategy for each component according to the preset optimization strategy mapping table includes:
[0021] For the target component in each component of the target model to be optimized, according to the frame type and structure type of the target component, determine the optimization strategy matching the target component according to the preset optimization strategy mapping table;
[0022] When there are multiple matching optimization strategies, select the target optimization strategy from the multiple matching optimization strategies according to the received optimization request conditions and the running time of the optimization strategy;
[0023] Call the target optimization strategy for the target component.
[0024] In one embodiment, after the step of optimizing each component according to the optimization strategy to complete the inference acceleration optimization of the target model, it further includes:
[0025] Record the overall optimization strategy for the inference acceleration optimization of the target model;
[0026] After receiving an optimization request for the target model, optimize the target model according to the overall optimization strategy.
[0027] In addition, to achieve the above object, the present application also proposes a model optimization device, which includes:
[0028] An identification module, configured to identify the framework type and structure type of each component in the target model to be optimized;
[0029] An optimization strategy matching module, configured to call a matching optimization strategy for each of the components according to a preset optimization strategy mapping table, where the optimization strategy is mapped and stored in the optimization strategy mapping table with the framework type and structure type of the component;
[0030] An optimization module, configured to optimize each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0031] In addition, to achieve the above object, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the model optimization method as described above.
[0032] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the model optimization method as described above are implemented.
[0033] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the model optimization method as described above are implemented.
[0034] The present application proposes a model optimization method. In the present application, first, each component in the target model to be optimized is identified, and its framework type and structure type are determined. Then, a preset optimization strategy mapping table is used, which associates different framework types and structure types with specific optimization strategies; according to the optimization strategy mapping table, the corresponding optimization strategy is automatically applied to each component of the target model without manual intervention, so as to optimize all components in the target model, complete the inference acceleration optimization of the target model, and improve the execution efficiency of the model in actual applications.
[0035] Compared with the traditional model optimization methods which require developers to manually select and adapt different optimization strategies for each component in a multi-framework model, this not only requires developers to have profound professional knowledge and rich experience, but also greatly increases the complexity and time cost of the model optimization work. In addition, manual optimization often fails to achieve global optimality, resulting in unsatisfactory optimization effects.
[0036] In this application, by automatically identifying the component framework and structure type, and automatically matching and executing the optimization strategy according to the preset optimization strategy mapping table, a fully automated process from identification to optimization is realized, which improves the optimization efficiency and also ensures the accuracy and pertinence of the optimization strategy. Thus, the problems of low optimization efficiency and insufficient flexibility faced by multi-framework models in inference acceleration optimization are effectively solved, enabling multi-framework models to execute inference tasks more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0038] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a schematic flowchart provided for the first embodiment of the model optimization method of the present application;
[0040] Figure 2 It is a schematic model optimization flowchart provided for the second embodiment of the model optimization method of the present application;
[0041] Figure 3 It is a schematic diagram of the user description component optimization rule interface structure provided for the second embodiment of the model optimization method of the present application;
[0042] Figure 4 It is a schematic diagram of the model structure provided for the second embodiment of the model optimization method of the present application;
[0043] Figure 5 It is a schematic diagram of the module structure of the model optimization device according to the embodiment of the present application;
[0044] Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the model optimization method according to the embodiment of the present application.
[0045] The realization of the purpose, functional features and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of this application and are not used to limit this application.
[0047] To better understand the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0048] The main solution of the embodiments of this application is: identifying the framework type and structure type of each component in the target model to be optimized; according to a preset optimization strategy mapping table, respectively calling a matching optimization strategy for each of the components, where the optimization strategy is mapped and stored in the optimization strategy mapping table with the framework type and structure type of the component; optimizing each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0049] Since existing artificial intelligence model frameworks, such as PyTorch, TensorFlow, and ONNX Runtime, describe the model structure and its running logic through different file formats, when a model includes multiple framework types, although a multi-framework model has advantages in model representation and execution, in terms of the inference acceleration optimization of the model, when performing model optimization, it is necessary to manually adapt different optimization engines for each component, resulting in low optimization efficiency and lack of flexibility.
[0050] The embodiments of this application provide a solution, which realizes a fully automated process from identification to optimization by automatically identifying the component framework and structure type and automatically matching and executing the optimization strategy according to a preset optimization strategy mapping table, improves the optimization efficiency, and also ensures the accuracy and pertinence of the optimization strategy, thereby effectively solving the problems of low optimization efficiency and lack of flexibility faced by multi-framework models in inference acceleration optimization, enabling multi-framework models to execute inference tasks more efficiently.
[0051] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, etc., or an electronic device capable of implementing the above functions. The following takes the model optimization terminal as an example to illustrate this embodiment and the following embodiments.
[0052] Based on this, the embodiments of this application provide a model optimization method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the model optimization method of this application.
[0053] In this embodiment, the model optimization method includes steps S10 to S30:
[0054] Step S10: Identify the framework type and structure type of each component in the target model to be optimized;
[0055] It should be noted that the target model is a multi-framework model, that is, different parts of the model adopt multiple different deep learning model frameworks such as PyTorch, ONNX Runtime, and TensorFlow (three mainstream model frameworks).
[0056] For the target model to be optimized, identify the framework type adopted by each component in the target model, such as PyTorch, TensorFlow, ONNX Runtime, etc., and their respective structure types, such as ControlNet (a neural network structure for image segmentation tasks), U-Net (a convolutional neural network structure), etc. This step is the basis for subsequent optimization strategy selection, ensuring that the most suitable optimization strategy can be applied to different types of components.
[0057] Step S20: According to the preset optimization strategy mapping table, call the matching optimization strategy for each of the components, where the optimization strategy is stored in the optimization strategy mapping table in a mapped manner with the framework type and structure type of the component;
[0058] According to the pre-defined optimization strategy mapping table, assign the corresponding optimization strategy to each component in the target model. This optimization strategy mapping table details the optimization strategies applicable to components of different framework types and structure types, ensuring that each component can obtain the best optimization plan.
[0059] It should be noted that in the optimization strategy mapping table, the optimization strategies can be divided into two categories: native optimization strategies and customized optimization strategies. For components with relatively unified model structure descriptions, the native optimization strategies supported by this format can be directly matched. For example, for components of the ONNX framework, the TensorRT deep learning inference engine can be directly used for optimization; for components with relatively specialized model structures, the customized optimization strategies can be directly matched. For example, components using the PyTorch framework or the TensorFlow framework, which contain the unique characteristics of the framework and are not easy to be transplanted across frameworks or use general optimization strategies, require customized optimization strategies customized for specific frameworks to make full use of the advantages of the framework.
[0060] Step S30: Optimize each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0061] After determining the optimization strategies of each component, implement optimization measures for each component one by one. The optimization measures include operator fusion, precision optimization, memory management improvement, parallel processing, etc.
[0062] In this embodiment, compared with the traditional model optimization method that requires developers to manually select and adapt different optimization strategies for each component in the multi-frame model, this not only requires developers to have deep professional knowledge and rich experience, but also greatly increases the complexity and time cost of model optimization. In addition, manual optimization is often difficult to achieve global optimization, resulting in unsatisfactory optimization results.
[0063] In this application, by automatically identifying component frameworks and structural types, and automatically matching and executing optimization strategies according to preset optimization strategy mapping tables, a fully automated process from identification to optimization is achieved, which improves optimization efficiency and ensures the accuracy and pertinence of optimization strategies, thereby effectively solving the problems of low optimization efficiency and insufficient flexibility faced by multi-framework models in reasoning acceleration optimization, enabling multi-framework models to perform reasoning tasks more efficiently.
[0064] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. On this basis, step S10 can include steps S101 to S102:
[0065] Step S101, for a target component in each component of the target model to be optimized, framework identification is performed on the target component by a preset component framework classification identifier to determine the framework type of the target component;
[0066] The preset component framework classification identifier is used to perform framework identification on each target component in the target model. The component framework classification identifier determines the framework type to which the target component belongs by analyzing the code or model representation of the target component.
[0067] For example, if the target component is implemented using the PyTorch framework, the component framework classification identifier can identify it as the PyTorch framework by checking specific API calls or component structures.
[0068] Step S102: Performing structural recognition on the target component by using a preset model structure classification identifier to determine the structural type of the target component.
[0069] Use the preset model structure classification identifier to analyze the internal structure of the target component and identify its structural type based on the component's functions and operator characteristics.
[0070] In a feasible embodiment, step S103 may be further included before step S102:
[0071] Step S103, when the target component has multiple framework types, the target component is structurally split according to different framework types to obtain various partial target components;
[0072] When the target component contains multiple framework types, the target component is structurally split according to different framework types to obtain each part of the target component. This splitting process ensures that each part of the target component can be independently subjected to subsequent structure recognition and optimization.
[0073] For example, if a component contains both the PyTorch and TensorFlow frameworks at the same time, the component is separated into two independent parts for separate processing.
[0074] On this basis, step S102 may include step S1021:
[0075] Step S1021, respectively perform structure recognition on each of the partial target components through a preset model structure classification recognizer to determine the respective structure types of each of the partial target components.
[0076] For each of the split partial target components, use the model structure classification recognizer to perform structure recognition to ensure that the structure type of each partial target component can be accurately recognized, so as to customize the most suitable optimization strategy for each partial target component.
[0077] For example, if a partial target component after splitting is a convolutional neural network structure implemented by PyTorch, then this structure is recognized and a matching optimization strategy is selected for this partial target component.
[0078] Specifically, each component in the target component may be built based on different frameworks, such as PyTorch, TensorFlow, etc. These frameworks have different representation methods and performance characteristics. The component framework classification recognizer is used to analyze and recognize the framework types used by each component in the target model. If a component uses multiple frameworks, the component framework classification recognizer will recognize these different parts and split the parts with different frameworks. Then, according to the specific structural characteristics of the component, the model structure classification recognizer classifies it into more specific categories, so as to select appropriate optimization strategies for each component in the target model according to the framework type and structure type respectively.
[0079] In this way, this embodiment can accurately identify the framework and structure type of each component in the target model, providing an efficient, flexible and automated optimization solution for multi-framework models, not only improving the accuracy of model optimization, but also significantly enhancing the optimization efficiency through the automated recognition and matching process.
[0080] Exemplarily, in a feasible embodiment, such as Figure 2As shown, when the given target model includes multiple different components (Component A, Component B, Component C, Component D), the component framework classification recognizer will obtain the main call structure that becomes the performance bottleneck of the component through the disassembling on the call, and identify the artificial intelligence model framework representation used by this structure. For the situation where multiple framework representations appear in a component, such as the ONNX, PyTorch, and TensorFlow frameworks, the component framework classification recognizer will perform further optimization on the parts with different frameworks / different structures in subsequent steps through structural split calls. Then, the model structure classification recognizer will further refine the classification of the component modules corresponding to the component frameworks passed by the component framework classification recognizer, so that some model specialization component compilation optimization methods, or even some cross-platform optimization methods, can be better applied.
[0081] Among them, the component framework classification recognizer determines and identifies different component framework representations through the call object categories provided by the programming language, recursively obtains the category calls that appear in the call, classifies the recognizable optimizable component framework representations, and during this classification process, the multiple component framework representations that appear in the same component will be automatically decomposed and enter the next stage of optimization; the model structure classification recognizer differentiates different model structures based on the operator features in the model graph representation and the existing main model structure names.
[0082] Finally, according to the component structure types and framework types in the target model, different optimization methods can be applied respectively: for components with relatively unified model structure descriptions (such as components applying the ONNX framework), some native optimization strategies (such as TensorRT), that is, the device native semantic optimization method in the figure, can be used; for components with relatively specialized model structures (such as components applying PyTorch or TensorFlow), customized optimization strategies (such as stable fast), that is, the cross-platform component customization tuning method in the figure, can be used, and the tuning execution engine is called according to the selected optimization strategy to achieve cross-platform component optimization of the target model.
[0083] In a feasible embodiment, steps S40 to S50 may also be included before step S20:
[0084] Step S40, based on a preset user description component optimization rule interface, receives the input optimization rules, where the optimization rules include the corresponding relationships between the framework type, structure type, and optimization strategy;
[0085] Based on a preset user description component optimization rule interface, this interface allows users to input specific optimization rules. The input optimization rules are a set of clear instructions that define the corresponding relationships between different framework types, structure types, and optimization strategies. Users can provide customized rules through this interface.
[0086] For example, based on the optimization rules received through the interface, it is specified that for a convolutional neural network structure implemented using a specific version of the PyTorch framework, a specific operator fusion optimization strategy should be adopted.
[0087] In addition, the user description component optimization rule interface can receive and parse natural language or input in a specific format to ensure the accurate entry of optimization rules.
[0088] Step S50: Create an optimization strategy mapping table according to the optimization rules.
[0089] Create an optimization strategy mapping table according to the optimization rules input by the user. This optimization strategy mapping table is used to store the mapping relationships between the framework type, structure type, and optimization strategy. By automatically processing the received optimization rules, they are converted into entries in the mapping table, and each entry clearly indicates which optimization strategy should be adopted for a specific combination of framework and structure types.
[0090] In this way, in this embodiment, the user is allowed to input custom optimization rules according to their own needs and the characteristics of the model, and ensure that these rules are understood and executed by the system, thereby improving the adaptability and effectiveness of the optimization process. At the same time, it also provides the user with control over the optimization process, enabling the optimization strategy to better meet specific business goals and performance requirements.
[0091] Exemplarily, in a feasible implementation manner, the structural block diagram of the user description component optimization rule interface is as Figure 3 shown. The framework type is described through the model framework representation type definition window, the structure type is described through the model structure category description window, and the type of the optimization strategy is selected as the device native compilation optimization method or the component custom compilation optimization method. Then, by inputting specific optimization strategies, the introduction of new rules can be realized, and the optimization requirements of users can be flexibly adapted.
[0092] Specifically, for new components, new optimization methods, and even new model framework representations encountered in the actual optimization scenario, they can all be described through the new component optimization rule method. During the actual process of model optimization, it is to traverse the like terms in the optimization rules in each individual step, classify them, and select appropriate optimization rules for optimization. For the situation where multiple rules are met, one of them is selected for execution according to the priority set by the rules. In order to be able to select the optimal optimization strategy, all optimization strategies can be evaluated to measure their running time under the corresponding request conditions and platform environments. Through this evaluation method, it can be ensured that it can adapt to new model components and model frameworks, and a better tuning method can be selected.
[0093] In a feasible embodiment, step S20 may include steps S201 to S203:
[0094] Step S201, for a target component among the components of the target model to be optimized, according to the framework type and structure type of the target component, and in accordance with a preset optimization strategy mapping table, determine the optimization strategy matched by the target component;
[0095] Deeply analyze each target component to be optimized in the target model to determine its framework type and structure type. Based on this information and combined with the optimization strategy mapping table, search for and determine the optimization strategy matched by the target component.
[0096] Step S202, when there are multiple matched optimization strategies, select a target optimization strategy from the multiple matched optimization strategies according to the received optimization request conditions and the running time of the optimization strategy;
[0097] When a target component corresponds to multiple matched optimization strategies, according to the received optimization request conditions, such as performance requirements, resource limitations, or specific user preferences, and the expected running time of various optimization strategies, select the most suitable target optimization strategy from the multiple matched strategies.
[0098] Step S203, call the target optimization strategy for the target component.
[0099] Call the corresponding target optimization strategy for the selected target component. In this way, the finally selected target optimization strategy takes into account the balance between optimization effect and efficiency, ensuring that the optimization result not only meets the user's needs but also can be completed within a reasonable time.
[0100] In a feasible embodiment, after step S30, steps S60 to S70 may further be included:
[0101] Step S60, record the overall optimization strategy for the inference acceleration optimization of the target model;
[0102] Step S70, after receiving an optimization request for the target model, optimize the target model according to the overall optimization strategy.
[0103] Record the overall optimization strategy for the inference acceleration optimization of the target model. This overall optimization strategy includes the specific optimization measures applied to each component of the target model and the decision basis for the selection strategy.
[0104] After receiving a new optimization request for the target model, automatically optimize the target model according to the recorded overall optimization strategy, allowing the system to quickly reuse the previous optimization strategy, thereby significantly improving the optimization efficiency and reducing the need for repetitive work.
[0105] Thus, in this embodiment, not only an accurate and efficient optimization strategy matching and execution are provided during the initial optimization, but also the memorization of optimization experience is achieved, enabling future optimization requests to be processed quickly and automatically, thereby improving the flexibility and scalability of multi-framework model optimization and reducing the complexity and time cost of the optimization work.
[0106] Exemplarily, in a feasible implementation manner, to facilitate understanding the implementation process of the model optimization method obtained by combining this embodiment with the above embodiment, this embodiment presents an implementation process of model optimization for the application pipeline of a text-to-image model + face-swapping model in an actual production application:
[0107] The model is as Figure 4 shown, where the model includes three inputs (target face shape, control image, and prompt), two components (Component A and Component B), and one output (target image). Among them, Component A is a structure of stablediffusion (a deep learning text-to-image generation model) with a controlNet structure under the pytroch framework, and Component B is a face-swapping model mainly based on Inswapper.onnx (an image inpainting and face-swapping model represented using the ONNX framework).
[0108] For Component A, after being recognized by the component framework classifier, the nn.module structure (a basic structure in the PyTorch framework) is recognized and classified as a model under the pytorch framework. Then, it is continuously processed by the model structure classifier, and the controlnet structure and unet structure are recognized. These two structures can be optimized using corresponding customized optimization strategies respectively; for Component B, the call to the onnxruntime library is recognized in the component framework classifier, so it is classified as a model under the onnx framework. For models of this type of framework, the existing tensorrt already supports corresponding optimizations, and the matching native optimization strategies can be directly used for optimization. Finally, the tuning execution engine is called according to the selected optimization strategy to complete the optimization of Component A and Component B one by one. In addition, the optimization of Component A and Component B is memorized, and when the model is optimized next time, the excessive latency caused by repeated compilation and optimization of the same type of components can be avoided.
[0109] In this way, this embodiment can disassemble the model representation differences and model structure diversification in multiple model components, and apply different optimization strategies for local optimization after disassembly, so as to integrate various existing optimization strategies, avoid additional optimization workload, solve the problem of inflexible optimization deployment, and has the advantages of strong deployment compatibility, flexibility and portability; in this embodiment, it is also adapted to the increasing components in the complex flow of artificial intelligence models and the expansion of corresponding new optimization methods, and can supplement the model optimization rules incrementally in a regular format. Through the memorized deployment optimization strategy, the optimal solution for the corresponding user requests and platform environment is selected, solving the problem of limited optimization scalability, and having the advantages of strong scalability and good overall optimization effect.
[0110] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the model optimization method of this application. Based on this technical concept, more forms of simple transformation are within the protection scope of this application.
[0111] The embodiment of this application also provides a model optimization device. Please refer to Figure 5 , the model optimization device includes:
[0112] The recognition module 10 is used to recognize the framework type and structure type of each component in the target model to be optimized;
[0113] The optimization strategy matching module 20 is used to call the matching optimization strategy for each component according to the preset optimization strategy mapping table, where the optimization strategy is mapped and stored in the optimization strategy mapping table with the framework type and structure type of the component;
[0114] The optimization module 30 is used to optimize each component according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0115] Optionally, the model construction module 10 is further used for:
[0116] For the target component in each component of the target model to be optimized, the framework of the target component is recognized through a preset component framework classification recognizer to determine the framework type of the target component;
[0117] The structure of the target component is recognized through a preset model structure classification recognizer to determine the structure type of the target component.
[0118] Optionally, the model optimization device further includes a splitting module (not shown), and the splitting module is used for:
[0119] When the framework type of the target component is multiple, the target component is structurally split according to different framework types to obtain each part of the target component;
[0120] Optionally, the model construction module 10 is further configured to:
[0121] Perform structure recognition on each of the partial target components through a preset model structure classifier to determine the structure type of each of the partial target components.
[0122] Optionally, the model optimization device further includes a mapping table creation module (not shown), and the mapping table creation module is configured to:
[0123] Receive an input optimization rule based on a preset user description component optimization rule interface, where the optimization rule includes the correspondence between the framework type and the structure type and the optimization strategy;
[0124] Create an optimization strategy mapping table according to the optimization rule.
[0125] Optionally, the optimization strategy matching module 20 is further configured to:
[0126] For a target component in each component of the target model to be optimized, determine an optimization strategy matched by the target component according to the framework type and the structure type of the target component and in accordance with a preset optimization strategy mapping table;
[0127] When there are multiple matched optimization strategies, select a target optimization strategy from the multiple matched optimization strategies according to the received optimization request conditions and the running time of the optimization strategy;
[0128] Invoke the target optimization strategy for the target component.
[0129] Optionally, the model optimization device further includes a memory module (not shown), and the memory module is configured to:
[0130] Record the overall optimization strategy for the inference acceleration optimization of the target model;
[0131] After receiving an optimization request for the target model, optimize the target model according to the overall optimization strategy.
[0132] The model optimization device provided by the embodiments of the present application adopts the model optimization method in the above embodiments, and can improve the efficiency and flexibility of model optimization. Compared with the prior art, the beneficial effects of the model optimization device provided by the embodiments of the present application are the same as those of the model optimization method provided by the above embodiments, and other technical features in the model optimization device are the same as the features disclosed in the method of the above embodiments, and will not be elaborated here.
[0133] An embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model optimization method in the first embodiment above.
[0134] Reference is made below Figure 6 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player: portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The model optimization device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0135] As Figure 6 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following devices may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various devices, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.
[0136] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0137] The electronic device provided by the embodiments of the present application adopts the model optimization method in the above embodiments, which can improve the efficiency and flexibility of model optimization. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiments of the present application are the same as those of the model optimization method provided by the above embodiments, and other technical features in this electronic device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0138] It should be understood that the various parts disclosed in the embodiments of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0139] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0140] The embodiments of the present application provide a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the model optimization method in the above embodiments.
[0141] The computer-readable storage medium provided by the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0142] The above computer-readable storage medium may be included in an electronic device; or may exist separately without being assembled into the electronic device.
[0143] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device is caused to: identify the framework type and structure type of each component in the target model to be optimized; call a matching optimization strategy for each of the components according to a preset optimization strategy mapping table, where the optimization strategy is mapped and stored in the optimization strategy mapping table with the framework type and structure type of the component; optimize each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model.
[0144] Computer program code for performing the operations of the embodiments of the present application may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based device for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0146] The modules described in the embodiments of the present application may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0147] The readable storage medium provided by the embodiments of the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned model optimization method, which can improve the efficiency and flexibility of model optimization. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the embodiments of the present application are the same as those of the model optimization method provided by the above embodiments, and will not be elaborated here.
[0148] An embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the model optimization method as described above when executed by a processor.
[0149] The computer program product provided by the embodiment of the present application can improve the efficiency and flexibility of model optimization. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the model optimization method provided by the above embodiment, and will not be elaborated here.
[0150] The above are only partial embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A model optimization method, characterized in that: The model optimization method comprises: Identify the framework type and structure type of each component in the target model to be optimized; According to a preset optimization strategy mapping table, a matching optimization strategy is respectively called for each of the components, wherein the optimization strategy is stored in the optimization strategy mapping table in a mapping manner with the framework type and the structure type of the component; Optimizing each of the components according to the optimization strategy to complete the reasoning acceleration optimization of the target model; The step of respectively calling a matching optimization strategy for each of the components according to a preset optimization strategy mapping table includes: For the target component in each component of the target model to be optimized, according to the framework type and structure type of the target component, according to the preset optimization strategy mapping table, determine the optimization strategy that matches the target component, wherein the optimization strategies in the optimization strategy mapping table include native optimization strategies and customized optimization strategies, and the native optimization strategy is matched to the target component whose framework type is the ONNX framework, and the customized optimization strategy is matched to the component whose framework type is the PyTorch framework or the Tensorflow framework; When there are multiple matching optimization strategies, a target optimization strategy is selected from the multiple matching optimization strategies according to the received optimization request conditions and the running time of the optimization strategies; The target optimization strategy is invoked for the target component.
2. The model optimization method according to claim 1, characterized in that: The step of identifying the framework type and structure type of each component in the target model to be optimized includes: For a target component in each component of the target model to be optimized, a preset component framework classification identifier is used to perform framework identification on the target component to determine the framework type of the target component; The target component is structurally identified by a preset model structure classification identifier to determine the structural type of the target component.
3. The model optimization method according to claim 2, characterized in that: Before the step of performing structural identification on the target component by a preset model structure classification identifier to determine the structural type of the target component, the step further includes: When the target component has multiple framework types, the target component is structurally split according to different framework types to obtain various partial target components; The step of performing structural identification on the target component by a preset model structure classification identifier to determine the structural type of the target component includes: The structure of each of the partial target components is respectively identified by a preset model structure classification identifier to determine the structural type of each of the partial target components.
4. The model optimization method according to claim 1, characterized in that: Before the step of respectively calling the matching optimization strategy for each of the components according to the preset optimization strategy mapping table, the method further includes: Based on a preset user description component optimization rule interface, receiving an input optimization rule, wherein the optimization rule includes a correspondence between a framework type and a structure type and an optimization strategy; An optimization strategy mapping table is created according to the optimization rules.
5. The model optimization method according to any one of claims 1 to 4, characterized in that: After the step of optimizing each of the components according to the optimization strategy to complete the inference acceleration optimization of the target model, the method further includes: Recording the overall optimization strategy for the inference acceleration optimization of the target model; After receiving the optimization request for the target model, the target model is optimized according to the overall optimization strategy.
6. A model optimization device, characterized in that: The model optimization device comprises: An identification module, used to identify the framework type and structure type of each component in the target model to be optimized; An optimization strategy matching module, used to call a matching optimization strategy for each of the components according to a preset optimization strategy mapping table, wherein the optimization strategy is stored in the optimization strategy mapping table in a mapping manner with the framework type and structure type of the component; An optimization module, used to optimize each of the components according to the optimization strategy to complete the reasoning acceleration optimization of the target model; The optimization strategy matching module is also used for: For the target component in each component of the target model to be optimized, according to the framework type and structure type of the target component, according to the preset optimization strategy mapping table, determine the optimization strategy that matches the target component, wherein the optimization strategies in the optimization strategy mapping table include native optimization strategies and customized optimization strategies, and the native optimization strategy is matched to the target component whose framework type is the ONNX framework, and the customized optimization strategy is matched to the component whose framework type is the PyTorch framework or the Tensorflow framework; When there are multiple matching optimization strategies, a target optimization strategy is selected from the multiple matching optimization strategies according to the received optimization request conditions and the running time of the optimization strategies; The target optimization strategy is invoked for the target component.
7. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the model optimization method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the model optimization method according to any one of claims 1 to 4 are implemented.
9. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the model optimization method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Deep learning neural network deployment system and method
CN110942139A
Method, apparatus and device for optimizing compiler based on tensor data calculation inference
US20240061661A1