Model optimization generation method and device, equipment, storage medium and program product

By performing structural analysis and multiple rounds of optimization of the AI ​​model, combined with the chip hardware characteristics, model codes adapted to different chips are generated, which solves the problems of low operating efficiency and poor compatibility of AI models on different chips, and achieves efficient and stable cross-platform operation.

CN120124696APending Publication Date: 2025-06-10INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510094954.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When existing AI models run on different chips, there are problems such as performance degradation, unstable operation or even inability to run normally, and poor model compatibility, which increases the complexity of development and deployment.

Method used

By performing structural analysis of the optimization model, structural information is obtained, and based on the operator optimization strategy and the hardware characteristics of the target chip, the model is optimized and adjusted multiple rounds, and finally a model code can be generated that can run on the target chip.

Benefits of technology

It improves the operating efficiency and compatibility of AI models on different chips, avoids the problem of unstable or inability to run on the target chips, and simplifies the development and deployment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124696A_ABST
    Figure CN120124696A_ABST
Patent Text Reader

Abstract

The invention provides a model optimization generation method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: receiving a to-be-optimized model; performing structured analysis on the to-be-optimized model to obtain structural information of the to-be-optimized model; according to the structure information, based on an operator optimization strategy, optimizing the to-be-optimized model to obtain a first optimization model; based on hardware characteristics of the target chip, adjusting the first optimization model to obtain a second optimization model; and compiling the second optimization model to generate a target model code. Through the mode, the model code capable of running on the target chip can be generated according to the hardware characteristics of the target chip, the model compatibility is improved, the problem that the model runs unstably and even cannot run normally on the target chip is avoided, model optimization is carried out according to the structure information of the model and the operator optimization strategy, and the model optimization efficiency is improved. The operation efficiency of the model can be improved, and efficient reasoning of the model on a target chip is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, storage medium and program product for model optimization and generation. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, AI models have been widely used in various scenarios, such as computer vision, natural language processing, intelligent monitoring and other fields. The inference and training of AI models require powerful computing resources, which has promoted the development of dedicated artificial intelligence acceleration chips. In recent years, various chips have gradually occupied an important position in the AI field with their efficient computing capabilities. However, due to significant differences that may exist in hardware architecture, operator support, memory management, etc. among different chips, when existing AI models are directly transplanted to different chips, they often face problems such as performance degradation, unstable operation, or even inability to run properly.

[0003] In the actual application of AI models, especially complex deep learning models (such as YOLOv5), the performance of AI models highly depends on the computing efficiency of the underlying hardware of the chips. Due to different designs of the underlying hardware among different chips, this directly leads to large differences in the running efficiency of the same AI model on different chips, and it is difficult for AI models to achieve consistent and excellent running effects on different chips. In order to achieve the efficient operation of AI models on different chips, for each type of chip, special optimization is required. This not only increases the complexity of AI model development and deployment, but also creates obstacles to the cross-platform compatibility of AI models. At the same time, this also poses higher requirements for developers. Developers need to deeply understand the structure of AI models and be familiar with the characteristics and tuning methods of different chips.

[0004] Therefore, the existing model optimization and generation methods have problems of low model running efficiency and poor model compatibility. Summary of the Invention

[0005] The present invention provides a method, device, equipment, storage medium and program product for model optimization and generation to solve the problems of low model running efficiency and poor model compatibility existing in the existing model optimization and generation methods.

[0006] The present invention provides a method for model optimization and generation, including: receiving a model to be optimized; performing structural parsing on the model to be optimized to obtain the structure information of the model to be optimized; based on the structure information and an operator optimization strategy, optimizing the model to be optimized to obtain a first optimized model; adjusting the first optimized model based on the hardware characteristics of the target chip to obtain a second optimized model; compiling the second optimized model to generate target model code; the target model code is model code that can run on the target chip.

[0007] A model optimization generation method provided by the present invention, wherein the structure information includes the computation graph of the model to be optimized, and the computation graph is used to represent the operators, hierarchical structure, hierarchical structure dependency relationship, data flow, and computing units of the model to be optimized; based on the structure information and an operator optimization strategy, the model to be optimized is optimized to obtain a first optimized model, including: optimizing the operators of the model to be optimized based on a preset optimization strategy to obtain target operators; the preset optimization strategy includes at least one of a Winograd convolution optimization strategy, a GEMM optimization strategy, a mixed-precision calculation strategy, and a sparsity utilization strategy; determining the optimal computation path of the model to be optimized based on the chip characteristics of the target chip; rewriting and optimizing the computation graph based on the target operators and the optimal computation path to obtain a target computation graph; and optimizing the model to be optimized based on the target computation graph to obtain a first optimized model.

[0008] A model optimization generation method provided by the present invention, based on the hardware characteristics of the target chip, adjusts the first optimized model to obtain a second optimized model, including: optimizing and adjusting the memory access mode and computing task scheduling method of the first optimized model based on the hardware characteristics of the target chip to obtain a third optimized model; and optimizing the third optimized model based on the target hardware accelerator of the target chip to obtain a second optimized model.

[0009] A model optimization generation method provided by the present invention, wherein the structure information includes the intermediate representation of the model to be optimized; after compiling the second optimized model to generate target model code, it further includes: performing cross-platform compatibility checks on the target model code and the intermediate representation based on interface specifications and API adaptation specifications to obtain a compatibility check result.

[0010] A model optimization generation method provided by the present invention, after performing cross-platform compatibility checks on the target model code and the intermediate representation based on interface specifications and API adaptation specifications to obtain a compatibility check result, it further includes: if the compatibility check result is that the target model code and the intermediate representation pass the compatibility check, then performing functional verification and performance testing on the target model code, and evaluating the running performance of the target model code based on a performance monitoring tool to obtain performance data during the running process of the target model code; and iteratively optimizing the computation path, scheduling strategy, and quantization parameters of the target model code based on the performance data to obtain optimized model code.

[0011] A model optimization generation method provided by the present invention, after iteratively optimizing the computation path, scheduling strategy, and quantization parameters of the target model code based on the performance data to obtain optimized model code, it further includes: selecting a target device from multiple devices; the target device is a device configured with the target chip; and deploying the optimized model code to the target device.

[0012] The present invention also provides a model optimization and generation device, including: a receiving module, configured to receive a model to be optimized; a parsing module, configured to perform structured parsing on the model to be optimized to obtain the structure information of the model to be optimized; an operator optimization module, configured to optimize the model to be optimized based on the structure information and an operator optimization strategy to obtain a first optimized model; a hardware characteristic adaptation module, configured to adjust the first optimized model based on the hardware characteristics of a target chip to obtain a second optimized model; and a code generation module, configured to compile the second optimized model to generate target model code. The target model code is model code that can run on the target chip.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for model optimization and generation as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for model optimization and generation as described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for model optimization and generation as described in any one of the above is implemented.

[0016] For the model optimization and generation method, device, equipment, storage medium, and program product provided by the present invention, first, structured parsing is performed on the model to be optimized to obtain the structure information of the model to be optimized. Then, based on the structure information and the operator optimization strategy, the model to be optimized is optimized to obtain a first optimized model. And based on the hardware characteristics of the target chip, the first optimized model is adjusted to obtain a second optimized model. After compiling the second optimized model, model code that can run on the target chip is obtained. Through the above method, model code that can run on the target chip can be generated according to the hardware characteristics of the target chip, improving model compatibility and avoiding problems such as unstable operation or even inability to operate normally of the model on the target chip. At the same time, by optimizing the model according to the structure information of the model and the operator optimization strategy, the operation efficiency of the model can be improved, which is beneficial to the efficient inference of the model on the target chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of the model optimization generation method provided by the present invention.

[0019] Figure 2 It is a schematic structural diagram of the AI model cross-platform operator optimization and automatic generation system provided by the present invention.

[0020] Figure 3 It is a schematic working flowchart of the AI model cross-platform operator optimization and automatic generation system provided by the present invention.

[0021] Figure 4 It is a schematic structural diagram of the model optimization generation device provided by the present invention.

[0022] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Please refer to Figures 1 to 3 , Figure 1 It is a schematic flowchart of the model optimization generation method provided by the present invention, Figure 2 It is a schematic structural diagram of the AI model cross-platform operator optimization and automatic generation system provided by the present invention, Figure 3 It is a schematic working flowchart of the AI model cross-platform operator optimization and automatic generation system provided by the present invention. In this embodiment, the model optimization generation method is applied to the AI model cross-platform operator optimization and automatic generation system. The model optimization generation method includes steps S110 to S150, and the specific steps are as follows: S110: Receive the model to be optimized.

[0025] As Figure 2 shown, the AI model cross-platform operator optimization and automatic generation system includes a model input and parsing module, a general operator optimization module, a hardware feature adaptation module, an automatic code generation and tuning module, a cross-platform compatibility management module, a model verification and performance monitoring module, and a user interface and configuration management module. Each module is responsible for its specific function and cooperates with each other to achieve the full-process optimization and generation from model input to cross-platform efficient execution.

[0026] Among them, the functions and working processes of each module are as follows: (1)Model Input and Parsing Module: Functions of the Model Input and Parsing Module: The user inputs a standardized AI model (such as YOLOv5) into the AI Model Cross-Platform Operator Optimization and Automatic Generation System. The Model Input and Parsing Module is responsible for parsing the computational graph, operators, hierarchical structure, and their dependencies of the AI model, and extracting the core computational units and data streams of the AI model.

[0027] Workflow of the Model Input and Parsing Module: The functions of the Model Input and Parsing Module can perform a structured analysis on the input AI model to generate an intermediate representation (IR), which can describe the computational requirements of the model and provide basic information for subsequent optimization and generation processes.

[0028] (2)General Operator Optimization Module: Functions of the General Operator Optimization Module: The General Operator Optimization Module can optimize common operators in the AI model (such as convolution, matrix multiplication, pooling, etc.), and adopt general optimization strategies to adapt to different chip architectures. Optimization strategies include but are not limited to Winograd convolution optimization strategy, GEMM optimization strategy, mixed-precision calculation strategy, and sparsity utilization strategy, etc.

[0029] Workflow of the General Operator Optimization Module: The General Operator Optimization Module can select the optimal computational path according to the chip characteristics, and rewrite and optimize the computational graph of the AI model to reduce the computational complexity of the AI model and improve the execution efficiency of the AI model.

[0030] (3)Hardware Feature Adaptation Module: Functions of the Hardware Feature Adaptation Module: For different chips, the Hardware Feature Adaptation Module is responsible for further optimizing the AI model optimized by the General Operator Optimization Module to adapt to the specific chip hardware architecture, including adjusting the computational task scheduling method of the AI model, optimizing the memory access mode, and utilizing the specific hardware accelerators of the chip, etc.

[0031] Workflow of the Hardware Feature Adaptation Module: The Hardware Feature Adaptation Module will automatically select the best execution path according to the hardware characteristics of the chip, in order to generate targeted model formats and codes subsequently, ensuring that the AI model can run efficiently on the target chip.

[0032] (4)Automatic Code Generation and Tuning Module: Functions of the Automatic Code Generation and Tuning Module: The Automatic Code Generation and Tuning Module is responsible for converting the optimized and hardware-adapted AI model into specific model codes that can run on the target chip. The Automatic Code Generation and Tuning Module combines cross-platform compilation technologies and intelligent tuning algorithms to generate efficient and executable codes and perform automated performance tuning.

[0033] Workflow of the Automatic Code Generation and Optimization Module: The automatic code generation and optimization module compiles the AI model that has been optimized and adjusted for hardware adaptation, generates efficient model code adapted to a specific chip architecture, automatically analyzes the performance of the model code during runtime, and further optimizes the execution strategy through a feedback mechanism to ensure that the AI model can achieve the best performance on different chip platforms.

[0034] (5) Cross-Platform Compatibility Management Module: Functions of the Cross-Platform Compatibility Management Module: The cross-platform compatibility management module can ensure that the generated model code and model format have good cross-platform compatibility and can run smoothly on different chip platforms.

[0035] Workflow of the Cross-Platform Compatibility Management Module: When generating model code, the cross-platform compatibility management module follows unified interface specifications and API (Application Programming Interface) adaptation specifications to ensure that the intermediate representation of the AI model can adapt to the hardware architectures of different chips, and achieves seamless connection for cross-platform operation through scheduling control and the API adaptation layer.

[0036] (6) Model Verification and Performance Monitoring Module: Functions of the Model Verification and Performance Monitoring Module: The model verification and performance monitoring module is responsible for testing and validating the generated AI model and model code to ensure the correctness and efficiency of the AI model on the target chip. At the same time, the model verification and performance monitoring module can monitor the running performance of the AI model in actual applications and collect performance data for further optimization.

[0037] Workflow of the Model Verification and Performance Monitoring Module: The model verification and performance monitoring module can automatically execute a series of tests, including functional verification, performance testing, etc., and evaluate the running situation of the model in real time through performance monitoring tools, and feedback the performance data to the automatic code generation and optimization module for continuous optimization.

[0038] (7) User Interface and Configuration Management Module: Functions of the User Interface and Configuration Management Module: The user interface and configuration management module provides users with a convenient operation interface and configuration management tools. Users can upload AI models, select target chip platforms, view optimization results, and adjust configuration parameters through this module.

[0039] Workflow of the User Interface and Configuration Management Module: The user interface and configuration management module needs to be closely integrated with other modules of the system. Users can customize the optimization process of the AI model through this module, manage and view optimization results and performance reports on different chip platforms.

[0040] Through the collaborative work of the above-mentioned modules, the AI model cross-platform operator optimization and automatic generation system can achieve the full-process automation from the input of the AI model to the optimization and code generation adapted to multiple chip platforms, providing users with a comprehensive solution to simplify deployment and improve performance, and ensuring that the AI model can run efficiently and stably on different chip platforms.

[0041] Understandably, as Figure 3 shown, based on the above system, users can use a standardized AI model as the model to be optimized, and input the model to be optimized into the AI model cross-platform operator optimization and automatic generation system through the user interface and the configuration management module; the AI model cross-platform operator optimization and automatic generation system can receive the model to be optimized.

[0042] S120: Perform a structural analysis on the model to be optimized to obtain the structural information of the model to be optimized.

[0043] After the AI model cross-platform operator optimization and automatic generation system receives the model to be optimized, the model input and parsing module can perform a structural analysis on the model to be optimized to obtain the structural information of the model to be optimized.

[0044] Among them, the process of structural analysis can generate an intermediate representation (IR) of the model to be optimized, extract key computing units, data streams, and dependency relationships in the model to be optimized, etc., laying a foundation for subsequent optimization and code generation.

[0045] S130: Based on the structural information and the operator optimization strategy, optimize the model to be optimized to obtain the first optimized model.

[0046] After completing the structural analysis and obtaining the structural information of the model to be optimized, the model to be optimized will be sent to the general operator optimization module; the general operator optimization module can optimize the common operators in the model to be optimized through a series of preset optimization strategies, such as Winograd algorithm optimization, GEMM optimization, mixed-precision calculation optimization (FP16 / INT8 quantization), and sparsity utilization applicable to convolution operations. At the same time, the general operator optimization module can rewrite and optimize the computational graph of the model to be optimized, and optimize the model to be optimized based on the rewritten and optimized computational graph to obtain the first optimized model, thereby reducing the computational complexity, improving the execution efficiency of the model, and ensuring the cross-platform executability of the optimized model.

[0047] S140: Based on the hardware characteristics of the target chip, adjust the first optimized model to obtain the second optimized model.

[0048] After completing the operator optimization and obtaining the first optimized model, the first optimized model will be sent to the hardware feature adaptation module; the hardware feature adaptation module can further adjust the first optimized model according to the hardware features of different chips.

[0049] Specifically, the hardware feature adaptation module can optimize the memory access mode of the first optimized model, adjust the calculation task scheduling method of the first optimized model according to the hardware features of the target chip, and use the specific hardware accelerator of the target chip to optimize the first optimized model to obtain the second optimized model, ensuring that the second optimized model can make full use of the hardware advantages of different chips and achieve efficient operation.

[0050] S150: Compile the second optimized model to generate the target model code.

[0051] The target model code is the model code that can run on the target chip.

[0052] After completing the hardware adaptation and obtaining the second optimized model, the second optimized model will be sent to the automatic code generation and tuning module; the automatic code generation and tuning module can compile the second optimized model, convert the second optimized model into the model code that can run on the target chip, and perform intelligent tuning.

[0053] Through the above process, the AI model cross-platform operator optimization and automatic generation system can optimize and compile the standardized AI model input by the user, generate efficient code adapted to different chip architectures, and dynamically adjust the execution strategy of the AI model through the performance feedback mechanism to ensure that the AI model can achieve the best performance on different chip platforms.

[0054] The model optimization and generation method provided in this embodiment first performs a structural analysis on the model to be optimized to obtain the structural information of the model to be optimized, then optimizes the model to be optimized based on the structural information and the operator optimization strategy to obtain the first optimized model, and adjusts the first optimized model based on the hardware features of the target chip to obtain the second optimized model. After compiling the second optimized model, the model code that can run on the target chip is obtained. In this way, the model code that can run on the target chip can be generated according to the hardware features of the target chip, improving the model compatibility and avoiding the problem that the model runs unstably or even cannot run normally on the target chip. At the same time, optimizing the model according to the structural information and operator optimization strategy of the model can improve the running efficiency of the model and is beneficial to the model to achieve efficient inference on the target chip.

[0055] In some embodiments, the structural information includes the computation graph of the model to be optimized, and the computation graph is used to characterize the operators, hierarchical structure, hierarchical structure dependency relationships, data flow, and computing units of the model to be optimized; based on the structural information and an operator optimization strategy, the model to be optimized is optimized to obtain a first optimized model, including: optimizing the operators of the model to be optimized based on a preset optimization strategy to obtain target operators; the preset optimization strategy includes at least one of a Winograd convolution optimization strategy, a GEMM optimization strategy, a mixed-precision computing strategy, and a sparsity utilization strategy; determining the optimal computation path of the model to be optimized based on the chip characteristics of the target chip; rewriting and optimizing the computation graph based on the target operators and the optimal computation path to obtain a target computation graph; and optimizing the model to be optimized based on the target computation graph to obtain a first optimized model.

[0056] In this embodiment, the structural information includes the computation graph of the model to be optimized, and the computation graph is used to characterize the operators, hierarchical structure, hierarchical structure dependency relationships, data flow, and computing units of the model to be optimized.

[0057] It should be noted that in the field of artificial intelligence, an operator is a core concept. In deep learning and neural networks, an operator specifically manifests as a basic unit that implements a specific mathematical operation or logical operation. An operator defines the basic computation steps in a neural network model, including but not limited to matrix multiplication, weighted summation, activation functions (such as ReLU, sigmoid, etc.), pooling operations, convolution, etc.

[0058] Specifically, after completing the structural analysis and obtaining the structural information of the model to be optimized, the model to be optimized will be sent to a general operator optimization module; the general operator optimization module can optimize the operators of the model to be optimized based on a preset optimization strategy to obtain target operators.

[0059] Among them, the preset optimization strategy includes at least one of a Winograd convolution optimization strategy, a GEMM optimization strategy, a mixed-precision computing strategy, and a sparsity utilization strategy.

[0060] Optionally, the general operator optimization module can use the Winograd convolution optimization strategy to optimize the convolution operation (a type of operator) to reduce the number of multiplications in convolution calculations, thereby accelerating the convolution operation.

[0061] Optionally, the general operator optimization module can adopt the GEMM optimization strategy to convert the convolution operation into matrix multiplication and achieve optimization by optimizing the matrix multiplication, thereby improving the efficiency of the model in convolution calculations.

[0062] Optionally, the general operator optimization module can further optimize the computational efficiency of the model to be optimized through strategies such as mixed-precision computing (FP16 / INT8 quantization) and sparsity utilization, and ensure that the model can achieve the best performance on different chip platforms through dynamic tuning.

[0063] Furthermore, the general operator optimization module can select the optimal computational path for the model to be optimized based on the chip characteristics of the target chip, rewrite and optimize the computational graph using the target operator and the optimal computational path to obtain the target computational graph, and optimize the model to be optimized based on the target computational graph to obtain the first optimized model, so as to reduce the computational complexity of the model and improve the execution efficiency of the model.

[0064] In some embodiments, based on the hardware characteristics of the target chip, the first optimized model is adjusted to obtain the second optimized model, including: based on the hardware characteristics of the target chip, optimizing and adjusting the memory access mode and computational task scheduling method of the first optimized model to obtain the third optimized model; optimizing the third optimized model based on the target hardware accelerator of the target chip to obtain the second optimized model.

[0065] After completing operator optimization and obtaining the first optimized model, the first optimized model will be sent to the hardware characteristics adaptation module; the hardware characteristics adaptation module can further adjust the first optimized model according to the hardware characteristics of different chips and optimize the execution strategy of the model.

[0066] Specifically, the hardware characteristics adaptation module can optimize the memory access mode of the first optimized model, adjust the computational task scheduling method of the first optimized model according to the hardware characteristics of the target chip, and tune the first optimized model using the specific hardware accelerator of the target chip to obtain the second optimized model, ensuring that the second optimized model can make full use of the hardware advantages of different chips and achieve efficient operation.

[0067] Optionally, the hardware characteristics adaptation module can optimize the data access mode of the first optimized model according to the memory bandwidth and cache structure of the target chip to reduce the latency of data transmission and improve the inference efficiency of the model.

[0068] Furthermore, the automatic code generation and tuning module can automatically compile and generate efficient model code adapted to different chip architectures according to the computational graph of the second optimized model, and dynamically adjust the computational path and execution strategy of the model based on the performance feedback of the model code during runtime to achieve the best performance.

[0069] In some embodiments, the structural information includes the intermediate representation of the model to be optimized; after compiling the second optimized model to generate the target model code, it further includes: performing cross-platform compatibility checks on the target model code and the intermediate representation based on the interface specification and API adaptation specification to obtain the compatibility check result.

[0070] Specifically, after the target model code is generated, the target model code and the intermediate representation will be sent to the cross-platform compatibility management module; the cross-platform compatibility management module can perform cross-platform compatibility checks on the target model code and the intermediate representation through the interface specification and the API adaptation specification, and obtain the compatibility check result, so as to ensure that the generated target model code and format have good cross-platform compatibility, can run stably on different chip platforms, and ensure that the intermediate representation of the model can be adapted to different chip platforms and be correctly executed in different chip hardware environments.

[0071] In some embodiments, after performing cross-platform compatibility checks on the target model code and the intermediate representation based on the interface specification and the API adaptation specification and obtaining the compatibility check result, it further includes: if the compatibility check result is that the target model code and the intermediate representation pass the compatibility check, then perform functional verification and performance testing on the target model code, and evaluate the running performance of the target model code based on the performance monitoring tool to obtain the performance data of the target model code during operation; based on the performance data, perform iterative optimization on the calculation path, scheduling strategy, and quantization parameters of the target model code to obtain the optimized model code.

[0072] Specifically, if it is determined that the target model code and the intermediate representation pass the compatibility check, the model verification and performance monitoring module can perform functional verification and performance testing on the target model code, and the model verification and performance monitoring module of the system will automatically perform functional verification and performance testing to ensure that the target model code can be correctly executed on the target chip platform and meet the expected performance indicators.

[0073] Meanwhile, the model verification and performance monitoring module can evaluate the running performance of the target model code in real time based on the performance monitoring tool, collect the performance data of the target model code during operation, and send the collected performance data to the automatic code generation and tuning module.

[0074] Furthermore, the automatic code generation and tuning module can perform iterative optimization on the calculation path, scheduling strategy, and quantization parameters of the target model code based on the performance data; this process can be iterated multiple times to continuously optimize the execution strategy of the model, gradually improve the running efficiency of the target model code on the target chip platform, and obtain the optimized model code to ensure that the final optimized model code can achieve the optimal performance on the target chip platform.

[0075] In some embodiments, after performing iterative optimization on the calculation path, scheduling strategy, and quantization parameters of the target model code based on the performance data to obtain the optimized model code, it further includes: selecting a target device from multiple devices; the target device is a device configured with the target chip; deploying the optimized model code to the target device.

[0076] After obtaining the optimized model code, the optimized model code and the optimized and verified target model will be output through the user interface and the configuration management module.

[0077] Furthermore, the user can select a target device (i.e., the target chip platform) from multiple devices (i.e., chip platforms) and deploy the optimized model code to the corresponding target device.

[0078] Optionally, the AI model cross-platform operator optimization and automatic generation system can provide detailed deployment guidelines and configuration options to ensure that the model can be quickly launched and run stably.

[0079] The model optimization and generation method provided in this embodiment has at least the following technical advantages compared with the prior art: (1) Strong cross-platform compatibility: The AI model cross-platform operator optimization and automatic generation system provided in this embodiment can automatically convert the AI model into model code adapted to different chips, enhancing the cross-platform compatibility of the AI model; through a unified interface specification and API adaptation layer, it can ensure that the generated model code can run stably and efficiently on different chip platforms.

[0080] (2) Significantly improve the inference efficiency of the model: Through various operator optimization strategies (such as Winograd convolution optimization, GEMM optimization, mixed-precision calculation, sparsity utilization, etc.), the inference efficiency of the AI model is significantly improved, enabling the optimized AI model to fully utilize the hardware computing capabilities of the corresponding chips on different chips, greatly reducing the computing time and resource consumption, and enhancing the overall performance.

[0081] (3) Have the ability of automatic optimization and generation: This embodiment constructs an end-to-end AI model cross-platform operator optimization and automatic generation system, which can realize the whole process automation from model parsing to code generation and tuning. The user only needs to input a standardized AI model, and the system can automatically output optimized models and codes adapted to different chip platforms, simplifying the development process, reducing the workload of manual tuning, and improving the development efficiency.

[0082] (4) Have intelligent hardware adaptation ability: Aiming at the hardware architecture characteristics of different chips, this embodiment introduces an intelligent hardware adaptation mechanism, which can automatically adjust the execution strategy of the model according to the hardware characteristics of the chip, optimize the memory access mode and the calculation task scheduling method, and ensure that the target model code can achieve the best performance on different chip platforms. This intelligent hardware adaptation mechanism effectively solves the problem of performance loss in model transplantation.

[0083] (5) Improve the flexibility and stability of model deployment: Through cross-platform compatibility management, the model code and model format generated in this embodiment can run seamlessly on different chip architectures, significantly improving the flexibility and stability of model deployment. Developers can easily deploy the same model on multiple chip platforms without having to perform separate optimizations and debugging for each chip platform.

[0084] (6) Reduce development costs and time: Through automated processes and cross-platform optimization strategies, this embodiment significantly reduces the porting and tuning costs of AI models on different chips and shortens the development cycle. Developers no longer need to deeply understand the underlying architecture of each chip and can complete complex optimization tasks simply through the tools provided by the system.

[0085] (7) Support the popularization of large-scale AI applications: The methods and systems of this embodiment can effectively support the cross-platform deployment of large-scale AI applications, which are particularly suitable for application scenarios that need to run in multiple hardware environments, such as intelligent monitoring, autonomous driving, medical image processing, etc. Through efficient operator optimization strategies and hardware adaptation mechanisms, this embodiment provides technical support for the wide application of different chips in the field of AI.

[0086] The present invention also provides a model optimization and generation device. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the model optimization and generation device provided by the present invention. In this embodiment, the model optimization and generation device includes a receiving module 410, a parsing module 420, an operator optimization module 430, a hardware feature adaptation module 440, and a code generation module 450.

[0087] The receiving module 410 is used to receive the model to be optimized.

[0088] The parsing module 420 is used to perform a structural parsing on the model to be optimized to obtain the structural information of the model to be optimized.

[0089] The operator optimization module 430 is used to optimize the model to be optimized based on the operator optimization strategy according to the structural information to obtain a first optimized model.

[0090] The hardware feature adaptation module 440 is used to adjust the first optimized model based on the hardware features of the target chip to obtain a second optimized model.

[0091] The code generation module 450 is used to compile the second optimized model to generate target model code; the target model code is model code that can run on the target chip.

[0092] In some embodiments, the structural information includes the computational graph of the model to be optimized, and the computational graph is used to characterize the operators, hierarchical structures, hierarchical structure dependencies, data flows, and computing units of the model to be optimized.

[0093] The operator optimization module 430 is configured to optimize the operators of the model to be optimized based on a preset optimization strategy to obtain target operators; the preset optimization strategy includes at least one of a Winograd convolution optimization strategy, a GEMM optimization strategy, a mixed-precision computing strategy, and a sparsity utilization strategy; determine the optimal computing path of the model to be optimized based on the chip characteristics of the target chip; rewrite and optimize the computational graph based on the target operators and the optimal computing path to obtain a target computational graph; optimize the model to be optimized based on the target computational graph to obtain a first optimized model.

[0094] In some embodiments, the hardware characteristic adaptation module 440 is configured to optimize and adjust the memory access mode and the computing task scheduling method of the first optimized model based on the hardware characteristics of the target chip to obtain a third optimized model; optimize the third optimized model based on the target hardware accelerator of the target chip to obtain a second optimized model.

[0095] In some embodiments, the structural information includes the intermediate representation of the model to be optimized.

[0096] The model optimization and generation device further includes a cross-platform compatibility management module.

[0097] The cross-platform compatibility management module is configured to perform a cross-platform compatibility check on the target model code and the intermediate representation based on the interface specification and the API adaptation specification to obtain a compatibility check result.

[0098] In some embodiments, the model optimization and generation device further includes a model verification and performance monitoring module.

[0099] The model verification and performance monitoring module is configured to, if the compatibility check result indicates that the target model code and the intermediate representation pass the compatibility check, perform a function verification and a performance test on the target model code, and evaluate the running performance of the target model code based on a performance monitoring tool to obtain performance data during the running of the target model code; perform iterative optimization on the computing path, the scheduling strategy, and the quantization parameters of the target model code based on the performance data to obtain an optimized model code.

[0100] In some embodiments, the model optimization and generation device further includes a user interface and configuration management module.

[0101] The user interface and configuration management module is configured to select a target device from multiple devices; the target device is a device configured with the target chip; deploy the optimized model code to the target device.

[0102] The present invention also provides an electronic device. Figure 5 is a schematic structural diagram of the electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call logical instructions in the memory 530 to execute the model optimization generation method.

[0103] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0104] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the model optimization generation method provided by the above-mentioned methods is implemented.

[0105] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model optimization generation method provided by the above-mentioned methods.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0107] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model optimization generation method, characterized in that: include: Receive the model to be optimized; Performing structural analysis on the model to be optimized to obtain structural information of the model to be optimized; According to the structural information, based on an operator optimization strategy, the model to be optimized is optimized to obtain a first optimized model; Based on the hardware characteristics of the target chip, adjusting the first optimization model to obtain a second optimization model; Compiling the second optimization model to generate target model code; The target model code is a model code that can be run on the target chip.

2. The model optimization generation method according to claim 1, characterized in that: The structural information includes a computational graph of the model to be optimized, where the computational graph is used to characterize operators, hierarchical structures, hierarchical structure dependencies, data flows, and computational units of the model to be optimized; The step of optimizing the model to be optimized based on the structural information and an operator optimization strategy to obtain a first optimized model includes: Based on a preset optimization strategy, the operator of the model to be optimized is optimized to obtain a target operator; the preset optimization strategy includes at least one of a Winograd convolution optimization strategy, a GEMM optimization strategy, a mixed precision calculation strategy, and a sparsity utilization strategy; Based on the chip characteristics of the target chip, determining the optimal calculation path of the model to be optimized; Based on the target operator and the optimal calculation path, rewrite and optimize the calculation graph to obtain a target calculation graph; Based on the target calculation graph, the model to be optimized is optimized to obtain the first optimized model.

3. The model optimization generation method according to claim 1, characterized in that: The adjusting the first optimization model based on the hardware characteristics of the target chip to obtain the second optimization model includes: Based on the hardware characteristics of the target chip, the memory access mode and the computing task scheduling method of the first optimization model are optimized and adjusted to obtain a third optimization model; Based on the target hardware accelerator of the target chip, the third optimization model is optimized to obtain the second optimization model.

4. The model optimization generation method according to claim 1, characterized in that: The structural information includes an intermediate representation of the model to be optimized; After compiling the second optimization model to generate the target model code, the method further includes: Based on the interface specification and the API adaptation specification, a cross-platform compatibility check is performed on the target model code and the intermediate representation to obtain a compatibility check result.

5. The model optimization generation method according to claim 4, characterized in that: The cross-platform compatibility check of the target model code and the intermediate representation based on the interface specification and the API adaptation specification, and obtaining the compatibility check result, further includes: If the compatibility check result is that the target model code and the intermediate representation pass the compatibility check, then functional verification and performance testing are performed on the target model code, and the running performance of the target model code is evaluated based on a performance monitoring tool to obtain performance data of the target model code during the running process; Based on the performance data, the calculation path, scheduling strategy and quantization parameters of the target model code are iteratively optimized to obtain an optimized model code.

6. The model optimization generation method according to claim 5, characterized in that: After the calculation path, scheduling strategy and quantization parameter of the target model code are iteratively optimized based on the performance data to obtain the optimized model code, the method further includes: Select a target device from a plurality of devices; the target device is a device configured with the target chip; Deploy the optimized model code to the target device.

7. A model optimization generation device, characterized in that: include: A receiving module, used for receiving a model to be optimized; An analysis module, used for performing structural analysis on the model to be optimized to obtain structural information of the model to be optimized; An operator optimization module, configured to optimize the model to be optimized according to the structural information and based on an operator optimization strategy to obtain a first optimized model; A hardware characteristic adaptation module, used for adjusting the first optimization model based on the hardware characteristics of the target chip to obtain a second optimization model; A code generation module, used for compiling the second optimization model to generate target model code; The target model code is a model code that can be run on the target chip.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the model optimization generation method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model optimization generation method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the model optimization generation method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Model compiling method and device, electronic equipment and storage medium

    CN120973382A

  • Model compiling method and device, electronic equipment and storage medium

    CN120973382B