Deep learning model multi-NPU end side one-button deployment system and method
Through the one-click deployment system and method of multiple NPU terminals for deep learning models, the problem of unified deployment of multiple NPU platforms is solved, an efficient and automated one-stop deployment process is realized, domestic NPUs are supported, the deployment threshold is lowered, and development efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202511367492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing deep learning model deployment solutions lack unified deployment capabilities for multiple NPU platforms, are difficult to adapt to the deployment tool chains and operator specifications of different manufacturers, lack a highly integrated and automated one-stop deployment process, and do not provide broad support for mainstream domestic NPUs. They also lack a simple and easy-to-use visual front-end interface, which increases the deployment threshold.
This paper provides a one-click deployment system and method for deep learning models on multiple NPUs. Through a front-end interaction module, a server-side control module, an edge-side execution module, and a result integration module, it realizes automated deployment of models, supports unified deployment on multiple NPU platforms, and integrates a visual front-end interface to reduce the operation threshold.
It enables automated model deployment on various NPU devices, improving deployment efficiency, supporting parallel deployment of multiple types of NPUs, providing a unified management and control platform, simplifying the analysis process, lowering the operational threshold, and expanding application scenarios.
Smart Images

Figure CN120848903A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning model deployment technology on multiple NPUs, and particularly to a one-click deployment system and method for deep learning models on multiple NPUs. Background Technology
[0002] In the field of artificial intelligence model deployment technology, deep learning models are increasingly widely used in embedded neural network processor (NPU) edge devices, covering multiple key scenarios such as intelligent monitoring, industrial inspection, and edge computing. With technological advancements, the application demand for heterogeneous NPU terminals continues to grow, but current deep learning model deployment technologies on multiple NPU edge devices still have significant limitations, making it difficult to meet the requirements for efficient, universal, and convenient deployment.
[0003] In existing technologies, some solutions have attempted to optimize the model deployment process. For example, invention application CN116841911A discloses a "model testing method, heterogeneous chip, device, and medium based on heterogeneous platforms." This method controls the model under test to perform inference on the edge side and acquire recorded data by calling the model testing program on the target testing device. After comparing the data with the server-side data, the test results are output. By using standardized testing interfaces, manual intervention is reduced, improving model testing efficiency. Invention application CN120145888A proposes a "TensorFlow-based NPU model optimization method, device, and storage medium." Through model parsing, conversion, and optimization processes, the TensorFlow model is automatically converted into an intermediate representation supported by the NPU. Deployment is completed by combining operator fusion, memory optimization, and quantization strategies, improving model conversion efficiency and making it suitable for rapid deployment and performance monitoring of large-scale models. In addition, invention application CN117785220A discloses an artificial intelligence one-stop deployment platform that integrates image processing, model generation and deployment modules to realize a closed loop of the entire process from model training to deployment, reduce operational complexity and help promote AI in edge computing scenarios.
[0004] However, the existing solutions still have shortcomings: First, most solutions only implement the testing or deployment process of the model under specific devices or frameworks, lacking unified deployment capabilities for multiple NPU platforms and making it difficult to adapt to deployment toolchains and operator specifications from different vendors. Second, existing methods typically design model deployment, testing, performance monitoring, quantization optimization, and other stages separately, failing to form a highly integrated, automated one-stop deployment process. Third, some solutions only support mainstream international NPU types (such as Nvidia), lacking broad support for mainstream domestic NPUs such as Huawei Ascend and Cambricon MLU, resulting in insufficient platform versatility and limiting their widespread application in various edge scenarios. Fourth, some existing solutions mainly exist in the form of toolchains, lacking a simple and easy-to-use visual front-end interface, making them unfriendly to non-professional users and increasing the deployment threshold.
[0005] Therefore, there is a need for a one-click deployment system for deep learning models across multiple NPUs on the edge, which has the ability to deploy uniformly across multiple NPU platforms, improves the platform's versatility, and meets the promotion and application needs of various edge scenarios. Summary of the Invention
[0006] To address the aforementioned problems, the present invention aims to provide a one-click deployment system and method for deep learning models across multiple NPU platforms, solving issues such as insufficient unified deployment capability for various NPU platforms, lack of highly integrated automated one-stop deployment process, insufficient support for mainstream domestic NPUs, and lack of simple and easy-to-use visual front-end interface in existing deep learning model deployment solutions.
[0007] This invention provides a one-click deployment system and method for deep learning models using multiple NPUs on the edge.
[0008] First aspect: A one-click deployment system for deep learning models using multiple NPUs on the edge, comprising:
[0009] The front-end interaction module provides a visual interface for uploading inference images and configuring the target model, NPU type, and deployment parameters.
[0010] The server-side control module receives the front-end configuration, calls the unified version container environment to perform model transformation, and pushes the transformed model, inference script and image data to the client-side execution module.
[0011] The edge execution module receives the file from the server-side control module, starts the inference script, completes model loading, performs NPU edge inference, obtains the inference results, and monitors NPU performance synchronously.
[0012] The results integration module receives inference results and performance data returned by the execution module on the receiving end, generates a deployment report, and provides feedback to the user through the front-end interface.
[0013] In one embodiment of the present invention, the service-side control module includes a container environment management unit, which configures dedicated toolchains and automated scripts for different NPU types through the Conda environment isolation mechanism to form a unified version container environment and realize the automated execution of model conversion.
[0014] The second aspect: A one-click deployment method for deep learning models using multiple NPUs on the edge, including:
[0015] S1. On the user side, based on the front-end visual operation interface, the system accepts inference images uploaded by users and performs front-end configuration of the target model, NPU type and deployment parameters.
[0016] S2. On the service side, receive the front-end configuration, call the unified version container environment to perform model conversion, and push the converted model, inference script and image data to the client side.
[0017] S3. On the client side, receive the file pushed by the server, start the inference script, complete the model loading, perform NPU data inference, obtain the inference results, and monitor the NPU performance simultaneously.
[0018] S4. On the user side, the inference results and performance data returned by the receiving end are used to generate a deployment report, which is then fed back to the user through the front-end interface.
[0019] In one embodiment of the present invention, step S1 includes the following steps:
[0020] S11. Start the front-end interactive visual interface and create an appropriate number of processes;
[0021] S12. Select the NPU type and target model type, and configure the front-end parameters;
[0022] S13. Determine the model source;
[0023] S14. Select the inference image directory;
[0024] S15. Configure SSH connection information on the server side and the client side.
[0025] In one embodiment of the present invention, step S2 includes the following steps:
[0026] S21. Receive front-end configuration parameters;
[0027] S22. Detect the container environment's running status and automatically start the container environment;
[0028] S23. Dynamically select the corresponding Conda environment within the container environment based on the front-end configuration parameters;
[0029] S24. Based on the model source, upload or download the model file. The container environment executes dynamic conversion commands according to the target model type, NPU type, and inference mode.
[0030] S25. Push the converted model, inference script, and image data to the client side.
[0031] In one embodiment of the present invention, S24 includes the step of:
[0032] S24a. Call the corresponding deployment script according to the NPU type, convert the model, generate inference scripts and preprocess image data;
[0033] S24b: Upload the transformation model, inference script, and preprocessed image data to the specified directory on the client side.
[0034] In one embodiment of the present invention, calling the corresponding deployment script in S24a includes:
[0035] Load the NPU toolchain and its dependencies, convert the model format based on the NPU toolchain and its dependencies, and perform model quantization.
[0036] In one embodiment of the present invention, step S3 includes the following steps:
[0037] S31. Initialize the NPU runtime environment;
[0038] S32. Receive files pushed by the server and allocate input / output buffers;
[0039] S33. Start the inference script and execute inference calculations;
[0040] S34. Obtain inference results and NPU performance data;
[0041] S35, release NPU resources.
[0042] In one embodiment of the present invention, S4 includes the step of:
[0043] S41. Receive and parse inference results and NPU performance data;
[0044] S42. Generate operator distribution charts and layer-by-layer time consumption analysis charts;
[0045] S43. Integrate charts and generate deployment reports;
[0046] S44. Display based on the front end.
[0047] Third aspect: An electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the steps of the method provided in the second aspect.
[0048] Fourth aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the second aspect.
[0049] The beneficial effects of this invention are:
[0050] 1. This invention integrates a front-end interactive interface, a server-side control module, a container environment management unit, an edge-side execution module, and a result integration module to achieve automated model deployment on various NPU devices. It enables one-stop automated deployment, eliminating the need for manual intervention throughout the entire process from front-end configuration to edge-side execution and result analysis, significantly improving efficiency.
[0051] 2. The deployment system and method of the present invention have heterogeneous NPU adaptation capabilities. Through containerization isolation and dedicated scripts, they shield the differences in hardware from different manufacturers and support the parallel deployment of multiple types of NPUs.
[0052] 3. The deployment system and method of the present invention provide a unified management and control platform, support multi-device collaborative scheduling, automatic collection and visualization of performance data, and simplify the analysis process.
[0053] 4. The deployment system and method of the present invention lower the operating threshold, and the combination of a visual front-end interface and plug-and-play installation method allows non-professional users to get started quickly, thus expanding the application scenarios. Attached Figure Description
[0054] Figure 1 This is a diagram of the automated deployment module of the deployment system of the present invention;
[0055] Figure 2 This is a flowchart illustrating the automated deployment process of the deployment method of the present invention.
[0056] Figure 3 This is a schematic diagram of the front-end module flow of the present invention;
[0057] Figure 4 This is a schematic diagram of the service-side control module of the present invention;
[0058] Figure 5 This is a schematic diagram of the container environment management unit process of the present invention;
[0059] Figure 6 This is a schematic diagram of the end-side execution module of the present invention;
[0060] Figure 7 This is a schematic diagram of the process for integrating the results of this invention;
[0061] Figure 8 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation
[0062] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar symbols denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0063] Existing solutions suffer from the following problems: lack of unified deployment capabilities for multiple NPU platforms, difficulty in adapting to deployment toolchains and operator specifications from different vendors, inability to form a highly integrated, automated one-stop deployment process, and lack of a simple and easy-to-use visual front-end interface.
[0064] To address the above problems, this invention provides a one-click deployment system and method for deep learning models using multiple NPUs on the edge.
[0065] First, let's introduce the hardware and software environments.
[0066] For the hardware environment, users can use ordinary computers or mobile terminal devices. The server-side equipment is configured with a high-performance computer as the control center, supporting virtual machine installation and containerized deployment; the edge-side equipment includes various NPU development boards, such as Ascend 310 and Rockchip RK3588, which communicate with the server via network connection; auxiliary equipment includes storage media and connecting cables to ensure stable data transmission and communication between systems.
[0067] The software environment includes a server-side operating system that supports Ubuntu virtual machines, and a multi-NPU deployment toolchain built using a pre-installed containerized environment. The client-side devices are configured with network connectivity and SSH service by burning the official base image. The data and models are as follows: the models to be deployed are deep learning model files in ONNX format; the inference dataset is an image dataset used for testing.
[0068] Example 1:
[0069] like Figure 1 As shown in the figure, this embodiment discloses a one-click deployment system for deep learning models using multiple NPUs on the client side. It includes: a front-end interaction module, a server-side control module, a client-side execution module, and a result integration module.
[0070] The front-end interaction module provides a visual interface for uploading inference images and configuring the target model, NPU type, and deployment parameters.
[0071] The front-end interaction module provides a visual operation interface, allowing users to select the target model and its source (e.g., a preset model library or local upload), select the target NPU type (e.g., Ascend 310, Rockchip RK3588, etc.), upload inference images and configure deployment parameters (e.g., quantization switch, process specification, etc.), and supports local caching of configuration information; the installer can be plugged and played via USB device, automatically completing the deployment of dependent environments.
[0072] The server-side control module receives the front-end configuration, calls the unified version container environment to perform model transformation, and pushes the transformed model, inference script and image data to the client-side execution module.
[0073] The service-side control module includes a container environment management unit. The container environment management unit uses the Conda environment isolation mechanism to configure dedicated toolchains and automated scripts for different NPU types, forming a unified version container environment and realizing the automated execution of model conversion.
[0074] The server-side control module verifies connectivity with the edge-side NPU device and container environment via the SSH protocol, calls the container environment management unit to perform model transformation, and then pushes the transformed model, inference script and data to the preset directory on the edge via the SCP protocol to trigger the deployment process.
[0075] The endpoint execution module receives the file from the server-side control module, starts the inference script, completes model loading, performs NPU data inference, obtains the inference results, and simultaneously monitors the NPU device performance.
[0076] The edge execution module reads the configuration through command line parameters, completes model loading, data preprocessing (such as size adjustment and normalization), performs NPU inference, and post-processes the inference results; and simultaneously starts an independent performance monitoring process to collect NPU performance data such as NPU / CPU utilization, temperature and inference time, as well as model inference performance data such as model operator distribution and layer-by-layer time.
[0077] The results integration module receives inference results and performance data returned by the execution module on the receiving end, generates a deployment report, and provides feedback to the user through the front-end interface.
[0078] The inference results (such as labeled images, classification probabilities, etc.) and performance data returned by the receiving end are used to generate a deployment report containing visual charts (model operator distribution, layer-by-layer time consumption), which is then fed back to the user through the front-end interface.
[0079] Example 2:
[0080] like Figure 2 As shown, this embodiment discloses a one-click deployment method for deep learning models using multiple NPUs on the edge, including the following steps:
[0081] S1. On the user side, based on the front-end visual operation interface, users upload inference images and accept front-end configurations of the target model, NPU type, and deployment parameters, such as... Figure 3 As shown:
[0082] First, install the front-end interactive program, complete the system environment configuration and shortcut creation, and start the front-end interactive visual interface.
[0083] Next, select the NPU type and target model type, and configure the system front-end, including:
[0084] The process management configuration allows you to create and manage multiple deployment process tabs, supporting concurrent multi-task deployment operations.
[0085] Device connection configuration: Configure the connection parameters of the server-side Ubuntu virtual machine, set the network connection information of the end-side NPU device, and verify the connection status and communication availability of the NPU device;
[0086] Deployment parameter configuration includes selecting the target NPU type, specifying the deep learning model to be deployed, and configuring model quantization and pre- and post-processing options.
[0087] Next, determine the model source. The model source can be a self-trained model or a preset model. If a self-trained model is used, upload the model. If a self-trained model is not used, use the preset model by configuring the deployment parameters.
[0088] Next, select the inference image directory, set the inference image data path, and determine the specific location of the image data used for inference.
[0089] Finally, configure the SSH connection information on the server side and the client side. The system automatically saves the configuration parameters and starts the one-click automated deployment process.
[0090] S2. On the server side, receive the frontend configuration, invoke the unified version container environment to perform model transformation, and push the transformed model, inference script, and image data to the client side, such as... Figure 4 As shown:
[0091] First, multi-process configuration management is initialized. When the system starts, a multi-process configuration management mechanism is established. A process isolation architecture is used to create an independent configuration instance for each tab, ensuring configuration independence and functional consistency between processes.
[0092] Then, the running status of the unified version container environment is detected via the SSH protocol, and the container environment is automatically started. Simultaneously, automated control of the Ubuntu virtual machine on the server side is implemented, establishing SSH connections, supporting multi-process concurrent connections, and implementing connection timeout control and abnormal reconnection mechanisms. The running status of the target container environment is detected via the SSH connection; if the container is not running, it is automatically started, verifying the availability of the container environment and ensuring that the subsequent deployment environment is ready.
[0093] Then, based on the NPU type, the corresponding Conda runtime environment within the container environment is dynamically selected, while maintaining the container environment mapping table to ensure that the dependency library versions are correct during deployment.
[0094] Then, the initial model file is automatically loaded, and image data preprocessing is performed. The container environment executes dynamic transformation commands based on the model type, NPU type, and inference mode.
[0095] Furthermore, the system checks if the required model file exists in the target path within the container environment. If not, it downloads the file automatically, supporting resumeable downloads and error retries. For different NPU device types, it implements dedicated image preprocessing algorithms for each NPU device, ensuring image format compatibility. Preprocessed images are automatically cleaned up after deployment.
[0096] Furthermore, commands can be dynamically assembled and executed based on parameters such as NPU type, model type, and inference mode, and advanced parameters such as inference-specific mode and custom configuration can be used.
[0097] Finally, the transformed model, inference script, and data are pushed to the edge.
[0098] Furthermore, the container environment executes dynamic transformation commands based on model type, NPU type, and inference mode, such as... Figure 5 As shown, including:
[0099] First, the Conda environment is automatically matched based on the front-end configuration parameters.
[0100] The system receives the front-end configuration parameters transmitted by the server-side control module, selects the corresponding Conda runtime environment according to the NPU device type, and completes the Conda virtual environment switching.
[0101] Based on model name mapping rules, the target working directory path is dynamically obtained to achieve standardized model name mapping and support user-defined path overriding.
[0102] Then, based on the NPU type, the corresponding deployment script is called to transform the model, generate inference scripts, and preprocess image data.
[0103] The process of calling the corresponding deployment script includes: loading the NPU toolchain and dependency libraries, converting the model format based on the NPU toolchain and dependency libraries, and performing model quantization.
[0104] The corresponding deployment script is invoked based on the NPU platform type, conversion parameters are set for different model types, conversion commands are executed asynchronously, and the conversion progress is monitored to ensure that the target format model file is successfully generated.
[0105] Furthermore, an SSH connection channel is established to create a persistent interactive session, enabling unified remote control of NPU devices across multiple platforms and supporting the continuous execution and automated orchestration of complex command sequences.
[0106] Finally, upload the transformed model, inference script, and preprocessed image data to the designated directory on the device. Upload the local dataset, target model, and inference script to the designated directory on the NPU device, and simultaneously clean up historical data in the directory.
[0107] S3. On the client side, receive the file pushed by the server, start the inference script, complete model loading, perform NPU data inference, obtain the inference results, and simultaneously monitor NPU performance, such as... Figure 6 As shown:
[0108] First, initialize the NPU runtime environment; perform NPU runtime environment initialization operations, initialize the corresponding runtime library according to the NPU platform type, establish device context management, and ensure the correct allocation and release of resources.
[0109] Then, it receives files pushed by the server, allocates input / output buffers, performs model file loading, obtains model identifiers for subsequent inference calls, and creates model descriptors to obtain detailed model information, including key parameters such as input / output dimensions, data types, and memory layout.
[0110] Pre-allocate device memory buffers according to the model's input and output specifications, obtain the input tensor dimensions, allocate buffers of the appropriate size in the NPU device memory, and establish a data transmission channel from the host to the device.
[0111] Meanwhile, the input image data is preprocessed by sequentially performing image reading, size adjustment, color space conversion, dimension transformation, data type conversion, and normalization to ensure that the input data meets the model requirements.
[0112] Then, the inference script is started, inference calculations are executed, the model and data transmitted by the container environment management unit are obtained and inference calculations are started, an asynchronous inference execution mechanism is implemented, batch data processing and pipeline optimization are supported, and the inference execution status is monitored at the same time.
[0113] Then, the inference results and NPU performance data are acquired, the inference results are transferred from the edge memory to the edge CPU and structured parsing is performed, the structured processing is performed according to the model output format, confidence filtering, coordinate transformation, non-maximum suppression and class decoding are performed, and the final detection results are generated.
[0114] Real-time collection of inference execution performance data, recording the total time from data transmission to inference completion, and collecting key performance indicators such as layer-by-layer execution time and operator memory usage during model inference.
[0115] Finally, release the resources occupied by the NPU. Perform a complete resource cleanup process, sequentially executing operations such as unloading the model, destroying the model descriptor, releasing device memory, and cleaning up dataset objects, to ensure that the NPU device state is completely reset.
[0116] S4. On the user side, the inference results and performance data returned from the receiving end are used to generate a deployment report, which is then fed back to the user through the front-end interface, such as... Figure 7 As shown:
[0117] First, the system receives and parses inference results and NPU performance data. After the inference execution is completed on the edge, a multi-channel data receiving mechanism is established to simultaneously acquire inference results and performance data, process different types of data transmission, and ensure the complete transmission of inference results and performance data.
[0118] Then, operator distribution charts and layer-by-layer time consumption analysis charts are generated. Data is classified, parsed, and format converted according to data type. Different data formats are identified and parsed accordingly. The inference result images are format verified and path managed. Performance data is parsed and converted into a structured format.
[0119] At the same time, a multi-process isolated data storage mechanism is established to ensure that data from different deployment tasks are managed independently, process-specific storage directories are created, and structured data objects are maintained in memory to support subsequent statistical analysis and visualization.
[0120] Then, the charts are integrated to generate a deployment report. Distributed charts are generated based on operator type and memory usage. Statistical analysis of operator data is performed to calculate the proportion and resource consumption of each type of operator. A pie chart is generated to show the proportion of operator types, and an interactive hover function is added.
[0121] Based on hierarchical execution time data, a time consumption analysis chart is generated. Operators are sorted according to memory usage or execution time to identify performance bottlenecks. A bar chart is used to display the memory consumption of each operator, and a mouse hover interaction function is implemented.
[0122] Create a multi-tab user interface that integrates analysis results from different dimensions. Use a GUI framework to build an interactive interface that allows users to switch between different analysis views, embed charts, and support chart zooming, scrolling, and interactive operations.
[0123] It integrates inference results, performance data, and visualization charts to generate deployment reports containing complete analytical information, summarizes various data metrics, forms structured report content, and supports formatted output and chart export functions.
[0124] Finally, the analysis results are displayed on the front end, providing interactive operation functions, establishing a user-friendly feedback mechanism, and supporting functions such as result viewing, data export, and report sharing, thus achieving a complete user experience from data analysis to result display.
[0125] This embodiment demonstrates an automated deployment system for deep learning models that integrates multi-NPU platform support and intelligent deployment algorithms. In a real heterogeneous hardware environment, the system, through a unified front-end interface and back-end deployment engine, achieves automatic model conversion, remote deployment, performance monitoring, and result visualization analysis across multiple NPU platforms. The system can run stably in complex multi-platform environments. Through the collaborative work of its core modules, it provides users with a fully automated service from model upload to performance report generation, effectively lowering the deployment threshold for deep learning models on different NPU hardware and improving development efficiency and resource utilization.
[0126] The present invention also provides an electronic device, Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 8 As shown, the electronic device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory, for example, to execute the following method:
[0127] S1. On the user side, based on the front-end visual operation interface, users can configure the target model, NPU type, uploaded inference images, and configuration deployment parameters.
[0128] S2. On the service side, receive the front-end configuration, call the unified version container environment to perform model conversion, and push the converted model, inference script and image data to the client side.
[0129] S3. On the client side, receive the file pushed by the server, start the inference script, complete the model loading, perform NPU data inference, obtain the inference results, and monitor the NPU performance simultaneously.
[0130] S4. On the user side, the inference results and performance data returned by the receiving end are used to generate a deployment report, which is then fed back to the user through the front-end interface.
[0131] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0132] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments, including, for example:
[0133] S1. On the user side, based on the front-end visual operation interface, users can configure the target model, NPU type, uploaded inference images, and configuration deployment parameters.
[0134] S2. On the service side, receive the front-end configuration, call the unified version container environment to perform model conversion, and push the converted model, inference script and image data to the client side.
[0135] S3. On the client side, receive the file pushed by the server, start the inference script, complete the model loading, perform NPU data inference, obtain the inference results, and monitor the NPU performance simultaneously.
[0136] S4. On the user side, the inference results and performance data returned by the receiving end are used to generate a deployment report, which is then fed back to the user through the front-end interface.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A one-click deployment system for deep learning models using multiple NPUs on the edge, characterized in that, include: The front-end interaction module provides a visual operation interface, allows users to upload inference images, and configure the target model, NPU type, and deployment parameters. The server-side control module receives the front-end configuration, calls the unified version container environment to perform model transformation, and pushes the transformed model, inference script and image data to the client-side execution module. The endpoint execution module receives the file from the server-side control module, starts the inference script, completes model loading, performs NPU data inference, obtains the inference results, and monitors NPU performance synchronously. The results integration module receives inference results and performance data returned by the execution module on the receiving end, generates a deployment report, and provides feedback to the user through the front-end interface.
2. The system according to claim 1, characterized in that, The server-side control module includes: The container environment management unit uses the Conda environment isolation mechanism to configure dedicated toolchains and automated scripts for different NPU types, forming a unified version container environment and enabling automated execution of model conversion.
3. A deployment method applied to the system described in claim 1 or 2, characterized in that, Including the following steps: S1. On the user side, based on the front-end visual operation interface, the system accepts inference images uploaded by users and performs front-end configuration of the target model, NPU type and deployment parameters. S2. On the service side, receive the front-end configuration, call the unified version container environment to perform model conversion, and push the converted model, inference script and image data to the client side. S3. On the client side, receive the file pushed by the server, start the inference script, complete the model loading, perform NPU client-side inference, obtain the inference results, and monitor the NPU performance simultaneously. S4. On the user side, the inference results and performance data returned by the receiving end are used to generate a deployment report, which is then fed back to the user through the front-end interface.
4. The deployment method according to claim 3, characterized in that, S1 includes the following steps: S11. Start the front-end interactive visual interface and start an appropriate number of processes; S12. Select the NPU type and target model type, and configure the front-end parameters; S13. Determine the model source; S14. Select the inference image directory; S15. Configure SSH connection information on the server side and the client side.
5. The deployment method according to claim 4, characterized in that, S2 includes the following steps: S21. Receive front-end configuration parameters; S22. Detect the container environment's running status and automatically start the container environment; S23. Dynamically select the corresponding Conda environment within the container environment based on the front-end configuration parameters; S24. Based on the model source, upload or download the model file. The container environment executes dynamic conversion commands according to the target model type, NPU type, and inference mode. S25. Push the converted model, inference script, and image data to the client side.
6. The deployment method according to claim 5, characterized in that, S24 includes the following steps: S24a. Call the corresponding deployment script based on the NPU type to convert the model, perform quantization, and preprocess image data; S24b: Upload the converted model, inference script, and preprocessed image data to the specified directory on the client side.
7. The deployment method according to claim 6, characterized in that, The S24a call to the corresponding deployment script includes: Load the NPU toolchain and its dependencies, convert the model format based on the NPU toolchain and its dependencies, and perform model quantization.
8. The deployment method according to claim 5, characterized in that, S3 includes the following steps: S31. Initialize the NPU runtime environment; S32. Receive files pushed by the server and allocate input / output buffers; S33. Start the inference script and execute inference calculations; S34. Obtain inference results and NPU performance data; S35, release NPU resources.
9. The deployment method according to claim 8, characterized in that, S4 includes the following steps: S41. Receive and parse inference results and NPU performance data; S42. Generate operator distribution charts and layer-by-layer time consumption analysis charts; S43. Integrate charts and generate deployment reports; S44. Display based on the front end.
Citation Information
Patent Citations
Model testing method based on heterogeneous platform, heterogeneous chip, equipment and medium
CN116841911A
Artificial intelligence one-stop deployment platform
CN117785220A
NPU model optimization method and device based on TensorFlow and storage medium
CN120145888A
Optimization method and device of machine learning model and electronic equipment
CN115564060A
Model end side deployment method and device, equipment and storage medium
CN115756516A
Cited By
End-to-end zero code AI model automatic training and deployment system and method
CN121050735A