Model deployment method capable of configuring model, model deployment end and target equipment end

By registering a unified interface for the model and packaging it into compressed files, the configuration chaos caused by inconsistent interfaces in model deployment is solved, the convenience of interface standardization and model management is achieved, and compatibility and security are improved.

CN120523477APending Publication Date: 2025-08-22VERISILICON TECH (SHANGHAI) CO LTD +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510606595.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When different models are deployed on the target device, configuration confusion caused by inconsistent interface definitions, serious compatibility problems, and affect management and scalability.

Method used

By registering a target interface instance for the model to be deployed, the preset unified interface is bound to the interface in the target interface file, and the model file is packaged into a compressed file, and the preset unified interface is used to call and manage it on the target device side.

Benefits of technology

It realizes standardization and unification of interfaces, reduces compatibility issues, facilitates model management and expansion, improves development and maintenance efficiency, and enhances the flexibility and security of model deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523477A_ABST
    Figure CN120523477A_ABST
Patent Text Reader

Abstract

The invention provides a model deployment method capable of configuring a model, a model deployment end and a target equipment end, and the model deployment method applied to the model deployment end comprises the steps: obtaining a to-be-deployed file and a target interface file of a to-be-deployed model; registering a target interface instance of the to-be-deployed model; compiling the registered target interface file into a model dynamic library; packaging the to-be-deployed file into a compressed file; and sending the model dynamic library and the compressed file to the target equipment end. According to the scheme, the target interface instance is registered for the to-be-deployed model, so that the preset unified interface is bound with the interface in the target interface file, an application program can call interface modes of different models through the unified interface, implementation details of the models can not be concerned any more, standardization and unification of the interface are facilitated, and the deployment efficiency is improved. The compatibility problem caused by model interface differences is avoided, the application program can integrate and call different models more flexibly, and management and expansion of the models are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model deployment method, a model deployment terminal, and a target device terminal for a configurable model. Background Art

[0002] With the rapid development of artificial intelligence technology, the importance of model deployment has become increasingly prominent. Model deployment is not only about applying trained algorithms to real-world scenarios, but also a key step in ensuring that these models can achieve maximum effectiveness in production environments.

[0003] To deploy AI models on target devices, technologies typically require a complete set of initialization, pre-processing, post-processing, and de-initialization methods for the model being deployed. However, different models often have unique interface definitions, and these inconsistent interfaces can lead to configuration confusion when managing multiple models on the target device. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model deployment method, a model deployment terminal and a target device terminal of a configurable model to solve the above problems.

[0005] In the first aspect, an embodiment of the present application provides a model deployment method for a configurable model, which is applied to a model deployment end, and the method includes: obtaining a to-be-deployed file and a target interface file of the to-be-deployed model; wherein the to-be-deployed model is a configurable model; registering a target interface instance of the to-be-deployed model to bind a preset unified interface to an interface method in the target interface file; wherein the interface method includes an initialization method, a preprocessing method, a post-processing method, and a deinitialization method of the to-be-deployed model; compiling the registered target interface file into a model dynamic library; packaging the to-be-deployed file into a compressed file; and sending the model dynamic library and the compressed file to a target device end to complete the deployment of the to-be-deployed model.

[0006] During the implementation of the above solution, by registering the target interface instance for the model to be deployed, the preset unified interface is bound to the interface in the target interface file, so that the application on the target device side can call the interface methods of different models through the unified interface, and no longer need to pay attention to the implementation details of the model. On the one hand, it is conducive to the standardization and unification of the interface, avoiding compatibility problems caused by differences in model interfaces, allowing applications to more flexibly integrate and call different models, and facilitating model management and expansion; on the other hand, the model deployment side can more clearly understand the interface information and functions of each model through the registration mechanism. On the target device side, the application can interact with the model through the unified interface, which simplifies the model calling logic, reduces the complexity of model management, and is conducive to improving development and maintenance efficiency.

[0007] In an implementation of the first aspect, the method further includes: creating a private data source file of the model to be deployed; wherein the private data source file is used to define the data structure of the private data of the model to be deployed and declare functions for accessing and operating the private data; compiling the private data source file into a private data dynamic library; wherein the private data dynamic library stores a private data structure; the application deployed on the target device end accesses and / or operates the private data structure by calling the preset unified interface; and sending the private data dynamic library to the target device end.

[0008] During the implementation of the above solution, the model's private data is encapsulated in a private data dynamic library. The application accesses and operates the private data through a preset unified interface, avoiding the risk of private data exposure, thereby reducing the risk of data leakage and tampering, and helping to improve the security of model deployment. On the other hand, the model's private data is encapsulated in a private data dynamic library, which reduces the coupling between the application and the model, and helps to improve the flexibility of model deployment.

[0009] In an implementation of the first aspect, the method further includes: compiling the preset unified interface file of the preset unified interface into an interface dynamic library; sending the interface dynamic library to the target device end, so that the target device end uses the interface dynamic library to call the interface method.

[0010] During the implementation of the above solution, the preset unified interface is separately encapsulated, which reduces the coupling between the model and the application, and is conducive to improving the flexibility of the above model deployment method; on the other hand, when the model needs to be replaced or updated, the new model only needs to follow the same preset unified interface definition, and the application can interact with the new model without modification, which reduces the cost and risk of model replacement and is conducive to further improving the flexibility of the above model deployment method.

[0011] In an implementation of the first aspect, packaging the file to be deployed into a compressed file includes: splitting the model structure file in the file to be deployed into multiple model structure slices; and packaging the multiple model structure slices into compressed files respectively.

[0012] In the implementation process of the above solution, by splitting the model structure file into multiple model structure slices, small files can be stored more flexibly in the limited storage space of the target device. The target device can also classify and store the model slices according to their types, functions, etc., which is conducive to improving the efficiency and flexibility of storage management. On the other hand, multiple model structure slices can be transmitted in parallel. If there is an interruption during the transmission process, only the incomplete or damaged model structure slices need to be retransmitted without retransmitting the entire model structure file, which is conducive to saving time and network resources. On the other hand, during the model inference process, the target device can gradually load the corresponding model slices according to actual needs to reduce memory usage and loading time.

[0013] In a second aspect, an embodiment of the present application provides a model deployment method for a configurable model, which is applied to a target device end. The method includes: receiving and storing a model dynamic library and a compressed file of a to-be-deployed model sent by a model deployment end; responding to a target model inference request of an application program, executing a target model inference step to complete the deployment of the to-be-deployed model;

[0014] Among them, the target model reasoning step includes: decompressing the compressed file and loading the model to be deployed into the memory; using a preset unified interface to call the interface method in the model dynamic library to reason about the model to be deployed; wherein, the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

[0015] During the implementation of the above solution, the application deployed on the target device can call different interface methods by calling the preset unified interface, which is conducive to the standardization and unification of the interface, avoiding compatibility issues caused by differences in model interfaces, and allowing the application to more flexibly integrate and call different models, facilitating model management and expansion.

[0016] In an implementation of the second aspect, the method further includes:

[0017] Receive and store the private data dynamic library sent by the model deployment end; wherein the private data dynamic library stores a private data structure; the private data dynamic library is compiled from the private data source file of the model to be deployed; the private data source file is used to define the data structure of the private data of the model to be deployed and declare functions for accessing and operating the private data; in response to the private data access request of the application, call the preset unified interface to access the private data structure in the private data dynamic library; in response to the private data operation request of the application, call the preset unified interface to operate the private data structure in the private data dynamic library; wherein the private data access request and the private data operation request are requests sent when the application calls the interface method.

[0018] During the implementation of the above solution, the private data structure supports applications to access and operate through a preset unified interface, so that external interfaces cannot directly access and operate private data, reducing the risk of data leakage and tampering, and helping to improve the security of private data.

[0019] In an implementation of the second aspect, storing the model dynamic library and the compressed file includes: classifying and storing the model dynamic library and the compressed file based on the model name of the model to be deployed.

[0020] In the implementation process of the above solution, the model name is used as the unique identifier. Classified storage according to the model name can quickly find the dynamic library and compressed files of a specific model, reducing the time and complexity of finding the target model file in a large number of files, which is conducive to improving file management efficiency; on the other hand, the dynamic library and compressed files of each model are stored in a classified manner. When replacing or updating the corresponding model, only the corresponding files need to be replaced, which is conducive to improving the update and maintenance efficiency of the model; on the other hand, when a new model needs to be added, it is only necessary to create a new directory according to the naming convention and store the corresponding files. There is no need to modify the storage structure in the target device end, which is conducive to improving the scalability of the above model deployment method.

[0021] In an implementation of the second aspect, before decompressing the compressed file and loading the model to be deployed into the memory, the method further includes: obtaining the model name of the model to be deployed based on the target model inference request; using a preset environment variable and the model name, obtaining the loading path of the model to be deployed; wherein the preset environment variable points to the storage path of the compressed file of the model to be deployed; and obtaining the compressed file of the model to be deployed based on the loading path.

[0022] During the implementation of the above solution, the application can flexibly find and load the model compressed file without hard-coding the storage path of the model file in the code. This allows the application to find the model compressed file in the corresponding directory by simply changing the value of the environment variable on different devices or systems, which is conducive to improving the flexibility of the above model deployment method.

[0023] In an implementation of the second aspect, decompressing the compressed file and loading the model to be deployed into the memory includes: dynamically decompressing the compressed file and dynamically loading the model to be deployed into the memory.

[0024] During the implementation of the above solution, the compressed file of the model is loaded into the memory only when needed, which reduces the memory usage when the application starts, and is beneficial to saving memory and storage space on the target device. On the other hand, the new model can be added to the target device quickly and easily by simply storing the model compressed file and dynamic library file in the specified directory. The application can dynamically load and use the new model at runtime without recompiling and deploying the entire application, which is beneficial to improving the flexibility and scalability of the above model deployment method. On the other hand, the model deployment is independent of the application. When updating the model, only the corresponding compressed file and dynamic library file need to be replaced without modifying the application code, which is beneficial to simplifying the model update and maintenance process.

[0025] In a third aspect, an embodiment of the present application further provides a model deployment method for a configurable model, comprising:

[0026] The model deployment end performs the following steps: obtaining a to-be-deployed file and a target interface file of a to-be-deployed model; wherein the to-be-deployed model is a configurable model; registering a target interface instance of the to-be-deployed model to bind a preset unified interface to an interface mode in the target interface file; wherein the interface mode includes an initialization mode, a pre-processing mode, a post-processing mode, and a de-initialization mode of the to-be-deployed model; compiling the registered target interface file into a model dynamic library; packaging the to-be-deployed file into a compressed file; and sending the model dynamic library and the compressed file to a target device end to complete the deployment of the to-be-deployed model.

[0027] The target device side performs the following steps: receiving and storing the model dynamic library and the compressed file of the model to be deployed sent by the model deployment side; executing the target model inference step in response to the target model inference request of the application to complete the deployment of the model to be deployed; wherein, the target model inference step includes: decompressing the compressed file and loading the model to be deployed into the memory; using the preset unified interface to call the interface method in the model dynamic library to infer the model to be deployed.

[0028] In a fourth aspect, an embodiment of the present application provides a model deployment terminal, including:

[0029] A file acquisition module, configured to acquire a to-be-deployed file and a target interface file of a to-be-deployed model; wherein the to-be-deployed model is a configurable model;

[0030] An interface instance registration module is used to register the target interface instance of the model to be deployed, so as to bind the preset unified interface with the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed;

[0031] A model dynamic library compilation module is used to compile the registered target interface file into a model dynamic library;

[0032] A compression module, used to package the files to be deployed into a compressed file;

[0033] The sending module is used to send the model dynamic library and the compressed file to the target device end to complete the deployment of the model to be deployed.

[0034] In a fifth aspect, an embodiment of the present application provides a target device, including:

[0035] The receiving module is used to receive and store the model dynamic library and compressed file of the model to be deployed sent by the model deployment end;

[0036] A model reasoning module, configured to respond to a target model reasoning request from an application program and execute a target model reasoning step to complete the deployment of the to-be-deployed model;

[0037] Among them, the target model reasoning step includes: decompressing the compressed file and loading the model to be deployed into the memory; using a preset unified interface to call the interface method in the model dynamic library to reason about the model to be deployed; wherein, the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

[0038] In the sixth aspect, an embodiment of the present application provides a model deployment system, which includes a model deployment end and a target device end, wherein the model deployment end is the model deployment end provided by the third aspect or any possible implementation method of the third aspect, and the target device end is the fourth aspect or any possible target device end of the fourth aspect.

[0039] In the seventh aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a memory and a communication bus, wherein the processor and the memory communicate with each other through the communication bus; the memory stores computer program instructions that can be executed by the processor, and when the computer program instructions are read and run by the processor, the method provided by the first aspect or any possible implementation of the first aspect or the second aspect or any possible implementation of the second aspect or the third aspect or any possible implementation of the third aspect is executed.

[0040] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are read and run by a processor, the method provided by the first aspect or any possible implementation of the first aspect or the second aspect or any possible implementation of the second aspect or the third aspect or any possible implementation of the third aspect is executed.

[0041] In a ninth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method provided by the first aspect or any possible implementation of the first aspect, or the second aspect or any possible implementation of the second aspect, or the third aspect or any possible implementation of the third aspect.

[0042] Other features and advantages of the present application will be described in the following description and, in part, will become apparent from the description or be understood by practicing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 A flow chart of a model deployment method for a configurable model applied to a model deployment terminal provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of the configuration parameters of a configuration file in an application scenario provided in an embodiment of the present application;

[0046] Figure 3A flow chart of a model deployment method for a configurable model applied to a target device provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram illustrating the relationship between the private data source file, private data dynamic library, private data structure, application, preset unified interface, and model dynamic library provided in an embodiment of the present application;

[0048] Figure 5 A flow chart of a model deployment method for a configurable model provided in an embodiment of the present application;

[0049] Figure 6 A schematic diagram of the structure of the model deployment terminal provided in an embodiment of the present application;

[0050] Figure 7 A schematic diagram of the structure of the target device provided in an embodiment of the present application;

[0051] Figure 8 A schematic diagram of the structure of the model deployment system provided in an embodiment of the present application;

[0052] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application and are therefore only examples and cannot be used to limit the scope of protection of the present application.

[0054] The embodiment of the present application provides a model deployment method for a configurable model. The method registers a target interface instance for the model to be deployed, thereby binding a preset unified interface to the interface in the target interface file, so that the application on the target device side can call the interface methods of different models through a unified interface, and no longer needs to pay attention to the implementation details of the model. On the one hand, it is conducive to the standardization and unification of the interface, avoiding compatibility issues caused by differences in model interfaces, allowing applications to more flexibly integrate and call different models, facilitating model management and expansion; on the other hand, the model deployment end can more clearly understand the interface information and functions of each model through the registration mechanism. On the target device side, the application can interact with the model through a unified interface, simplifying the model calling logic, reducing the complexity of model management, and helping to improve development and maintenance efficiency. The above-mentioned model deployment method involves a model deployment end 100 and a target device end 200, wherein the model deployment end 100 refers to the environment or platform for developing and preparing the model so that the model can run on the target device, which can be a workstation, server or computing cluster. The main functions of the model deployment end 100 include: model development and training, model optimization, model file preparation, testing and verification, etc. The target device 200 is the device or system that actually runs and uses the deployed model, such as an embedded device. The main functions of the target device 200 include running applications, receiving model files, loading models, and performing model inference.

[0055] The following describes the model deployment methods applied to the model deployment end 100 and the target device end 200:

[0056] See Figure 1 The embodiment of the present application provides a model deployment method for a configurable model applied to a model deployment terminal 100, including:

[0057] Step S110: Obtain the to-be-deployed file and target interface file of the to-be-deployed model, wherein the to-be-deployed model is a configurable model.

[0058] The configurable models mentioned above are models that can be flexibly configured, deployed, and used. When the target device 200 is an embedded device, the models to be deployed may be computer vision models and computer auditory models. Computer vision models include image classification models, object detection models, and text detection models. Computer auditory models include speech recognition models and speech detection models.

[0059] The above-mentioned files to be deployed may include model structure files and configuration files, wherein the model structure file refers to a file containing the model architecture, which describes the model's hierarchical structure, the type of each layer, the connection method, and other information. Generally speaking, the model structure file in the above-mentioned files to be deployed refers to the structure file of the artificial intelligence model that has completed training and optimization (or, completed training, optimization, and quantization). The main functions of the model structure file include: (1) Model reconstruction: Rebuilding the structure of the model on the target device end 200 for reasoning; (2) Model conversion: Converting models between different deep learning frameworks through the model structure file. The model structure file can be a structure file derived from the training framework, or a structure file obtained after conversion using a model conversion tool.

[0060] The configuration file mentioned above refers to a file containing model configuration parameters. It defines various parameters and settings for the model's runtime, such as input and output dimensions, data types, pre-processing flags, and post-processing flags. The configuration file provides the necessary parameter settings for the model's operation, ensuring that the model correctly processes input data and generates output results. By modifying the configuration file, the model's behavior and performance can be adjusted without changing the model's structure or code. Configuration files can be manually written by the model developer or automatically generated by tools or scripts.

[0061] The above-mentioned model structure file can adopt the nb format, and the configuration file can adopt the json (JavaScript Object Notation) format. JavaScript Object Notation is designed based on a subset of ECMAScript and is an open standard file format and data exchange format. Both the model structure file and the configuration file can be named as the model name. The configuration file includes the configuration of input parameters and output parameters, where the input parameters include the number, width, height, number of channels, arrangement, format, pre-processing flags, etc. of the input tensors (tensors). The output parameters include the number, format, post-processing flags, etc. of the output tensors (tensors). For example, when deploying a face detection model, the model name is named roi (Region of Interest), then the model structure file can be named model_roi.nb, and the configuration file can be named model_roi.json. In a certain application scenario, the parameter configuration contained in the configuration file model_roi.json is as follows Figure 2 As shown, in this application scenario, the model configuration parameters include:

[0062] (1) Model name: roi.

[0063] (2) Preprocessing parameters, including:

[0064] Number of input data: 1;

[0065] The 0th data attribute includes:

[0066] Preprocessing flag: off; data width: 1920; data height: 1080; data channels: 3; data format: RGB.

[0067] (3) Post-processing parameters, including:

[0068] Number of output data: 1;

[0069] The 0th data attribute includes:

[0070] Post-processing flag: on; data format: rgb.

[0071] The target interface file is an interface file defined and implemented by the model deployment end 100 for a specific model. It is typically a source code file that implements the model's interface methods, including initialization, preprocessing, postprocessing, and deinitialization methods. These methods provide the necessary functionality and logic for the model to run on the target device end 200. The target interface file can be manually written by the model developer or generated using a predefined template or script.

[0072] Step S120: registering the target interface instance of the model to be deployed to bind the preset unified interface with the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed.

[0073] It should be pointed out that the interface mode in the above-mentioned target interface file actually refers to the interface method in the target interface file, while the initialization mode, preprocessing mode, post-processing mode and deinitialization method refer to the above-mentioned initialization method, preprocessing method, post-processing method and deinitialization method respectively.

[0074] The above-mentioned target interface instance provides a unified calling interface for each interface method, and matches the callback function of each model through different function names. The registration process of the target interface instance can be carried out in the model configuration framework deployed on the model deployment end. The model configuration framework is mainly used to register, configure and manage the target interface instance on the model deployment end 100. The above-mentioned preset unified interface refers to a set of interface specifications pre-defined in the model configuration framework, including the name of the interface, parameter list and return type, etc. These interface specifications are unified, that is, all models need to follow this specification to implement their own interface methods. The registration process of the above-mentioned target interface instance refers to the process of binding (or associating) the specific interface method in the target interface file with the preset unified interface. After registration, when the application deployed on the target device end 200 calls the preset unified interface, it actually executes the corresponding specific method in the target interface file. For example, the above-mentioned registration process:

[0075] The model class is registered in the OSML (Open Source Machine Learning) unified interface. The OSML unified interface defines the initialization method interface as int osml_init(void*handle), the deinitialization method interface as int osml_deinit(void*handle), the preprocessing method interface as int osml_pre_process(void*handle), and the postprocessing method interface as int osml_post_process(void*handle). The application deployed on the target device end 200 can initialize the model by calling the initialization method interface osml_init and allocate relevant resources for model operation. When the preprocessing flag of the input data is on, the input data of the model is preprocessed by calling the preprocessing method interface osml_pre_process. If the preprocessing flag of the input data is off, it means that the model does not need to perform preparatory work before formal processing, and the preprocessing method interface returns empty. The output data of the model is post-processed by calling the postprocessing method interface osml_post_process. Deinitialization is performed by calling the deinitialization method interface osml_deinit.

[0076] By registering in the model configuration framework, different models have the same calling method, which facilitates unified management and expansion. On the other hand, centralized management of model interfaces and parameters in the model configuration framework is conducive to the standardization and automation of model configuration.

[0077] Step S130: compile the registered target interface file into a model dynamic library;

[0078] The above-mentioned dynamic library (Dynamic Library) is a library file that is loaded into memory only when the program is running. Unlike static libraries, the code of a dynamic library is shared between multiple programs, rather than being copied into the executable file when each program is compiled. This sharing mechanism can save disk space and memory resources. The application in the target device end 200 can use the dlopen function to open the dynamic library and execute the relevant initialization methods, preprocessing methods, post-processing methods, and deinitialization methods based on the handle returned by the dynamic library. If the handle is empty, no execution is required.

[0079] When deploying the face detection model ROI, for example, step S130 compiles the registered model_library_roi.cpp source file into a model dynamic library file, model_roi.so. This dynamic library file is loaded at runtime, saving application size. The model_library_roi.cpp source file includes the post-registration initialization method, the pre-processing method (osml_pre_process), the post-processing method (osml_post_process), and the de-initialization method (osml_deini).

[0080] Step S140: Pack the files to be deployed into a compressed file;

[0081] The compression format of the above compressed files can be zip, tar, etc. The purpose of packaging them into compressed files is to reduce the size of each model file to reduce the transmission time of the files to be deployed.

[0082] Optionally, the above step S140 includes: splitting the model structure file in the file to be deployed into multiple model structure slices; and packaging the multiple model structure slices into compressed files respectively.

[0083] Exemplarily, the splitting method of splitting the model structure file into multiple model structure slices may be any of the following:

[0084] (1) Split by fixed size: Split the model structure file evenly into multiple small files according to a preset fixed size (such as 1MB, 2MB, etc.);

[0085] (2) Split according to the model structure: Split according to the model structure (such as different layers or modules in the model), for example, split the convolutional layer and the fully connected layer into independent files;

[0086] (3) Hybrid splitting: Combine fixed size and model structure hierarchy to split. For example, first perform a preliminary split according to the model structure, and then further split each part according to a fixed size.

[0087] When the target device 200 performs model inference, it can decompress the model structure files in sequence according to the order of the model files (such as index order or time order) to load them into the memory, and reorganize them in sequence to restore the structure of the original model file for inference.

[0088] The above solution splits the model structure file into multiple model structure slices, and small files can be stored more flexibly in the limited storage space of the target device end 200. The target device end 200 can also classify and store the model slices according to their types, functions, etc., which is beneficial to improving the efficiency and flexibility of storage management; on the other hand, multiple model structure slices can be transmitted in parallel. If an interruption occurs during the transmission process, only the incomplete or damaged model structure slices need to be retransmitted, and there is no need to retransmit the entire model structure file, which is beneficial to saving time and network resources; on the other hand, during the model inference process, the target device end 200 can gradually load the corresponding model slices according to actual needs, reducing memory usage and loading time.

[0089] Step S150: Send the model dynamic library and the compressed file to the target device 200 to complete the deployment of the model to be deployed.

[0090] The above-mentioned model dynamic library and compressed files can be transmitted to the target device end 200 through network transmission or physical storage medium transmission. Among them, the network transmission method, for example, uses the file transfer protocol FTP, the secure copy protocol SCP, the hypertext transfer protocol HTTP, the hypertext transfer security protocol HTTPS, and other network transmission protocols to transmit the model dynamic library and compressed files to the target device end 200. The physical storage medium transmission method, for example, uses a physical storage medium such as an SD card or a USB flash drive to transmit the model dynamic library and compressed files to the target device end 200.

[0091] Generally, the target device end 200 also needs some private parameters and data of the model to run the above-mentioned model to be deployed. The private parameters and data of the model generally refer to the parameters and data used inside the model. These parameters and data are usually not exposed to the outside and can only be accessed and modified through specific interfaces or methods. Private parameters and data generally include: configuration parameters, internal states and specific data of the model. The purpose of setting private parameters and data is to protect the internal implementation details of the model and prevent direct external access and modification, thereby ensuring the stability and security of the model. However, the model deployment scheme in the related art generally directly accesses and modifies the relevant private data, which makes the security of the model poor. Based on this, the embodiment of the present application provides the following scheme:

[0092] Optionally, the model deployment method applied to the model deployment terminal 100 further includes: creating a private data source file for the model to be deployed; wherein the private data source file is used to define the data structure of the private data of the model to be deployed and declare functions for accessing and operating the private data; compiling the private data source file into a private data dynamic library; wherein the private data dynamic library stores private data structures; an application deployed on the target device terminal 200 accesses and / or operates the private data structures by calling a preset unified interface; and sending the private data dynamic library to the target device terminal 200. This embodiment is, for example:

[0093] First, create a private data source file osml_private_if.h for the model to be deployed. The functions of this source file include:

[0094] (1) Define the data structure of private data: The private data source file osml_private_if.h can define the private data structure of the model, which is used to store specific parameters and information required by the model at runtime;

[0095] (2) Declare related functions: The private data source file osml_private_if.h can declare functions used to access and operate private data, such as the function for obtaining the output tensor size, the function for obtaining the tensor arrangement, etc.

[0096] Then, compile the private data source file osml_private_if.h into a private data dynamic library, which stores the private data structure osml_private. It is understood that the private parameters and data required by different models need to be specified in the model initialization method and the corresponding fields added to the private data structure osml_private for unified management.

[0097] Finally, the private data dynamic library is sent to the target device 200. The method of sending the private data dynamic library to the target device 200 is similar to the method of sending the model dynamic library and the compressed file to the target device 200, which will not be described in detail in this embodiment.

[0098] In addition, the application can only access and / or operate the private data structure through the interface corresponding to the initialization method in the preset unified interface. For example, the application fills the private data structure by calling the initialization method.

[0099] The private data of the model in the above solution is encapsulated in a private data dynamic library. The application accesses and operates the private data through a preset unified interface, avoiding the risk of private data exposure, thereby reducing the risk of data leakage and tampering, and helping to improve the security of model deployment; on the other hand, the private data of the model is encapsulated in a private data dynamic library, which reduces the coupling between the application and the model, and helps to improve the flexibility of model deployment.

[0100] Optionally, the model deployment method applied to the model deployment terminal 100 further includes: compiling the preset unified interface file of the preset unified interface into an interface dynamic library; sending the interface dynamic library to the target device terminal 200, so that the target device terminal 200 uses the interface dynamic library to call the interface method. This implementation method is, for example:

[0101] The implementation of the above-mentioned OSML interface is packaged in a separate source file, and then the OSML interface source file is compiled into an interface dynamic library, and the interface dynamic library is sent to the target device end 200, so that the target device end 200 can use the preset unified interface in the interface dynamic library to call the interface method.

[0102] The above solution encapsulates the preset unified interface separately, reducing the coupling between the model and the application, which is conducive to improving the flexibility of the above model deployment method; on the other hand, when the model needs to be replaced or updated, the new model only needs to follow the same preset unified interface definition, and the application can interact with the new model without modification, reducing the cost and risk of model replacement, which is conducive to further improving the flexibility of the above model deployment method.

[0103] Based on the same inventive concept, an embodiment of the present application also provides a model deployment method for a configurable model applied to a target device end 200. The application deployed on the target device end 200 can call different interface methods by calling a preset unified interface, which is conducive to the standardization and unification of the interface, avoiding compatibility issues caused by differences in model interfaces, allowing the application to more flexibly integrate and call different models, and facilitating the management and expansion of the model.

[0104] See Figure 3 The model deployment method applied to the target device 200 includes:

[0105] Step S210: receiving and storing the model dynamic library and compressed file of the model to be deployed sent by the model deployment end 100, wherein the model to be deployed is a configurable model.

[0106] Optionally, the above step S210 stores the model dynamic library and compressed files, including: based on the model name of the model to be deployed, the model dynamic library and compressed files are classified and stored. For example, in this embodiment, each model corresponds to an independent model structure file (nb) and configuration file (json), and the model names are different from each other. Therefore, the model structure file and configuration file can be compressed into a compressed file named after the model name, and the compression format is such as tar, zip, and classified and stored according to the model name. The database and compressed file of each model can be stored in the same directory of the target device end 200 so that the model can be loaded during inference. Assuming that two models a and b are to be deployed, the models correspond to model structure files a.nb and b.nb and configuration files a.json and b.json. Classified storage according to model name means that the compressed file a.tar compressed with a.nb and a.json is stored on the target device end 200, and the compressed file b.tar compressed with b.nb and b.json is stored on the target device end 200.

[0107] The target device 200's file system can create a dedicated directory (e.g., / models) and store all model compressed files in this directory. When performing model inference, the application can search for the corresponding compressed file in this directory by the model name, then decompress and load the corresponding model.

[0108] The above scheme uses the model name as a unique identifier, and the classified storage according to the model name can quickly find the dynamic library and compressed files of a specific model, reducing the time and complexity of finding the target model file in a large number of files, which is conducive to improving file management efficiency; on the other hand, the dynamic library and compressed files of each model are stored in a classified manner. When replacing or updating the corresponding model, only the corresponding files need to be replaced, which is conducive to improving the update and maintenance efficiency of the model; on the other hand, when a new model needs to be added, it is only necessary to create a new directory and store the corresponding files according to the naming convention, without modifying the storage structure in the target device end 200, which is conducive to improving the scalability of the above model deployment method.

[0109] Step S220: In response to the target model inference request of the application, the target model inference step is executed to complete the deployment of the model to be deployed.

[0110] The above target model reasoning steps include:

[0111] Decompress the compressed file and load the model to be deployed into memory; use the preset unified interface to call the interface method in the model dynamic library to infer the model to be deployed; the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

[0112] The model deployment method applied to the target device 200 is independent of the application. According to the needs of the model, the relevant attributes of the model can be obtained and modified through the relevant interface. The application deployed on the target device 200 can start the model inference of the corresponding model by the model name.

[0113] Optionally, before decompressing the compressed file and loading the model to be deployed into memory, the model deployment method applied to the target device 200 further includes: obtaining the model name of the model to be deployed based on the target model inference request; obtaining the loading path of the model to be deployed using a preset environment variable and the model name; wherein the preset environment variable points to the storage path of the compressed file of the model to be deployed; and obtaining the compressed file of the model to be deployed based on the loading path. This embodiment is, for example:

[0114] The compressed files of all models are stored in a separate directory, which is separate from the mirror directory in the target device 200. By copying the path information of the separate directory to the mirror directory, the search directory for the model is specified by concatenating the environment variable with the model name. The environment variable points to the separate directory where the model compressed files are stored. For example, on the target device 200, setting the environment variable to / image / models means that the compressed files of the model are stored in the / image / models directory. When the application needs to load the model model_roi, the value of the environment variable / image / models is concatenated with the name of the model to be loaded, model_roi, to form the complete model compressed file path. The concatenated model compressed file path is / image / models / model_roi.zip.

[0115] The application in the above solution can flexibly find and load the model compressed file without hard-coding the storage path of the model file in the code. This allows the application to find the model compressed file in the corresponding directory by simply changing the value of the environment variable on different devices or systems, which helps to improve the flexibility of the above model deployment method.

[0116] Optionally, the decompressing the compressed file and loading the model to be deployed into the memory includes: dynamically decompressing the compressed file and dynamically loading the model to be deployed into the memory.

[0117] The above dynamic loading means that the model file is dynamically loaded into the memory as needed when the application is running, rather than statically linking or loading the model file into the application when the program is compiled or started.

[0118] During the implementation of the above solution, the compressed file of the model is loaded into the memory only when needed, which reduces the memory usage when the application is started, and is beneficial to saving the memory and storage space of the target device end 200; on the other hand, the new model can be added to the target device end 200 conveniently and quickly, and only the model compressed file and dynamic library file need to be stored in the specified directory. The application can dynamically load and use the new model at runtime without recompiling and deploying the entire application, which is beneficial to improving the flexibility and scalability of the above model deployment method; on the other hand, the model deployment is independent of the application. When updating the model, only the corresponding compressed file and dynamic library file need to be replaced without modifying the application code, which is beneficial to simplifying the model update and maintenance process.

[0119] The following describes the solution for the target device 200 to access and operate private data:

[0120] Optionally, the model deployment method applied to the target device 200 further includes: receiving and storing a private data dynamic library sent by the model deployment end 100; wherein the private data dynamic library stores a private data structure; the private data dynamic library is compiled from a private data source file of the model to be deployed; the private data source file is used to define the data structure of the private data of the model to be deployed and declare functions for accessing and operating the private data;

[0121] In response to a private data access request from an application, calling a preset unified interface to access a private data structure in a private data dynamic library;

[0122] In response to a private data operation request from an application, calling a preset unified interface to operate a private data structure in a private data dynamic library;

[0123] Among them, the private data access request and the private data operation request are requests sent when the application calls the interface method. This implementation method is for example:

[0124] If the deployed deep learning model is a face detection model, which belongs to a computer vision model and is named roi (Region of Interest), a new private data source file osml_private_if.h needs to be created on the model deployment end 100. This source file defines the private data structure of the above-mentioned roi model and declares functions for accessing and manipulating private data. The model deployment end 100 compiles the private data source file into a private data dynamic library and sends it to the target device end 200. In response to the application's private data access request and data operation request, the target device end 200 calls a preset unified interface to access and operate the private data structures in the private data dynamic library. For example: in the post-processing method (osml_post_process) of the above-mentioned face detection model roi, the roi model needs to perform cleanup work after formal processing, so the output tensor can be added to the private data structure osml_private. In the post-processing process, the size and arrangement of the tensor must be obtained. Therefore, the two functions OSML_GetOutTensorSize(void*handle) and OSML_GetLayout(void*handle) can be declared in the header file of the preset unified interface and implemented in the private data dynamic library file. In addition, the binary model file reading interface uint8_t*get_binary_file(size_t*size) and the model configuration parameter reading interface uint8_t*get_json_file(size_t*size) can also be declared in the preset unified interface for the model dynamic library to call.

[0125] The following describes the relationship between private data source files, private data dynamic libraries, private data structures, applications, preset unified interfaces, and model dynamic libraries:

[0126] See Figure 4 The private data dynamic library is compiled from private data source files and stores private data structures. Applications call interface methods in the model dynamic library by calling the method interface in the preset unified interface. When accessing and manipulating private data structures, the interface methods in the model dynamic library call the private data structure operation and access functions in the preset unified interface to operate and access the private data structures stored in the private data dynamic library.

[0127] In addition, the private data structure in the private data dynamic library can only be accessed and operated when the application calls the interface method, thereby improving the security of the private data structure.

[0128] The private data structure in the above solution supports application access and operation through a preset unified interface, making it impossible for external interfaces to directly access and operate private data, reducing the risk of data leakage and tampering, and helping to improve the security of private data.

[0129] See Figure 5 Based on the same inventive concept, the embodiment of the present application further provides a model deployment method for a configurable model, including:

[0130] The model deployment terminal 100 performs the following steps:

[0131] Step S110: Obtain the to-be-deployed file and target interface file of the to-be-deployed model; wherein the to-be-deployed model is a configurable model;

[0132] Step S120: registering the target interface instance of the model to be deployed, so as to bind the preset unified interface to the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed;

[0133] Step S130: compile the registered target interface file into a model dynamic library;

[0134] Step S140: Pack the files to be deployed into a compressed file;

[0135] Step S150: Send the model dynamic library and the compressed file to the target device 200 to complete the deployment of the model to be deployed;

[0136] The target device 200 performs the following steps:

[0137] Step S210: receiving and storing the model dynamic library and compressed file of the model to be deployed sent by the model deployment terminal 100;

[0138] Step S220: In response to the target model inference request of the application, executing the target model inference step to complete the deployment of the to-be-deployed model;

[0139] The above-mentioned target model inference steps include: decompressing the compressed file and loading the model to be deployed into the memory; using the preset unified interface to call the interface method in the model dynamic library to infer the model to be deployed; wherein the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

[0140] See Figure 6 Based on the same inventive concept, the embodiment of the present application further provides a model deployment terminal 100, including:

[0141] The file acquisition module 101 is used to acquire the to-be-deployed files and target interface files of the to-be-deployed model;

[0142] The interface instance registration module 102 is used to register the target interface instance of the model to be deployed, so as to bind the preset unified interface with the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed;

[0143] The model dynamic library compiling module 103 is used to compile the registered target interface file into a model dynamic library;

[0144] Compression module 104, used to package the files to be deployed into compressed files;

[0145] The sending module 105 is used to send the model dynamic library and the compressed file to the target device to complete the deployment of the model to be deployed.

[0146] See Figure 7 Based on the same inventive concept, the embodiment of the present application further provides a target device 200, including:

[0147] The receiving module 201 is used to receive and store the model dynamic library and compressed file of the model to be deployed sent by the model deployment end;

[0148] The model reasoning module 202 is used to respond to the target model reasoning request of the application program and execute the target model reasoning step to complete the deployment of the to-be-deployed model;

[0149] Among them, the target model reasoning step includes: decompressing the compressed file and loading the model to be deployed into the memory; using the preset unified interface to call the interface method in the model dynamic library to reason about the model to be deployed; among them, the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

[0150] See Figure 8 Based on the same inventive concept, an embodiment of the present application also provides a model deployment system 300 , which includes a model deployment end 100 and a target device end 200 .

[0151] Figure 9 This is a schematic diagram of an electronic device provided in an embodiment of the present application. Figure 9 The electronic device 400 includes a processor 410, a memory 420, and a communication interface 430. These components are interconnected and communicate with each other via a communication bus 440 and / or other forms of connection mechanisms (not shown).

[0152] The memory 420 includes one or more (only one is shown in the figure), which may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The processor 410 and other possible components can access the memory 420 and read and / or write data therein.

[0153] The processor 410 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 410 can be a GPU (Graphics Processing Unit), a CPU (Central Processing Unit), an AI (Artificial Intelligence), an NPU (Neural Network Processing Unit), an ISP (Image Signal Processor), a DPU (Display Processing Unit), a VPU (Video Processing Unit), a data processing core of a DSP (Digital Signal Processor), etc., or it can be a processor chip used in scenarios such as some large-scale data operations. The above is only an example and should not be a limitation to this application.

[0154] Communication interface 430 includes one or more (only one is shown in the figure) interfaces that can be used to communicate directly or indirectly with other devices to exchange data. For example, communication interface 430 can be an Ethernet interface; a mobile communication network interface, such as a 3G, 4G, or 5G network interface; or other types of interfaces that have data transmission and reception capabilities.

[0155] One or more computer program instructions can be stored in the memory 420, and the processor 410 can read and execute these computer program instructions to implement the model deployment method and other desired functions applied to the model deployment end 100 or the model deployment method and other desired functions applied to the target device end 200 provided in the embodiment of the present application.

[0156] I understand. Figure 9 The structure shown is for illustration only. The electronic device 400 may also include Figure 9 More or fewer components than shown, or with Figure 9 Different configurations shown. Figure 9 Each component shown in the figure can be implemented using hardware, software, or a combination thereof. For example, the electronic device 400 can be a single server (or other device with computing processing capabilities), a combination of multiple servers, a cluster of a large number of servers, etc., and can be both a physical device and a virtual device.

[0157] The present application also provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are read and executed by a computer processor, the model deployment method and other desired functions applied to the model deployment terminal 100 or the model deployment method and other desired functions applied to the target device terminal 200 provided in the present application are executed. For example, the computer-readable storage medium can be implemented as Figure 9 The memory 420 in the electronic device 400.

[0158] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0159] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] Based on the same inventive concept, an embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned model deployment method applied to the model deployment end 100 or the model deployment method applied to the target device end 200.

[0161] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Those skilled in the art will appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A model deployment method for a configurable model, characterized in that: Applied to the model deployment end, the method includes: Obtaining a to-be-deployed file and a target interface file of a to-be-deployed model; wherein the to-be-deployed model is a configurable model; Registering the target interface instance of the model to be deployed to bind the preset unified interface to the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed; Compile the registered target interface file into a model dynamic library; Packing the files to be deployed into a compressed file; The model dynamic library and the compressed file are sent to the target device to complete the deployment of the model to be deployed.

2. The model deployment method of the configurable model according to claim 1, characterized in that: The method further comprises: Creating a private data source file for the model to be deployed; wherein the private data source file is used to define the data structure of the private data of the model to be deployed and to declare functions for accessing and operating the private data; Compiling the private data source file into a private data dynamic library; wherein the private data dynamic library stores a private data structure; and an application deployed on the target device accesses and / or operates the private data structure by calling the preset unified interface; The private data dynamic library is sent to the target device end.

3. The model deployment method of the configurable model according to claim 1, characterized in that: The method further comprises: Compiling the preset unified interface file of the preset unified interface into an interface dynamic library; The interface dynamic library is sent to the target device end, so that the target device end uses the interface dynamic library to call the interface method.

4. The model deployment method of a configurable model according to any one of claims 1 to 3, characterized in that: The step of packaging the files to be deployed into a compressed file includes: Splitting the model structure file in the file to be deployed into multiple model structure slices; The plurality of model structure slices are packaged into compressed files respectively.

5. A model deployment method for a configurable model, characterized in that: Applied to the target device, the method includes: Receive and store the model dynamic library and compressed file of the model to be deployed sent by the model deployment end; wherein the model to be deployed is a configurable model; In response to a target model inference request of an application program, executing a target model inference step to complete the deployment of the to-be-deployed model; The target model reasoning step includes: Decompress the compressed file and load the model to be deployed into memory; The model to be deployed is inferred by calling the interface method in the model dynamic library using a preset unified interface; wherein the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

6. The model deployment method of the configurable model according to claim 5, characterized in that: The method further comprises: Receive and store a private data dynamic library sent by the model deployment end; wherein the private data dynamic library stores a private data structure; the private data dynamic library is compiled from the private data source file of the model to be deployed; the private data source file is used to define the data structure of the private data of the model to be deployed and declare functions for accessing and operating the private data; In response to a private data access request from the application, calling the preset unified interface to access the private data structure in the private data dynamic library; In response to a private data operation request of the application, calling the preset unified interface to operate the private data structure in the private data dynamic library; The private data access request and the private data operation request are requests sent when the application calls the interface method.

7. The model deployment method of the configurable model according to claim 5, characterized in that: Storing the model dynamic library and the compressed file includes: Based on the model name of the model to be deployed, the model dynamic library and the compressed file are classified and stored.

8. The model deployment method of the configurable model according to claim 7, characterized in that: Before decompressing the compressed file and loading the to-be-deployed model into memory, the method further includes: Based on the target model inference request, obtaining the model name of the to-be-deployed model; Using a preset environment variable and the model name, obtain the loading path of the model to be deployed; wherein the preset environment variable points to the storage path of the compressed file of the model to be deployed; Based on the loading path, the compressed file of the model to be deployed is obtained.

9. The model deployment method of a configurable model according to any one of claims 5 to 8, characterized in that: The decompressing the compressed file and loading the model to be deployed into the memory includes: The compressed file is dynamically decompressed, and the model to be deployed is dynamically loaded into the memory.

10. A model deployment method for a configurable model, characterized in that: The method comprises: The model deployment side performs the following steps: Obtaining a to-be-deployed file and a target interface file of a to-be-deployed model; wherein the to-be-deployed model is a configurable model; Registering the target interface instance of the model to be deployed to bind the preset unified interface to the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed; Compile the registered target interface file into a model dynamic library; Packing the files to be deployed into a compressed file; Sending the model dynamic library and the compressed file to the target device to complete the deployment of the model to be deployed; The target device performs the following steps: Receive and store the model dynamic library and the compressed file of the model to be deployed sent by the model deployment end; In response to a target model inference request of an application program, executing a target model inference step to complete the deployment of the to-be-deployed model; Among them, the target model reasoning step includes: decompressing the compressed file and loading the model to be deployed into the memory; using the preset unified interface to call the interface method in the model dynamic library to reason about the model to be deployed.

11. A model deployment terminal, characterized in that: include: A file acquisition module, configured to acquire a to-be-deployed file and a target interface file of a to-be-deployed model; wherein the to-be-deployed model is a configurable model; An interface instance registration module is used to register the target interface instance of the model to be deployed, so as to bind the preset unified interface with the interface mode in the target interface file; wherein the interface mode includes the initialization mode, pre-processing mode, post-processing mode and de-initialization mode of the model to be deployed; A model dynamic library compilation module is used to compile the registered target interface file into a model dynamic library; A compression module, used to package the files to be deployed into a compressed file; The sending module is used to send the model dynamic library and the compressed file to the target device end to complete the deployment of the model to be deployed.

12. A target device, characterized in that: include: A receiving module, configured to receive and store the model dynamic library and compressed file of the model to be deployed sent by the model deployment end; wherein the model to be deployed is a configurable model; A model reasoning module, configured to respond to a target model reasoning request from an application program and execute a target model reasoning step to complete the deployment of the to-be-deployed model; Among them, the target model reasoning step includes: decompressing the compressed file and loading the model to be deployed into the memory; using a preset unified interface to call the interface method in the model dynamic library to reason about the model to be deployed; wherein, the interface method includes the initialization method, preprocessing method, post-processing method and deinitialization method of the model to be deployed.

13. A model deployment system, characterized in that: The system includes a model deployment end and a target device end, wherein the model deployment end is the model deployment end as described in claim 11, and the target device end is the target device end as described in claim 12.

14. An electronic device, characterized in that: include: A processor, a memory and a communication bus, wherein the processor and the memory communicate with each other via the communication bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1 to 10 by calling the program instructions.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when executed by a computer, enable the computer to perform the method according to any one of claims 1 to 10.