Ai model deployment method and apparatus, electronic device, medium and program product

By determining the conversion path and performance evaluation, the AI model is automatically deployed to the target device, solving the problem that AI model deployment requires multi-party collaboration, and achieving a fast and stable deployment process.

WO2025145435A1PCT designated stage expired Publication Date: 2025-07-10SIEMENS AG +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/070881
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

In the prior art, AI models cannot be deployed directly on field devices, and AI engineers and automation engineers need to work together, resulting in large time and resource overhead.

Method used

An AI model deployment method is provided, by receiving the AI model trained by the user, determining the format supported by the target inference device, determining the conversion path, and converting the AI model into a format supported by the target device, and finally performing performance evaluation on the target device to determine the optimal model.

Benefits of technology

It realizes rapid and automated AI model deployment, reduces manpower investment, and ensures the stable operation of target equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024070881_10072025_PF_FP_ABST
    Figure CN2024070881_10072025_PF_FP_ABST
Patent Text Reader

Abstract

An AI model deployment method, an electronic device, a medium and a computer program product, relating to the field of artificial intelligence. The AI model deployment method comprises: receiving an AI model in a first format that has been trained by a user; determining a target inference device; determining at least one second format of the AI model that can be supported by the target inference device; according to the first format and the second format, determining a plurality of conversion paths; by means of each conversion path among the plurality of conversion paths, converting the AI model in the first format, so as to obtain a plurality of AI models in the second format; and deploying the plurality of AI models in the second format to the target inference device to perform performance evaluation, so as to determine an optimal AI model.
Need to check novelty before this filing date? Find Prior Art

Description

AI model deployment method, device, electronic device, medium and program product Technical Field

[0001] The embodiments of the present application mainly relate to the field of artificial intelligence, and in particular to a deployment method, electronic device, medium, and computer program product of an artificial intelligence (AI) model. Background Art

[0002] With the continuous deepening of research and application of artificial intelligence in various fields of society, the application of artificial intelligence in the industrial field has become more and more extensive. Various field devices such as industrial computers, edge devices, and programmable logic controllers (PLCs) use artificial intelligence for reasoning during factory production operations.

[0003] Field devices typically have limited computing power and are often unable to directly run a variety of AI models. Engineers designing AI models often focus on the model itself and fail to fundamentally understand the target device's requirements. Automation engineers, while familiar with the target device, lack a deep understanding of AI. Consequently, deploying trained AI models to field devices still requires collaboration between AI and automation engineers, which undoubtedly consumes significant time and resources.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a method, device, electronic device, medium and program product for deploying an AI model. Through the embodiments of the present application, the AI ​​model can be conveniently and quickly deployed to the target inference device at the factory site.

[0006] In a first aspect, a method for deploying an AI model is provided, comprising: receiving an AI model in a first format that has been trained by a user; determining the target inference device; determining at least one second format of the AI ​​model that the target inference device can support; determining multiple conversion paths based on the first format and the second format; converting the AI ​​model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format; deploying the multiple AI models in the second format to the target inference device for performance evaluation to determine the optimal AI model.

[0007] In the second aspect, a deployment device for an AI model is provided, including: a receiving module, configured to receive an AI model in a first format that has been trained by a user; a first determination module, configured to determine the target inference device; a second determination module, configured to determine at least one second format of the AI ​​model that the target inference device can support; a third determination module, configured to determine multiple conversion paths based on the first format and the second format; a conversion module, configured to convert the AI ​​model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format; and an evaluation module, configured to deploy the multiple AI models in the second format to the target inference device for performance evaluation to determine the optimal AI model.

[0008] In a third aspect, an electronic device is provided, comprising: at least one memory configured to store computer-readable code; and at least one processor configured to call the computer-readable code and execute each step of the method provided in the first aspect.

[0009] In a fourth aspect, a computer-readable medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the processor executes each step in the method provided in the first aspect.

[0010] In a fifth aspect, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes computer-executable instructions, which, when executed, cause at least one processor to perform the steps in the method provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The following figures are intended only to illustrate and explain the embodiments of the present application and are not intended to limit the scope of the embodiments of the present application.

[0012] FIG1 is a schematic diagram of a method for deploying an AI model according to an embodiment of the present application;

[0013] FIG2 is a schematic diagram of a deployment device for an AI model according to an embodiment of the present application;

[0014] FIG3 is a schematic diagram of an electronic device according to an embodiment of the present application.

[0015] Explanation of reference numerals 100: AI model deployment method 101-106: method steps 20: AI model deployment device 21: receiving module 22: first determination module 23: second determination module 24: third determination module 25: conversion module 26: evaluation module 300: electronic device 301: processor 302: communication interface 303: memory 304: communication bus 305: program DETAILED DESCRIPTION

[0016] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that discussing these embodiments is intended only to enable those skilled in the art to better understand and implement the subject matter described herein, and is not intended to limit the scope of protection, applicability, or examples set forth in the claims. The functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the embodiments of the present application. Various examples may omit, replace, or add various processes or components as needed. For example, the described method may be performed in an order different from the described order, and various steps may be added, omitted, or combined. In addition, features described relative to some examples may also be combined in other examples.

[0017] As used herein, the term "including" and its variations are open terms meaning "including but not limited to". The term "based on" means "based at least in part on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other definitions may be included below, whether explicit or implicit. Unless the context clearly indicates otherwise, the definition of a term is consistent throughout the specification.

[0018] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0019] FIG1 is a schematic diagram of a method for deploying an AI model according to an embodiment of the present application. As shown in FIG1 , the method 100 for deploying an AI model includes:

[0020] Step 101: Receive an AI model in a first format that has been trained by a user.

[0021] Step 102: Determine the target inference device.

[0022] Optionally, based on the AI ​​model in the first format, a corresponding target inference device is recommended by an intelligent operation and maintenance platform, such as AIOps (Artificial Intelligence Operations) or MLOps (Machine Learning Operations). Optionally, the target inference device is determined by searching all target inference devices connected to the system. Optionally, the target inference device can be determined by a user-specified target inference device.

[0023] Step 103: Determine at least one second format of the AI ​​model that the target inference device can support.

[0024] Optionally, based on properties of the target inference device, at least one second format of the AI ​​model that the target inference device can support is determined.

[0025] Optionally, different target inference devices can support different numbers of model formats, so the second format may be one or more different ones. The attributes of the target inference device refer to the types of processing chips it contains, such as one or more of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a video processing unit (VPU), and a tensor processing unit (TPU). The more general the processing chips contained in the target inference device, such as CPU, GPU, etc., the more second formats it can support.

[0026] Step 104: Determine multiple conversion paths according to the first format and the second format.

[0027] Optionally, among the preset multiple model converters, for example, including from format A to format B, from format B to format C, from format C to format D, from format D to format E, and from format B to format D, traversing all conversion paths starting from the first format (assuming format A) and with the second format as the target format (assuming format D) is to retrieve all possible conversion paths. In this example, the possible conversion paths include:

[0028] Path 1: From format A to format B, then from format B to format C, and then from format C to format D;

[0029] Path 2: From format A to format B, and then from format B to format D.

[0030] Step 105 : Convert the AI ​​model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format.

[0031] In one embodiment, before step 104, all pre-set plug-ins are traversed to determine the formats that each plug-in can convert. Next, multiple conversion paths are determined based on the first format and the second format. Each of the multiple conversion paths includes at least one plug-in. Finally, the AI ​​model in the first format is converted using the plug-in included in each of the multiple conversion paths to obtain multiple AI models in the second format.

[0032] Optionally, the preset plug-in includes at least one, and a plug-in may include a program conversion from one format to another, or may also include multiple conversion programs, such as conversion from format A to format B, conversion from format B to format C, conversion from format A to format C, and so on.

[0033] Step 106: Deploy multiple AI models in the second format to the target inference device for performance evaluation to determine the optimal AI model.

[0034] The optimal AI model includes one that requires less inference time and / or consumes fewer resources on the target inference device. Optionally, the inference time required on the target inference device may be prioritized. Optionally, the consumed resources include CPU usage and memory occupancy. Optionally, weights are set for the required inference time and consumed resource metrics, and the corresponding weighted averages are used to evaluate performance.

[0035] This embodiment of the application starts with a user-trained AI model in a first format and targets an AI model format supported by the target inference device. Multiple conversion paths are then determined. Each of these conversion paths converts the trained AI model from the first format into multiple AI models in a second format. These multiple AI models in the second format are then deployed on the target inference device for performance evaluation, thereby determining the optimal AI model.

[0036] An embodiment of the present application provides a method for automatically deploying AI models to a target inference device, which not only reduces the manpower input of AI engineers and automation engineers, but also can quickly and accurately determine the optimal conversion path, thereby ensuring the stable operation of the target inference device.

[0037] FIG2 is a schematic diagram of an AI model deployment device according to an embodiment of the present application. As shown in FIG2 , the AI ​​model deployment device 20 includes:

[0038] The receiving module 21 is configured to receive an AI model in a first format that has been trained by a user.

[0039] The first determining module 22 is configured to: determine the target inference device.

[0040] The second determination module 23 is configured to determine at least one second format of the AI ​​model that can be supported by the target inference device.

[0041] The third determining module 24 is configured to determine a plurality of conversion paths according to the first format and the second format.

[0042] The conversion module 25 is configured to: convert the AI ​​model in the first format through each conversion path of the multiple conversion paths to obtain multiple AI models in the second format.

[0043] The evaluation module 26 is configured to: deploy the multiple AI models in the second format to the target inference device for performance evaluation to determine the optimal AI model.

[0044] Optionally, the AI ​​model deployment device also includes at least one plug-in, the configuration file of the plug-in specifies the supported model format conversion types, and the plug-in is used for converting the AI ​​model format.

[0045] Furthermore, different plug-ins can run on the same machine as the system or on different machines, ensuring that the entire system can operate on a single machine, a cluster of multiple servers, or a cloud platform with elastic computing capabilities. Different plug-ins can integrate the basic runtime environment they rely on, or use a shared integrated runtime environment.

[0046] The embodiments of this application can help users easily and quickly deploy AI models to target inference devices on factory sites.

[0047] FIG3 is a schematic diagram of an electronic device according to an embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device. As shown in FIG3 , the electronic device 300 may include: a processor 301, a communications interface 302, a memory 303, and a communication bus 304.

[0048] The processor 301 , the communication interface 302 , and the memory 303 communicate with each other via the communication bus 304 .

[0049] The communication interface 302 is used to communicate with other electronic devices or servers.

[0050] The processor 301 is configured to execute the program 302 , and specifically may execute the relevant steps in any one of the aforementioned method embodiments.

[0051] Specifically, the program 305 may include program codes, which include computer operation instructions.

[0052] Processor 301 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.

[0053] The memory 303 is used to store the program 303. The memory 303 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0054] The program 305 can be specifically used to enable the processor 301 to execute any one of the multiple method embodiments in the aforementioned embodiments.

[0055] The specific implementation of each step in program 305 can refer to the corresponding descriptions of the corresponding steps and units in the aforementioned AI model deployment method embodiment, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the aforementioned method embodiment, and will not be repeated here.

[0056] The present application also provides a computer-readable storage medium storing instructions for causing a machine to perform any of the multiple method embodiments described herein. Specifically, a system or device equipped with a storage medium can be provided, wherein the storage medium stores software program code that implements the functions of any of the above-described embodiments, and a computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.

[0057] In this case, the program code read from the storage medium itself can realize the function of any one of the above embodiments, so the program code and the storage medium storing the program code constitute part of this application.

[0058] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0059] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to perform any corresponding operation in the above-mentioned multiple method embodiments.

[0060] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0061] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or can be implemented as software or computer code that can be stored in a recording medium (such as CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or can be implemented as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware (such as ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the method shown here.

[0062] It should be noted that not all steps and modules in the above processes and system structure diagrams are required, and certain steps or modules can be omitted according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or may be implemented by certain components in multiple independent devices.

[0063] In the above embodiments, the hardware module can be implemented mechanically or electrically. For example, a hardware module can include a permanent dedicated circuit or logic (such as a dedicated processor, FPGA or ASIC) to complete the corresponding operation. The hardware module can also include programmable logic or circuits (such as a general-purpose processor or other programmable processors), which can be temporarily set by software to complete the corresponding operation. The specific implementation method (mechanical method, or dedicated permanent circuit, or temporarily set circuit) can be determined based on cost and time considerations.

[0064] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.

[0065] Nouns and pronouns referring to persons in this patent application are not limited to a specific gender.

Claims

1. A method for deploying an AI model, comprising: - Receiving (101) an AI model in a first format trained by a user; - Determining (102) the target inference device; - Determining (103) at least one second format of the AI model supported by the target inference device; - Determining (104) multiple conversion paths according to the first format and the second format; - Converting (105) the AI model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format; - Deploying the multiple AI models in the second format to the target inference device for performance evaluation (106) to determine the optimal AI model.

2. The method according to claim 1, wherein, - Before determining (104) multiple conversion paths according to the first format and the second format, the method further comprises: - Traversing all preset plugins to determine the formats that each plugin can be used for conversion; - The determining (104) multiple conversion paths according to the first format and the second format includes: - Determining multiple conversion paths according to the first format and the second format; wherein each conversion path in the multiple conversion paths includes at least one plugin; - The converting (105) the AI model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format includes: - Converting the AI model in the first format through the plugins included in each of the multiple conversion paths to obtain multiple AI models in the second format.

3. The method according to claim 1, wherein, The determining (103) at least one second format of the AI model supported by the target inference device includes: - Determining at least one second format of the AI model supported by the target inference device according to the attributes of the target inference device.

4. The method according to claim 1, wherein The optimal AI model includes: - An AI model that requires less inference time and / or consumes fewer resources on the target inference device.

5. The method according to claim 1, wherein The determining (102) the target inference device includes: - Recommending a corresponding target inference device according to the AI model in the first format through AIOps or MLOps.

6. An apparatus for deploying an AI model, comprising: - A receiving module (21) configured to receive an AI model in a first format trained by a user; - A first determining module (22) configured to determine the target inference device; - A second determining module (23) configured to determine at least one second format of the AI model supported by the target inference device; - A third determining module (24) configured to determine multiple conversion paths according to the first format and the second format; - A conversion module (25) configured to convert the AI model in the first format through each of the multiple conversion paths to obtain multiple AI models in the second format; - An evaluation module (26), configured to: deploy the multiple AI models in the second format to the target inference device for performance evaluation to determine the optimal AI model.

7. An electronic device (300), comprising: A processor (301), a communication interface (302), a memory (303), and a communication bus (304), where the processor (301), the memory (303), and the communication interface (302) complete mutual communication through the communication bus (304); The memory (303) is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the method for deploying the AI model according to any one of claims 1-5.

8. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for deploying the AI model according to any one of claims 1-5.

9. A computer program product, the computer program product being tangibly stored on a computer-readable medium and including computer-executable instructions, and the computer-executable instructions, when executed, cause at least one processor to execute the method for deploying the AI model according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and equipment for accelerating deployment of AI model

    CN112799680A

  • Model deployment method based on artificial intelligence platform and related equipment

    CN114861836A

  • Model conversion method and device and related equipment

    CN115238895A

  • Vehicle-mounted terminal model deployment method, device, equipment and medium

    CN115562707A

  • Model deployment method, related device and storage medium

    CN116954631A