Edge AI model-oriented automatic deployment and updating method and system

Through automated deployment and update methods and systems, edge device information is managed uniformly, intelligent push and update are realized, and model operation environment is automatically configured, which solves the problems of inefficiency of traditional deployment and update methods and confusing version management, and realizes efficient and stable deployment and update of edge AI models.

CN119987810AActive Publication Date: 2025-05-13KUAIJI XINYUN (QINGDAO) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510080643.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Traditional edge AI model deployment and update methods are difficult to meet the complex edge computing environment, resulting in low-efficiency deployment, untimely updates, and chaotic version management, affecting system stability and reliability.

Method used

It provides an automated deployment and update method and system for edge AI models. By uniformly managing edge device information, intelligent push and update are realized, and the model operation environment is automatically configured, inference services are optimized, and API interfaces and documents are automatically generated.

Benefits of technology

It significantly improves the deployment and update efficiency of edge AI models, reduces operation and maintenance costs, improves the stability and availability of services, and ensures the timeliness and accuracy of model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987810A_ABST
    Figure CN119987810A_ABST
Patent Text Reader

Abstract

The invention provides an edge AI model-oriented automatic deployment and updating method and system, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the unified management of the information of edge devices needing to be deployed by a user; transmitting the change parameters of the remote server model to the target edge device based on an automatic pushing mechanism of a time period; through intelligent detection, automatic script generation and remote execution, unattended configuration of the operation environment of the edge device model is realized; by deploying a platform tool on the edge device, a configuration file is intelligently generated, and seamless deployment and optimization of a model reasoning service on the edge device are realized; and automatically generating an API interface and an API document of the edge device model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and system for automated deployment and updating of edge AI models. Background Art

[0002] With the rapid development of artificial intelligence technology, edge computing, as a distributed computing architecture, pushes computing and data storage from central nodes to the edge of the network to improve response speed and reduce network bandwidth requirements. Edge AI model deployment is an important part of edge computing, which enables intelligent analysis and decision-making to be performed near the data source, thereby reducing data transmission delays and bandwidth consumption.

[0003] On the one hand, traditional model deployment and update methods are difficult to meet the complexity of edge computing environments. Edge devices often have limited resources, including computing power, storage capacity, network bandwidth, etc., and the devices are extremely scattered, with complex models and different operating systems. Manual deployment and update of models one by one not only consumes a lot of manpower and time costs, but is also prone to configuration errors, resulting in the model not being able to run normally.

[0004] On the other hand, business scenarios have extremely high requirements for the timeliness of edge AI models. In intelligent security monitoring scenarios, once a new threat identification algorithm is updated, it needs to be quickly deployed to the edge devices corresponding to each camera in order to respond to new security risks in a timely manner; on industrial production lines, if the product defect detection model is updated late, it may cause a large number of defective products to be produced, resulting in huge economic losses. Traditional deployment and update operations require manual download of the complete model and then upload it to the edge for deployment. This process is not only cumbersome but also time-consuming, especially for large models with large model parameters. Due to the need to transmit the complete model file, the network bandwidth consumption is large, and the update process is time-consuming, resulting in the model update often not being timely enough to meet the needs of rapid iteration. Therefore, how to update and accurately grasp the push time to complete the model update without affecting the normal business operation of the edge device is a problem that needs to be solved urgently.

[0005] Furthermore, with the frequent iteration of the model, the problem of chaotic version management has gradually become prominent. If the historical version cannot be properly preserved and the change log cannot be clearly recorded, it will be difficult to quickly roll back to the available version when the new model is unstable or does not match the business, which will seriously affect the stability and reliability of the system. In addition, different edge devices may run different operating systems, which increases the complexity of deployment. These problems seriously affect the efficiency and flexibility of model deployment, especially in scenarios that require frequent updates and large-scale deployment. Developers have to spend a lot of time on basic operation and maintenance work instead of focusing on model optimization and innovation. Summary of the invention

[0006] Based on the above problems, the present application provides an automated deployment and update method and system for edge AI models, which can realize unified management of edge device information, intelligently push and update according to device characteristics and business needs, and automatically configure the model operating environment, optimize reasoning services, and automatically generate API interfaces and documents, thereby meeting the diverse needs of edge AI model deployment.

[0007] The purpose of this application is achieved by the following technical solutions:

[0008] In a first aspect, the present application provides an automated deployment and update method for edge AI models, the method comprising:

[0009] Obtain edge device information through the user interface; manage the edge device information that users need to deploy in a unified manner through grouping and label settings;

[0010] Based on the time period, the automatic push mechanism transmits the change parameters of the remote server model to the target edge device;

[0011] Through intelligent detection, automated script generation and remote execution, unattended configuration of the edge device model operating environment is achieved;

[0012] By deploying platform tools on edge devices and intelligently generating configuration files, seamless deployment and optimization of model inference services on edge devices can be achieved;

[0013] Automatically generate API interfaces and API documentation for edge device models.

[0014] On the other hand, the present application provides a system for implementing the method for automatic deployment and update of edge AI models of the present application, the system comprising:

[0015] The edge service management module is used to obtain edge device information through the user interface; through grouping and label settings, the edge device information that users need to deploy is uniformly managed;

[0016] The model push module is used to automatically push the parameters of the remote server model to the target edge device based on the time period.

[0017] The environment configuration module is used to achieve unattended configuration of the edge device model operating environment through intelligent detection, automated script generation and remote execution;

[0018] The inference service deployment module is used to deploy platform tools on edge devices and intelligently generate configuration files to achieve seamless deployment and optimization of model inference services on edge devices;

[0019] The API automatic generation module is used to automatically generate API interfaces and API documents for edge device models.

[0020] In a third aspect, the present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in and executable on the memory, wherein the processor executes the program to implement the steps of any method described in the present application.

[0021] The beneficial effects of the present invention include: by centrally managing edge device information, including IP addresses, operating system types, and hardware configurations, the complexity of device management and maintenance is greatly simplified. Unified management not only improves work efficiency, but also ensures the accuracy and consistency of information; the automatic push mechanism based on time periods can intelligently identify and transmit the changing parameters of the remote server model to the target edge device. This feature not only reduces manual intervention, but also ensures the timeliness and accuracy of model updates. At the same time, through intelligent detection, automated script generation and remote execution, unattended configuration of the edge device model operating environment is achieved, further improving deployment efficiency; according to the actual load and business needs of the edge device, the optimal time period for model push is intelligently determined, thereby avoiding service interruptions or performance degradation caused by updates during business peaks; in addition, through intelligent generation of configuration files, seamless deployment and optimization of model inference services on edge devices are achieved, ensuring effective resource utilization and high availability of services. A new version number is generated for each model update, and historical version information is retained, making rollback operations simple and reliable; at the same time, the maximum number of edge device models to be retained is intelligently determined based on factors such as edge device capacity, computing resources, model capacity, and energy consumption, which not only ensures the diversity of the model, but also avoids the need to re-update the model from the remote service during backtracking, and avoids excessive resource occupation. When the edge device environment configuration fails, it can automatically roll back to the stable state before configuration and record the environment configuration log, ensuring the stability and reliability of the system, while facilitating problem troubleshooting and subsequent optimization. It automatically generates API code that complies with the RESTful specification based on the input and output characteristics of the model, and automatically generates API documents, which not only simplifies the API development and deployment process, but also improves the availability and ease of use of the API. When the model is updated, the relevant API is automatically updated to ensure the consistency of the API and the model.

[0022] In summary, this application significantly improves the deployment and update efficiency of edge AI models through automation, intelligence and optimization, reduces operation and maintenance costs, and improves the stability and availability of services. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the method for automatic deployment and update of edge AI models provided in an embodiment of the present application;

[0024] Figure 2 This is a schematic diagram of a process of updating a remote service model to an edge device provided in an embodiment of the present application;

[0025] Figure 3 It is a schematic diagram of the process of obtaining the changing parameters provided in the embodiment of the present application;

[0026] Figure 4 It is a schematic diagram of the push time period determination process provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] Below, the present application is further described in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.

[0028] See attached Figure 1 , the present application provides an automated deployment and update method for edge AI models, the method comprising:

[0029] Obtain edge device information through the user interface; manage the edge device information that users need to deploy in a unified manner through grouping and label settings;

[0030] Based on the time period, the automatic push mechanism transmits the change parameters of the remote server model to the target edge device;

[0031] Through intelligent detection, automated script generation and remote execution, unattended configuration of the edge device model operating environment is achieved;

[0032] By deploying platform tools on edge devices and intelligently generating configuration files, seamless deployment and optimization of model inference services on edge devices can be achieved;

[0033] Automatically generate API interfaces and API documentation for edge device models.

[0034] The working principle and effect of the above technical solution are as follows: users input or select relevant information of edge devices, such as device type, model, IP address, location, etc. through a dedicated user interface. Receive and store this information to provide basic data for subsequent edge AI model deployment and update; manage uniformly through grouping and labeling: users can group and label edge devices according to actual needs, such as grouping by geographic location, function type, department, etc. Through labeling, users can filter and locate target edge devices more quickly and improve management efficiency.

[0035] Real-time monitoring of AI model changes on remote servers, such as model version, parameter updates, performance optimization, etc.

[0036] When a model change is detected, the system records these changed parameters and prepares to transmit them to the target edge device. Users can set a specific time period, such as the off-peak period at night, to automatically push the model change parameters; according to the time period set by the user, the model change parameters on the remote server are automatically transmitted to the target edge device, or the system can automatically set the time period to transmit the model change parameters on the remote server to the target edge device; before the model is deployed, the operating environment of the target edge device is intelligently detected, including device performance, storage space, network connection status, etc., and based on the detection results, it is evaluated whether the hardware and software requirements for model deployment are met. If the deployment requirements are met, the deployment script is automatically generated according to the specific configuration of the target edge device and user needs. Through the remote execution mechanism, the system sends the deployment script to the target edge device and automatically executes the script to complete the configuration of the model operating environment. Deploy a dedicated platform tool on the edge device to receive, parse and execute model deployment instructions from the remote server. The platform tool is also responsible for monitoring the running status of the model reasoning service to ensure the stability and efficiency of the service. According to the model configuration information on the remote server and the actual situation of the target edge device, the configuration file is intelligently generated; the configuration file contains the parameters and settings required for the model reasoning service to ensure that the model can run correctly on the edge device; send the configuration file and model file to the target edge device and deploy them seamlessly through the platform tool. During the deployment process, optimize the model reasoning service, such as adjusting resource allocation, optimizing algorithm parameters, etc., to improve the performance and efficiency of the service. Automatically generate the API interface of the edge device model according to the model's functions and user needs. The API interface provides access to the model reasoning service and parameter settings to facilitate user integration and calling; automatically generate detailed API documents, including interface descriptions, parameter explanations, return value descriptions, etc.; the API document provides users with clear usage guides and reference information to help users better understand and use the API interface of the edge device model.

[0037] To sum up, the automated deployment and update method for edge AI models provided in this application realizes the unified management of edge device information, automatic push of model change parameters, unattended configuration of the model operating environment, seamless deployment and optimization of model reasoning services, and automatic generation of API interfaces and API documents through a series of automated and intelligent means. It greatly improves the deployment and update efficiency of edge AI models, reduces operation and maintenance costs, and provides strong support for the widespread application of edge AI.

[0038] In some embodiments, the step of obtaining edge device information through a user interface and uniformly managing edge device information that a user needs to deploy through grouping and tag settings includes:

[0039] Obtain edge device information through a user interface; the edge device information includes an IP address, an operating system type, and a hardware configuration of the edge device;

[0040] Automatically detect the operating system type of edge devices;

[0041] Adjust edge device model deployment strategies based on operating system characteristics;

[0042] Automatically install the Agent on the edge device and maintain communication with the server;

[0043] Regularly detect the status of edge devices and report to the management end in real time; and display monitoring alarm information through a visual UI interface;

[0044] Edge devices can be classified and managed through grouping and labeling.

[0045] The working principle of the above technical solution is:

[0046] The overall system architecture of this method is:

[0047] Server: As the central management node, it is responsible for receiving user instructions, storing device information, processing monitoring data, etc.

[0048] Agent: A client program installed on an edge device, responsible for collecting device information, executing server instructions, and reporting monitoring data.

[0049] User Interface (UI): Provides a visual interface for functions such as user registration, device management, and monitoring viewing.

[0050] The specific principles are:

[0051] The user registers an account and logs in to the system through the UI interface; the user enters the IP address, operating system type, hardware configuration and other information of the edge device to register the device; the system receives the information entered by the user, performs preliminary verification and storage; provides a multi-operating system compatibility detection mechanism, and the system can automatically detect the operating system type of the edge device (such as Linux, Windows, etc.). According to the detected operating system type, the system adjusts the deployment strategy and selects the appropriate Agent installation method and communication protocol. The edge device receives the installation package and automatically installs the Agent; after the Agent is installed, it establishes a connection with the server and maintains communication based on the authentication information provided by the system (such as SSH key or WinRM credentials). During the data transmission process, for Linux edge hosts, SSH key authentication is used for remote connection; for Windows systems, the WinRM (Windows Remote Management) protocol is used.

[0052] The agent can regularly check the online status and resource usage (CPU, memory, disk, etc.) of edge devices and report them to the management end in real time. Users can view monitoring-level alarm information based on the visual UI interface. The server analyzes the received monitoring data in real time and triggers alarm information according to the preset alarm rules.

[0053] The alarm information is displayed to the user through the UI interface, and the user can view the detailed alarm details and the list of affected devices.

[0054] Users can group and label devices through the UI interface to facilitate large-scale device management and batch operations; the grouping and labeling functions allow users to classify and manage devices based on their geographic location, usage and other information, thereby improving management efficiency.

[0055] The effect of the above technical solution is that users can register, manage and monitor all edge devices on a unified platform without logging into multiple systems or tools separately. The system can automatically detect the device operating system type and adjust the deployment strategy, and automatically install the Agent, which greatly reduces the time and workload of manual configuration. Through the device grouping and labeling functions, users can easily perform batch operations on large-scale devices, such as upgrading software, modifying configurations, etc., which improves management efficiency. The system supports multiple operating system types, including Linux and Windows, and can automatically detect and adapt to the characteristics of different operating systems to ensure smooth access and management of devices; according to the operating system type of the device, it can automatically select the appropriate communication protocol (such as SSH key authentication or WinRM protocol) for data transmission to improve the flexibility and compatibility of the system; the system uses security authentication methods such as SSH key authentication and WinRM protocol to ensure the security and integrity of the data transmission process; the agent can regularly check the online status and resource usage of the device, and report to the management end in real time when an abnormality is found, helping users to find and solve problems in time and improve the reliability and stability of the system; the system can collect and analyze the resource usage of the device in real time, such as CPU, memory, disk, etc., to help users understand the operating status and resource bottlenecks of the device; based on monitoring data, the system can intelligently schedule and optimize resource allocation to improve the operating efficiency and resource utilization of the device; the modular design is adopted to facilitate subsequent functional expansion and upgrade. Users can intuitively view device information, monitoring data and alarm information through a visual UI interface, reducing the difficulty of operation and maintenance costs.

[0056] In summary, the unified edge service management system provides users with a comprehensive, efficient and secure edge device management solution by improving management efficiency, enhancing system compatibility, improving security and reliability, optimizing resource utilization, and facilitating expansion and maintenance.

[0057] See attached Figure 2 In some embodiments, the time period-based automatic push mechanism transmits the change parameters of the remote server model to the target edge device; including:

[0058] Obtain the model and target edge device to be deployed through the model selection and configuration interface;

[0059] Compare the parameter differences between the remote service model and the edge device model to obtain the changed parameters;

[0060] Setting a push time period, and transmitting the change parameter to the edge device within the push time period;

[0061] Generate a new version number for each model update, retain the historical version information on the remote server, and select the retained model version on the edge device;

[0062] In response to the rollback request, rolling back the edge device model to a target version;

[0063] Monitor the model update status in real time and record the change log.

[0064] The working principle and effect of the above technical solution are as follows: In the automatic push mechanism based on time period, the change parameters of the remote server model are efficiently transmitted to the target edge device. The realization of this process depends on the following key steps and components:

[0065] Allow users to select the model to be deployed and the target edge device through the user interface.

[0066] Users can set the start and end time of the model push on this interface, or select a specific periodic push time period (such as 2 am to 4 am every day); or the system automatically sets the push time period;

[0067] First, the latest model parameters on the remote server are obtained and compared with the model parameters currently running on the edge device; through algorithm analysis, the differences between the two, that is, the changing parameters, are identified.

[0068] The task queue is used to manage push tasks to ensure that the changed parameters are pushed to the edge devices in an orderly manner according to the priority and load conditions within the specified time period; in order to cope with large-scale concurrent pushes, the system adopts a concurrency control strategy, such as limiting the number of devices that push at the same time to maintain system stability; only the changed parameter parts are transmitted instead of the entire model, thereby greatly reducing the amount of data transmission. For large models, the system adopts the method of parameter block transmission and appends a checksum after each block of data to ensure the accuracy of the incremental update.

[0069] Each push generates a new version number and updates it to the edge device. The server retains all historical version information so that users can roll back versions when needed. Edge devices can selectively retain which version of the model to quickly roll back calls without having to transfer again. Maintain a parameter change log to record the specific changes of each update in detail, including the time of update, updated parameters, updated version, etc. Provide a real-time monitoring interface where users can view key information such as push progress and success rate. Automatically record detailed push logs, including the start time and end time of the push task, the list of pushed devices, and the push results (success / failure). These log information is crucial for troubleshooting and performance optimization.

[0070] Through the above process, the changed parameters of the remote server model can be transmitted to the target edge device efficiently and accurately, while providing rich version management and logging functions, providing strong support for users' model deployment and operation and maintenance work.

[0071] See attached Figure 3 In some embodiments, the step of comparing the parameter differences between the remote service model and the edge device model to obtain the changed parameters includes:

[0072] For the lightweight model of the edge device, a parameter mapping table is constructed to obtain the mapping relationship between the edge device model and the corresponding remote service full model;

[0073] If there are multiple versions of lightweight models on the edge device, compare the latest remote service full model with the full model versions corresponding to the multiple versions of lightweight models on the edge device to obtain the smallest parameter difference;

[0074] According to the parameter mapping table and the minimum parameter difference, the change parameters of the lightweight model of the edge device are obtained.

[0075] The working principle of the above technical solution is: first, perform structural analysis on the full model in the remote service and the lightweight model on the edge device; this includes but is not limited to identifying the layers of the model, the parameters of each layer (such as weights and biases), activation functions, etc.; based on the structural similarity and functional correspondence between the two, define the mapping rules between parameters; for example, the lightweight model may be derived from the full model through pruning or quantization, then it is necessary to create a mapping based on these conversion rules. Using the above rules, find the corresponding item in the full model for each lightweight model parameter, and record this mapping relationship in a parameter mapping table. The table can be a data structure containing the lightweight model parameter ID, the corresponding full model parameter ID, and any necessary conversion information (such as scaling factors).

[0076] For edge devices with multiple versions of lightweight models, determine the minimum parameter difference. Identify all available lightweight model versions on the edge device; map each lightweight model version back to the corresponding full model version according to the parameter mapping table. If possible, use the stored historical full model version directly to save computing resources. Compare the latest remote service full model with each reconstructed full model version and calculate the parameter difference between the two. This process can be a layer-by-layer, parameter-by-parameter comparison of numerical differences, or a more advanced algorithm can be used to measure the change in overall model performance; select the smallest one from all calculated differences, that is, the version closest to the latest remote service full model. This version of the lightweight model is considered to be the most suitable baseline for the current update; extract the specific changed parameters of the lightweight model according to the parameter mapping table and the minimum parameter difference; determine which parameters have changed based on the results of the minimum parameter difference; use the parameter mapping table to map these changes back to the specific parameter locations of the lightweight model. This is because what needs to be updated in the end is the lightweight model on the edge device, not the full model. Based on the located changed parameters, an incremental update package is generated, which only contains those parameters that have actually changed. For large models, these parameters can be further processed in blocks to ensure transmission efficiency and accuracy. Before sending the incremental update package, a verification step is performed to ensure the correctness of the update. Once confirmed, the incremental update package can be pushed to the edge device for application.

[0077] The effects of the above technical solution are as follows: by constructing a parameter mapping table and transmitting only the changed parameters, the amount of data that needs to be transmitted is greatly reduced, which not only speeds up the update speed, but also reduces bandwidth consumption; since only the changed parts need to be processed, the edge device can apply the update faster, reducing downtime and resource usage; edge devices usually have limited resources, and by incremental updates instead of re-downloading the entire model, storage space can be effectively saved; only the changed parts are updated on the edge device, which reduces unnecessary computing work and improves the response speed and energy efficiency of the device. In some cases, the latest full remote service model may not be updated based on the last model, but based on an earlier version (such as the last model). If the latest full remote service model is directly compared with the most recently deployed lightweight model on the edge device, it may result in a large parameter difference, thereby increasing the amount and complexity of the incremental update data. Therefore, by comparing multiple versions of lightweight models, the version closest to the latest full remote service model is selected for update, which ensures the stability and compatibility of the model update, reduces the risk of update failure, and reduces the number of changed parameters that need to be transmitted. The update process is completed faster, reducing the waiting time and resource usage on the edge device.

[0078] See attached Figure 4In some embodiments, the setting of a push time period, during which the change parameter is transmitted to the edge device, includes:

[0079] Determine the predicted transmission time of the current changed parameter based on the changed parameter and the historical transmission time;

[0080]

[0081] Among them, T cy is the predicted transmission time of the changed parameter; D c is the size of the parameter of this change; D i is the size of the i-th change parameter in the historical statistical data, T i is the actual transmission time of the i-th change parameter in the historical data; T iy is the predicted transmission time of the i-th changed parameter in the historical data; n is the number of model updates; wi is the first weight coefficient; q is the index, which is a positive integer from 1 to n.

[0082] According to the predicted transmission time of the current change parameter, a first time window is set; the first time window T w Satisfies: k1×T cy ≤T w ≤k2×T cy , where k1 and k2 are coefficients, 1≤k1≤3, 3≤k2≤10, k1 <k2;

[0083] Divide a day into multiple first time windows; predict the load and business needs of edge devices in each first time window based on historical data;

[0084] Determine the push time period based on the edge device load and business requirements of the first time window.

[0085] In some embodiments, determining the push time period according to the edge device load and business requirements includes:

[0086] If the business demand allows the update, and the load of the edge device in the first time window satisfies the requirement of being less than the preset load, the first time window is marked;

[0087] If P consecutive first time windows are marked, the P consecutively marked first time windows are taken as a first time window group;

[0088] If there are P marked first time windows, and the time interval between any two adjacent marked windows is less than a preset interval threshold, the P first time windows are combined into a first time window group;

[0089] The initial time of the first time window group closest to the remote service model update is used as the initial transmission start time.

[0090] If the transmission start time of multiple edge devices is the same, the transmission tasks are first sorted according to the priority;

[0091] In the same priority level, the second sorting is performed according to the transmission amount of each transmission task from small to large;

[0092] According to the second sorting result and the current network load, dynamically determine the number of tasks that can be processed simultaneously in each batch;

[0093] If at a certain moment there are multiple tasks in the second order (such as the first 5) that can be safely processed under the current network conditions, these tasks will start transmission immediately;

[0094] For subsequent tasks (such as the 6th and later), before each new task starts, check whether the current network load and the remaining first time window groups meet the preset conditions;

[0095] If a previous task has been completed, and the addition of a new task (such as the sixth one) will not cause the network load to exceed the safety threshold, and the duration of the remaining first time window group of the task still meets the requirements of the predicted transmission duration, the transmission of the next task will be started immediately; if the new task does not meet the requirements, it will be skipped and the next task to be processed will be evaluated instead;

[0096] If all subsequent tasks meet the preset duration requirements but fail to meet the network load requirements, wait for more resources to be released;

[0097] Once the network load allows the addition of new tasks without exceeding the safety threshold, the system will recheck the remaining pending tasks and immediately start the eligible tasks;

[0098] For tasks that cannot be transmitted due to resource limitations in the first time window group, they are marked as "pending" and their original priority, transmission volume and other information are recorded; if a task cannot be transmitted in more than two time window groups, its priority and / or the order of the second sort is increased.

[0099] The working principle of the above technical solution is:

[0100] First, the predicted transmission time of the parameter being changed is calculated based on the size of the parameter being changed and historical statistical data (including the size of the parameter being changed in previous times and its actual transmission time, as well as the predicted transmission time of the parameter being changed in previous times, which may not be used directly in this decision-making process but can be used to evaluate the accuracy of the prediction model); the formula comprehensively considers the linear relationship between the parameter size and the transmission time as well as the deviation of historical transmission to improve the accuracy of the prediction. By introducing a weight coefficient to adjust the impact of each historical transmission on the current prediction, the weight coefficient adopts an exponential decay form, so that the closer the historical data is to the present, the higher the weight. When calculating the predicted transmission time, the transmission performance of the recent model update will be given priority consideration. This is consistent with the actual network environment and the dynamic changes in equipment performance, because the recent situation is more relevant to the current prediction, and can effectively reduce the prediction deviation caused by the interference of long-term historical data, so that the predicted time is closer to the actual value; the deviation between the actual time and the predicted time in the historical transmission is taken into account, and the prediction value is corrected from different dimensions in conjunction with the estimation based on the historical transmission efficiency in the first half, to make up for the shortcomings of a single estimation method and further improve the prediction accuracy; k1 and k2 are coefficients used to set the range of the first time window. Based on the predicted time, a reasonable time range is given to ensure that the update process has sufficient flexibility to accommodate possible network fluctuations, temporary equipment freezes and other unexpected situations; ensure that the transmission has enough buffer time but is not too long.

[0101] Divide a day into multiple such first time windows to evaluate the load and business needs of edge devices in different time periods. Use historical data to predict the expected load and business needs of edge devices in each first time window. The expected load includes key indicators such as CPU usage and memory usage. Check the first time windows one by one. If the business needs of this period allow updates, that is, the update operation will not interfere with the development of core business, and the device load is lower than the preset load, then mark the window as available; if there are P consecutive first time windows marked as available, then combine these consecutive windows into a first time window group, indicating that it is more appropriate to update during these time periods; if there are P marked windows, and the time interval between them is less than the preset interval threshold, these windows can also be combined into a first time window group to utilize fragmented available time.

[0102] From all marked and combined window groups, the initial time of the first time window group that is closest to the remote service model update is selected as the start time of the transmission. This can ensure that the model update is as timely as possible while reducing the impact on the normal operation of the edge device.

[0103] The effects of the above technical solution are as follows: through the prediction formula based on historical data and the current size of the changing parameters, the system can more accurately estimate the time required for transmission, reduce unnecessary waiting time, and ensure efficient use of resources; by considering the difference between the historical transmission duration and the predicted duration, the impact of short-term fluctuations on the prediction results can be effectively reduced, providing a more stable prediction performance, and by combining historical transmission data and exponential decay weights, the transmission duration of the current changing parameters can be more accurately predicted, reducing the prediction deviation caused by changes in network conditions. The time window set according to the predicted transmission duration is flexible enough to adapt to network fluctuations and changes in the processing power of edge devices, ensuring that the transmission process can be carried out under optimal conditions, thereby improving overall efficiency. Using historical data to predict the load and business needs of edge devices allows the system to transmit during periods of low load, avoiding resource contention and performance degradation that may result from transmission during peak hours. By marking and combining the first time window, forming the first time window group, and selecting the group closest to the remote service model update as the starting point, the update timeliness and device stability are taken into account. On the one hand, the update will not be delayed due to excessive waiting for the ideal window, so that the new model change parameters can empower the edge device as soon as possible; on the other hand, with the help of the window group strategy, it is also guaranteed that the device is in a state suitable for update at the beginning of the push, ensuring the continuity and stability of key services, while optimizing the use of resources; by intelligently selecting the push time period, the system reduces the risk of transmission failure and improves the reliability of the system; it can dynamically adjust the push strategy according to historical data and real-time conditions to adapt to different network environments, edge device configurations and business needs; through the combination of priority sorting and transmission volume sorting, it ensures the effective use of network resources; high-priority and small-transmission tasks are processed first, reducing the long-term occupation of network bandwidth by large tasks and improving the overall resource utilization efficiency; dynamically determine the number of tasks that can be processed simultaneously in each batch, and adjust according to the real-time network load. This enables the system to flexibly respond to different network conditions and avoid service interruptions or delays caused by network congestion. Before each new task starts, check whether the current network load and the remaining first time window group meet the preset conditions to ensure that the task is started at the most appropriate time without wasting time or causing resource overload.Tasks that fail to be transmitted on time are marked and relevant information is recorded. When transmission has not occurred for more than two time window groups, their priority and / or second sort order is increased, ensuring that critical tasks will not be postponed indefinitely, thereby enhancing system reliability and user experience. When subsequent tasks meet the preset duration requirements but do not meet the network load requirements, the system will wait for more resources to be released before evaluating the start conditions, effectively reducing unnecessary waiting time and potential network congestion problems. The network status and task completion status are continuously monitored. Once the network load allows new tasks to be added without exceeding the safety threshold, eligible tasks are immediately started. The immediate response mechanism improves the overall transmission efficiency and shortens the total time for task completion. "Pending" tasks are comprehensively sorted together with new tasks to ensure that unfinished tasks are not ignored, while also allowing new tasks to be started as soon as conditions permit, achieving a balance between fairness and efficiency.

[0104] In some embodiments, selecting a model version to retain at the edge device includes:

[0105] Determine the maximum number of edge device models to be retained based on the capacity of the edge device, the computing resources of the edge device, the capacity of the model, and the energy consumption of the model;

[0106]

[0107] Among them, g max The maximum number of edge device models to be retained; M compute The average computing resources required for a single edge model to run; this can be obtained through historical data or model evaluation; M capacity is the average capacity of a single model; M energy is the average energy consumption of a single model when running on an edge device; E compute is the computing resource of the edge device (expressed in some quantitative indicators, such as the number of floating-point operations per second FLOPS); energy is the total energy consumption that the edge device can bear in a specific time period (such as a day); R ratio is the reserved resource ratio, ranging from (0,1), α is the coefficient, ranging from (0,1); is the total capacity of the currently reserved m models, is the total computing resource requirement of the currently reserved m models, is the total energy consumption of the currently reserved m models; floor() is rounded down; E capacity is the total capacity of the edge devices;

[0108] Determine the model score based on the historical prediction results of the edge model and the rollback selection results;

[0109] Determine the model version retained by the edge device based on the model score and the model update frequency;

[0110] The model score is:

[0111] S u =x1×Pa u +x2×Pmax u +x3×Pmin u +x4×Rc u -x5×Rbc u

[0112] Among them, x1, x2, x3, x4, x5 are the second weight coefficients, Pacc u Pmax is the average accuracy of the u-th model in the past h predictions; u is the maximum accuracy of the u-th model in the past h predictions; pmin u Rc is the minimum accuracy of the u-th model in the past h predictions; u Rbc is the normalized value of the number of times of rolling back from other models to the u-th model; u is the normalized value of the number of rollbacks from the u-th model to other models;

[0113] According to the order of model scores from high to low, the top g models are selected as the model versions retained by the edge device;

[0114]

[0115] g min is the preset minimum number of retained models; fa is the model update frequency, fy is the preset update frequency, g max It is the maximum number of edge device models to be retained. min() is the minimum value and ceiling is rounded up.

[0116] The working principle of the above technical solution is: the computing resources, total capacity, and tolerable total energy consumption of the edge device are taken as input; the average computing resources, average capacity, and average energy consumption required for the operation of a single edge model are calculated; the available amount of each resource (capacity, computing resources, energy consumption) is calculated, and the smallest available amount of the three resources is taken as the bottleneck resource considering the reserved resource ratio, and the maximum number of retained models is calculated based on the bottleneck resource and the adjustment factor α; the reserved resource ratio sets the resource ratio reserved by the edge device to prevent full load.

[0117] The performance stability of the model is evaluated based on the average accuracy, maximum accuracy, and minimum accuracy of the model in the past h predictions. The number of times the model is rolled back and the number of times it is rolled back from the model are counted and normalized.

[0118] According to the order of model scores from high to low, the top g models are selected as the model versions retained by the edge devices.

[0119] If the update frequency is lower than the preset frequency, keep g min models; if the update frequency is higher than the preset frequency, the number of retained models is increased according to the frequency difference, but not more than g max .

[0120] The effect of the above technical solution is: the maximum number of models that the edge device can safely retain is accurately calculated through a formula, ensuring the effective use of device resources and avoiding waste or excessive occupation of resources.

[0121] Through the model scoring mechanism, the prediction accuracy, stability and rollback behavior of the model are comprehensively considered, and the model version with excellent performance and reasonable resource consumption is selected for retention, achieving a balance between performance and resource consumption; the reserved resource ratio ensures that there are enough resources to cope with sudden demands even in extreme cases, improving the stability and reliability of the system. A new version number is generated for each model update, and historical version information is retained, so that the edge device can quickly roll back to the previous stable version as needed, improving the flexibility and security of model updates. Model versions with stable performance and few rollbacks are retained on edge devices first, reducing the risks caused by unstable model updates; by selecting the optimal model version to be retained on the edge device, efficient response of the service is ensured and user experience is improved; retaining multiple model versions on the edge device ensures that even if there is a problem with the new version model, it can quickly switch to other stable versions, ensuring business continuity and stability.

[0122] In some embodiments, the unattended configuration of the edge device model operating environment is achieved through intelligent detection, automated script generation and remote execution; including:

[0123] Match the Agent program of the edge device according to the edge operating system type;

[0124] Using the Agent program to intelligently detect the operating environment of the edge device, identify installed software and libraries; and identify missing dependencies;

[0125] Receive server instructions from the remote controller, and based on model dependency requirements, the Agent program intelligently selects the appropriate version, resolves version conflicts or compatibility issues, and generates and executes an environment installation script adapted to the edge device;

[0126] When the edge device environment configuration fails, it automatically rolls back to the stable state before the configuration, records the environment configuration log and uploads the environment configuration log to the server.

[0127] The working principle and effect of the above technical solution are as follows: First, identify the operating system type of the edge device (such as Linux distribution, Windows IoT, etc.). Automatically select and deploy the most suitable Agent program according to the operating system type. The Agent program has the following functions:

[0128] Automatically detect installed software and libraries on edge devices and identify missing dependencies.

[0129] Automatically generate an adaptive installation script based on the model's dependency requirements and the environment of the edge device.

[0130] When encountering version conflicts or compatibility issues, it can intelligently select the appropriate version or provide solutions.

[0131] Use the Agent program to fully detect the operating environment of the edge device to ensure that all necessary dependencies are correctly identified: After the Agent program is started, it automatically scans the current environment of the edge device and records the installed software, libraries and their version information. Based on the requirements of the model to be deployed, analyze the required dependencies and compare them with the existing environment to identify any missing or incompatible dependencies. Check whether the software and library versions in the existing environment conflict with the requirements of the new model to prevent potential problems in advance.

[0132] Send environment installation instructions to the edge through remote commands, automatically generate and execute environment installation scripts, including: after the remote controller receives the environment installation request from the user or system, it sends the instruction to the Agent program on the corresponding edge device. The Agent program automatically generates an adaptive installation script based on the received instructions and model dependency requirements combined with the environment monitoring results. The Agent program executes the generated installation script to automatically install the necessary operating environment (such as Python, PyTorch, etc.) on the terminal system, and use package management tools suitable for the operating system (such as apt, yum, pip, etc.); during the installation process, the Agent program records the detailed installation progress and status in real time, and uploads these logs to the server. Users can view the installation progress on the web page to ensure the stability and recoverability of the environment configuration, record each configuration operation in detail, and support automatic rollback when failure occurs.

[0133] The Agent program records the operation content, time and results in detail in each configuration step to form a complete environment configuration log. Real-time monitoring of status changes during the configuration process, timely detection and handling of possible problems. If an error or failure occurs during the configuration process, the Agent program can automatically roll back to the previous stable state to ensure that the existing service will not be affected by the configuration failure. All configuration logs will be uploaded to the server for subsequent auditing and troubleshooting.

[0134] In some embodiments, the configuration files are intelligently generated by deploying platform tools on edge devices to achieve seamless deployment and optimization of model inference services on edge devices; including:

[0135] Deploy a suitable version of the platform tool on the edge device according to the edge device information;

[0136] Generate an optimal configuration file based on model characteristics and edge device resources; the configuration file includes resource allocation and concurrency settings;

[0137] Remotely start the inference service, monitor the service status in real time, and set up a heartbeat monitoring mechanism;

[0138] Automatically diagnose and recover from failures in starting or abnormal operation of reasoning services;

[0139] Collect and analyze the operating data of edge device inference services, and provide optimization suggestions based on the analysis results.

[0140] The working principle and effect of the above technical solution are as follows: first, the operating system type, hardware configuration (such as CPU architecture, memory size, storage space, etc.) and other relevant information of the edge device are collected. Based on the above information, the developed automatic deployment script will automatically select the platform tool version such as Ollama and Pytorch that is most suitable for the device. Ensure that the selected version is compatible with the device hardware and operating system to maximize performance and stability. Through the remote execution mechanism, the selected version of the platform tool is automatically deployed to the edge device to reduce manual intervention and speed up deployment. Perform feature analysis on the model to be deployed, including but not limited to model size, input and output format, computational complexity, etc. Evaluate the available resources of the edge device, such as the number of CPU cores, memory capacity, network bandwidth, etc. Use the configuration file automatic generator to comprehensively consider the model characteristics and device resource conditions to generate the optimal configuration file. The configuration file includes but is not limited to the reasonable allocation of CPU, memory and other resources to ensure the efficient operation of the model inference service. According to the device capabilities and model requirements, set the optimal concurrent processing capability to improve throughput and response speed. Remotely start the inference service deployed on the edge device from the central management system through remote commands or API interfaces. Introduce a real-time monitoring mechanism to continuously monitor the status of the inference service, including key indicators such as CPU usage, memory usage, and network traffic. Set up a heartbeat monitoring mechanism to regularly check the health of the service to ensure that it remains online. If an anomaly is detected, immediately trigger an alarm or take appropriate measures.

[0141] Once the monitoring mechanism finds that the service fails to start or runs abnormally, it immediately triggers the error handling process. The system automatically diagnoses the fault, analyzes logs, error codes, and other related data, and determines the cause of the problem. Based on the diagnosis results, the system attempts to automatically fix the problem, such as restarting the service, adjusting resource configuration, rolling back to a stable version, etc. For certain types of errors, the system can retry multiple times according to the preset strategy until the service returns to normal or the maximum number of retries is reached.

[0142] The system will regularly collect the operating data of the inference service, including but not limited to performance indicators, resource utilization, latency, etc. Through the built-in data analysis module, the collected data is deeply analyzed to identify potential performance bottlenecks and improvement points. Based on the analysis results, the system provides users with specific performance optimization suggestions, such as adjusting configuration parameters, upgrading hardware, optimizing model structure, etc. Users can make adjustments based on these suggestions to further improve the performance and efficiency of the inference service, forming a closed loop of continuous improvement.

[0143] Through the above method, not only the seamless deployment of model inference services on edge devices is achieved, but also comprehensive monitoring, fault handling and performance optimization functions are provided, which reduces manual intervention and improves the reliability and efficiency of services.

[0144] In some embodiments, the automatically generating an API interface and an API document of an edge device model includes:

[0145] Automatically generate API code that complies with RESTful specifications based on the input and output characteristics of the model;

[0146] Through the API program automatic deployer, the automatically generated API code is encapsulated into a container and automatically pushed to the cloud platform; the API service is started through containerization;

[0147] Remotely call the API service through the edge Agent program;

[0148] Automatically generate API documentation for edge device models;

[0149] When the model is updated, the related APIs are automatically updated.

[0150] The working principle of the above technical solution is: according to different types of models (such as image classification, text generation, etc.), corresponding API templates are provided, including: establishing an API template library containing multiple model types (such as image classification, text generation, speech recognition, etc.), each template predefines common API endpoints, request methods, parameter structures and response formats; according to the model type provided by the user, the most suitable API template is automatically matched to ensure that the generated API code complies with the RESTful specification and meets the requirements of specific models.

[0151] Automatically generate API code that complies with RESTful specifications based on the input and output features of the model, including: extracting input and output features from the model, including data format, parameter type, expected results, etc., and automatically generating complete API code based on the selected API template and extracted feature information. The code includes functional modules such as routing definition, request processing logic, data validation, and error handling. Ensure that the generated API code strictly follows RESTful design principles, such as using HTTP verbs (GET, POST, PUT, DELETE, etc.), resource path naming rules, etc.

[0152] The automatically generated API program code is encapsulated as a container and automatically pushed to the PaaS cloud platform, started in a containerized form, and remotely called based on the edge agent, including: packaging the generated API code and its dependencies into a Docker image or other container format to ensure that it can run consistently in any environment that supports containers. The container image is automatically pushed to the PaaS cloud platform (such as a Kubernetes cluster) through the API program automatic deployer; the deployer is responsible for managing the container life cycle, including operations such as starting, stopping, and expanding. Start the API service in a containerized form on the cloud platform to ensure the isolation and stability of the service; remotely call the API service through the edge agent program to achieve seamless mounting of the cloud and edge reasoning services. The user's reasoning service request is forwarded to the edge AI model for processing, and the result is automatically transmitted back to the cloud after the reasoning is completed.

[0153] Automatically generate API documentation for edge device models to ensure that the documentation is detailed and easy to understand, including: Use Swagger or similar tools to automatically generate API documentation. These tools can parse comments and metadata in API code and automatically generate detailed documentation content. The generated API documentation includes interface descriptions, parameter descriptions, request examples, response formats, etc., to ensure that developers can quickly understand and use the API. Whenever the API code changes, the document generation system automatically regenerates the latest documentation to ensure that the documentation is always synchronized with the actual API.

[0154] Generate a new version number for each API update and retain historical version information. Each version has an independent API path (such as / v1, v2, etc.) to ensure that different versions can exist at the same time without interfering with each other. When designing a new version of the API, ensure that existing clients are not affected. For non-destructive changes, try to maintain the availability of the old version; for destructive changes, provide a reasonable transition plan. When the model is updated, the API code automatic generator will automatically detect changes and update related APIs. The new API version will be automatically deployed and released according to the established process to ensure a smooth and uninterrupted update process.

[0155] The present application provides a system for implementing the method for automatically deploying and updating edge AI models in the embodiments of the present application, the system comprising:

[0156] The edge service management module is used to obtain edge device information through the user interface; through grouping and label settings, the edge device information that users need to deploy is uniformly managed;

[0157] The model push module is used to automatically push the parameters of the remote server model to the target edge device based on the time period.

[0158] The environment configuration module is used to achieve unattended configuration of the edge device model operating environment through intelligent detection, automated script generation and remote execution;

[0159] The inference service deployment module is used to deploy platform tools on edge devices and intelligently generate configuration files to achieve seamless deployment and optimization of model inference services on edge devices;

[0160] The API automatic generation module is used to automatically generate API interfaces and API documents for edge device models.

[0161] The present application also provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method described in the present application are implemented.

[0162] This application is explained from the perspectives of purpose of use, effectiveness, progress and novelty, and has met the functional enhancement and usage requirements emphasized by the Patent Law. The above description and drawings of this application are only the preferred embodiments of this application, and are not intended to limit this application. Therefore, all structures, devices, features, etc. that are similar or identical to this application, that is, all equivalent replacements or modifications made in accordance with the scope of the patent application of this application, should fall within the scope of protection of the patent application of this application.

Claims

1. The method for automatic deployment and update of edge AI models is characterized by: The method comprises: Obtain edge device information through the user interface; manage the edge device information that users need to deploy in a unified manner through grouping and label settings; Based on the time period, the automatic push mechanism transmits the change parameters of the remote server model to the target edge device; Through intelligent detection, automated script generation and remote execution, unattended configuration of the edge device model operating environment is achieved; By deploying platform tools on edge devices and intelligently generating configuration files, seamless deployment and optimization of model inference services on edge devices can be achieved; Automatically generate API interfaces and API documentation for edge device models.

2. The method according to claim 1, characterized in that The automatic push mechanism based on time period transmits the change parameters of the remote server model to the target edge device; including: Compare the parameter differences between the remote service model and the edge device model to obtain the changed parameters; Setting a push time period, and transmitting the change parameter to the edge device within the push time period; Generate a new version number for each model update, retain historical version information on the remote server, and select the retained model version on the edge device; In response to the rollback request, rolling back the edge device model to a target version; Monitor the model update status in real time and record the change log.

3. The method according to claim 2, characterized in that The step of comparing the parameter differences between the remote service model and the edge device model to obtain the changed parameters includes: For the lightweight model of the edge device, a parameter mapping table is constructed to obtain the mapping relationship between the edge device model and the corresponding remote service full model; If there are multiple versions of lightweight models on the edge device, compare the latest remote service full model with the full model versions corresponding to the multiple versions of lightweight models on the edge device to obtain the smallest parameter difference; According to the parameter mapping table and the minimum parameter difference, the change parameters of the lightweight model of the edge device are obtained.

4. The method according to claim 2, characterized in that: The setting of a push time period, during which the change parameter is transmitted to the edge device, comprises: Determine the predicted transmission time of the current changed parameter based on the changed parameter and the historical transmission time; According to the predicted transmission duration of the current change parameter, a first time window is set; the first time window T w Satisfies: k1×T cy ≤T w ≤k2×T cy , where k1 and k2 are coefficients, 1≤k1≤3, 3≤k2≤10, k1 <k2;T cy The predicted transmission time of the changed parameter; Divide a day into multiple first time windows; predict the load and business needs of edge devices in each first time window based on historical data; Determine the push time period based on the edge device load and business requirements of the first time window.

5. The method according to claim 4, characterized in that Determining the push time period according to the edge device load and business requirements includes: If the business demand allows the update, and the load of the edge device in the first time window satisfies the requirement of being less than the preset load, the first time window is marked; If P consecutive first time windows are marked, the P consecutively marked first time windows are taken as a first time window group; If there are P marked first time windows, and the time interval between any two adjacent marked windows is less than a preset interval threshold, the P first time windows are combined into a first time window group; The initial time of the first time window group closest to the remote service model update is used as the initial transmission start time.

6. The method according to claim 2, characterized in that The model version selected to be retained on the edge device includes: Determine the maximum number of edge device models to be retained based on the capacity of the edge device, the computing resources of the edge device, the capacity of the model, and the energy consumption of the model; Determine the model score based on the historical prediction results of the edge model and the rollback selection results; Determine the model version retained by the edge device based on the model score and the model update frequency; According to the order of model scores from high to low, the top g models are selected as the model versions retained by the edge device; g min is the preset minimum number of retained models; fa is the model update frequency, fy is the preset update frequency, g max It is the maximum number of edge device models to be retained. min() is the minimum value and ceiling is rounded up.

7. The method according to claim 1, characterized in that The unattended configuration of the edge device model operating environment is achieved through intelligent detection, automated script generation and remote execution; including: Match the Agent program of the edge device according to the edge operating system type; Using the Agent program to intelligently detect the operating environment of the edge device, identify installed software and libraries; and identify missing dependencies; Receive server instructions from the remote controller, and based on model dependency requirements, the Agent program intelligently selects the appropriate version, resolves version conflicts or compatibility issues, and generates and executes an environment installation script adapted to the edge device; When the edge device environment configuration fails, it automatically rolls back to the stable state before the configuration, records the environment configuration log and uploads the environment configuration log to the server.

8. The method according to claim 1, characterized in that: The method deploys platform tools on edge devices and intelligently generates configuration files to achieve seamless deployment and optimization of model inference services on edge devices; including: Deploy a suitable version of the platform tool on the edge device according to the edge device information; Generate an optimal configuration file based on model characteristics and edge device resources; the configuration file includes resource allocation and concurrency settings; Remotely start the inference service, monitor the service status in real time, and set up a heartbeat monitoring mechanism; Automatically diagnose and recover from failures in starting or abnormal operation of reasoning services; Collect and analyze the operating data of edge device inference services, and provide optimization suggestions based on the analysis results.

9. The method according to claim 1, characterized in that: The API interface and API document of the edge device model are automatically generated, including: Automatically generate API code that complies with RESTful specifications based on the input and output characteristics of the model; Through the API program automatic deployer, the automatically generated API code is encapsulated into a container and automatically pushed to the cloud platform; the API service is started through containerization; The API service is remotely called through the edge Agent program; and the API document of the edge device model is automatically generated.

10. A system for implementing any of the edge AI model automated deployment and update methods as claimed in claims 1-9, characterized in that: The system comprises: The edge service management module is used to obtain edge device information through the user interface; through grouping and label settings, the edge device information that users need to deploy is uniformly managed; The model push module is used to automatically push the parameters of the remote server model to the target edge device based on the time period. The environment configuration module is used to achieve unattended configuration of the edge device model operating environment through intelligent detection, automated script generation and remote execution; The inference service deployment module is used to deploy platform tools on edge devices and intelligently generate configuration files to achieve seamless deployment and optimization of model inference services on edge devices; The API automatic generation module is used to automatically generate API interfaces and API documents for edge device models.

Citation Information

Patent Citations

  • Artificial intelligence one-stop deployment platform

    CN117785220A

  • Method, apparatus, electronic device and readable storage medium for deploying application

    US20220091834A1

  • Method and system for deploying artificial intelligence model in edge device, device, and medium

    WO2024221461A1

Cited By

  • AI quality inspection system based on cloud upgrade

    CN120723779A

  • AI model automatic deployment platform based on containerization technology

    CN120743426A