A lightweight AI model deployment method for cameras, routers and gateways
Patent Information
- Application Number
- CN202611099098.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-29
AI Technical Summary
高精度模型依赖大规模浮点运算,其算力与存储带宽需求通常超出摄像头等前端低功耗处理器(如ARM Cortex-A系列)的承载能力,导致推理延迟增加、设备功耗急剧上升,甚至无法满足实时性要求
[0021]3、针对单模型推理任务,对计算图进行算子级分析(提取计算类型、参数量、FLOPs、内存访问量等),以最小化端到端延迟为目标,以各设备当前可用算力、内存、网络带宽及延迟上限为约束,采用动态规划或图分割算法将完整模型按层或按算子拆分为N个可独立执行的子模型片段;随后,构建“设备-片段时延矩阵”,以最小化流水线总周期时间为优化目标,结合负载均衡度约束和端到端延迟约束,采用启发式搜索算法为各片段匹配最优边缘设备,实现模型计算负载与异构硬件资源的亲和性调度(如将NPU友好型算子部署至搭载NPU的设备,将CPU密集型算子部署至算力最强的网关);该机制突破了传统静态部署的限制,将“端-边-网”的异构算力整合为统一、协同的推理资源池,使前端摄像头可借助后端网关算力分担深层网络计算,充分发挥算力梯度互补优势,显著提升系统整体吞吐量与资源利用率。
Smart Images

Figure CN122845446A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing and artificial intelligence model deployment technology, and in particular to a lightweight AI model deployment method for cameras, routers and gateways. Background Technology
[0002] With the deepening integration of the Internet of Things (IoT) and artificial intelligence (AI), heterogeneous edge device clusters, represented by smart cameras, edge routers, and smart gateways, have been widely deployed in scenarios such as smart security, smart homes, and vehicle-to-everything (V2X) communication, becoming key infrastructure supporting these scenarios. In such systems, cameras undertake front-end perception tasks such as facial recognition, behavior analysis, and target detection; routers and gateways, leveraging their integrated AI processing units (such as neural network processing units, NPUs), are gradually evolving into edge computing nodes with local inference capabilities, thus forming a distributed computing architecture that coordinates "end-edge-network".
[0003] However, efficiently deploying deep learning models (especially convolutional neural networks with a large number of parameters) in the aforementioned distributed edge environments with significant differences in computing power levels and heterogeneous hardware architectures still faces the following technical challenges, which restrict the improvement of the overall system performance.
[0004] (1) There is a gap between the model's computational characteristics and heterogeneous hardware resources. High-precision models rely on large-scale floating-point operations, and their computing power and storage bandwidth requirements usually exceed the carrying capacity of low-power processors (such as ARM Cortex-A series) in front-end devices such as cameras, resulting in increased inference latency, a sharp increase in device power consumption, and even failure to meet real-time requirements. Existing model optimization methods are relatively crude, usually using a globally uniform pruning rate or quantization accuracy, without differentiated adaptation based on the computing power levels of different devices in the edge device group. The phenomena of model redundancy on high-computing-power devices and model overload on low-computing-power devices coexist.
[0005] (2) The deployment strategy is rigid and fails to effectively utilize the computing power gradient and heterogeneous complementary advantages between devices. The current deployment methods are polarized: First, the complete model is deployed independently on each device without distinguishing the computing power level of the front-end camera and the back-end gateway. Low-end devices are overloaded and high-end devices are idle. Moreover, there is a lack of task collaboration between devices, resulting in low overall inference efficiency. Second, static and uniform model pruning or quantization is performed without considering the different tolerances of different devices for accuracy loss and inference speed. This results in some devices still having model redundancy and other devices experiencing excessive accuracy drops, making it difficult to achieve a globally optimal dynamic balance between "accuracy and efficiency". In addition, the existing solutions cannot adaptively adjust the model version or accuracy level when facing dynamically changing network environments and device loads, further exacerbating the uneven utilization of resources.
[0006] (3) The system-level inference pipeline is missing, and there is a lack of effective task collaboration and computation offloading mechanisms between devices. Currently, each edge device usually completes the entire inference link independently, and there is no vertical partitioning and horizontal offloading of computation tasks between cameras, gateways, and routers. Deep learning models themselves have a hierarchical structure, and the computational requirements of shallow feature extraction (such as backbone network) and deep semantic inference (such as neck network and detection head) are significantly different. The former has a relatively controllable amount of computation, while the latter has higher requirements for matrix multiplication and addition operations. Existing systems lack automated and low-overhead task partitioning and pipeline scheduling mechanisms based on model computation stages (such as backbone network, neck network, and detection head) or operator types, which makes it difficult for the high-performance computing power of the backend gateway to effectively share the load of the frontend camera, thus limiting the improvement of the overall system throughput and resource utilization.
[0007] (4) The model has weak lifecycle management capabilities and insufficient dynamic adaptability and maintainability. The resource status of edge nodes (such as network bandwidth, concurrent load, and chip temperature) is dynamic and time-varying. However, after the model is deployed, the structure, accuracy, and computational load of the existing solution remain fixed. It is impossible to elastically switch the model version or adjust the accuracy level according to the real-time resource status. Overload is likely to occur during peak periods, while computing power is wasted during off-peak periods. When the model is iterated and updated in the cloud, the existing solution relies on the complete distribution of large-granular model files to all edge nodes in the network, which consumes a lot of network resources and affects service continuity. In addition, the existing system generally lacks effective monitoring of the performance of edge models. Problems such as inference accuracy decay, data distribution drift, and accumulation of difficult cases are difficult to detect and repair in a timely manner. The model update cycle is long, and the version management of the cloud and the edge is disconnected, making it difficult to form an efficient continuous iterative optimization closed loop.
[0008] In summary, existing deployment solutions have significant shortcomings in areas such as fine-grained model and hardware adaptation, heterogeneous computing power collaboration mechanisms, dynamic elastic scheduling, and model operation and maintenance flexibility. These shortcomings result in low overall utilization of edge computing power, difficulty in reconciling inference latency and detection accuracy, and poor system maintainability, becoming key bottlenecks restricting the large-scale and sustainable deployment of AI technology in complex IoT scenarios. Therefore, providing a lightweight AI model deployment method for cameras, routers, and gateways to improve inference efficiency, resource utilization, collaborative performance, and operational manageability in heterogeneous edge environments has become an urgent technical problem to be solved in this field. Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a lightweight AI model deployment method for cameras, routers and gateways, so as to improve inference efficiency, resource utilization, collaborative performance and operation and maintenance manageability in heterogeneous edge environments.
[0010] This invention provides a lightweight AI model deployment method for cameras, routers, and gateways, comprising the following steps: Step S1: Edge Device Capability Awareness and Modeling: The capability awareness module deployed on the gateway or router actively detects the hardware specifications and real-time resource status of each edge device in the local area network through the local area network protocol, and builds a device capability profile and dynamic resource pool; the edge device is a camera, router or gateway. Step S2: Lightweight Model Generation and Optimization: Obtain the original AI model and deployment constraints. Based on hardware-aware neural network structure search technology, use the device capability profile as constraints to automatically search for the optimal sub-network structure and generate a basic model library containing multiple compression granularities. Perform mixed-precision quantization and knowledge distillation on the models in the basic model library to output multiple versions of lightweight models adapted to edge devices with different computing power levels. Step S3: Adaptive Model Deployment and Scheduling: Based on the device capability profile and task QoS requirements of each edge device, match and deploy the corresponding version of the lightweight model for each edge device; for single model inference tasks, perform automatic model sharding and pipeline orchestration, split the complete model into independently executable sub-models by layer or by operator, and schedule them to multiple edge devices for collaborative pipeline inference; Step S4: Dynamic management of model runtime: The inference service deployed on each edge device loads the corresponding version of the lightweight model and performs local inference or collaborative inference; continuously monitors the real-time resource status and inference performance indicators of each edge device, and switches the model version of the edge device online when the device load or inference performance triggers the preset threshold conditions. Step S5: Model lifecycle maintenance and iterative update: After the cloud management platform generates a new version of the optimized model, it calculates the differential update package between the new version of the optimized model and the lightweight model currently running on the edge device. The differential update package is pushed to the target edge device via OTA. The target edge device receives the differential update package and completes the hot replacement of the model in memory without interruption of the inference service.
[0011] Furthermore, in step S1, the hardware specifications include the number and frequency of CPU cores, the model and computing power of GPU or NPU, memory capacity, storage space and power consumption limits; the real-time resource status includes CPU utilization, GPU utilization or NPU utilization, memory usage, temperature and power consumption.
[0012] Furthermore, in step S2, generating a basic model library with multiple compression granularities specifically includes: using neural network structure search technology, with the computing power constraints, memory constraints, and latency constraints of edge devices as search space constraints, automatically searching and generating sub-network structures of three granularities: ultra-lightweight, lightweight, and standard, respectively adapted to the computing power levels of cameras, routers, and gateways.
[0013] Furthermore, in step S2, the mixed precision quantization includes: calculating the sensitivity of each network layer in the model in the basic model library to quantization error; selecting INT8 precision mode or FP16 precision mode for each network layer according to the sensitivity and the hardware architecture characteristics of the target edge device; and performing hardware-aware quantization parameter calibration in conjunction with the calibration dataset. The knowledge distillation includes: using the original AI model or a large cloud model as the teacher model, using the sub-network structure generated by neural network structure search as the student model, and using the mid-layer feature maps and output soft labels of the teacher model to guide the training of the student model, so that the student model maintains more than 95% of the inference accuracy of the teacher model under the condition of compression ratio exceeding 70%.
[0014] Furthermore, in step S3, the automatic fragmentation and pipeline orchestration of the execution model specifically includes: Step S31: Dynamic partitioning at the fragment granularity: Obtain the computation graph of the single model to be deployed, perform operator-level analysis on the computation graph, and extract the computation type, parameter quantity, floating-point operation count, memory access volume, and tensor shape information of each operator; based on the operator-level analysis results, with the goal of minimizing end-to-end inference latency, and constrained by the current available computing power, available memory, network transmission bandwidth, and end-to-end latency limit of each edge device, use dynamic programming or graph partitioning algorithms to divide the computation graph of the single model into N independently executable sub-model fragments, where N≥2 and N does not exceed the total number of edge devices participating in collaborative inference; minimize the sum of the output intermediate feature data of each sub-model fragment; Step S32: Device-Segment Affinity Matching and Optimal Deployment Decision: In real time, acquire the device capability profiles of each edge device participating in collaborative inference. The device capability profiles include the CPU computing power margin, NPU dedicated computing power margin, memory margin, communication bandwidth, and current task queue depth of each edge device. Calculate the estimated inference latency and intermediate feature data transmission latency of each sub-model segment on different edge devices to construct a device-segment latency matrix. With minimizing the total pipeline cycle time as the optimization objective and the constraints of load balancing and end-to-end latency of each device as conditions, use a heuristic search algorithm to match the optimal target edge device for each sub-model segment and generate the optimal deployment mapping scheme. Step S33: Lightweight Transmission of Intermediate Feature Data: A feature compression adapter is set at the connection point of each of the sub-model segments. The feature compression adapter performs at least one of the following lightweight processing on the intermediate feature map output by the current sub-model segment at the sending end before transmitting it to the next sub-model segment: (i) adaptive channel pruning or downsampling based on feature importance evaluation; (ii) feature map quantization compression based on the accuracy sensitivity analysis of the next sub-model segment; (iii) dynamic sparsity transmission based on transmission bandwidth awareness. A feature recovery adapter is set at the receiving end, and performs recovery or approximate recovery operation on the received lightweight intermediate feature data before inputting it into the corresponding sub-model segment. Step S34: Asynchronous pipeline parallel execution and bubble elimination: Each sub-model fragment is deployed to the corresponding matched target edge device. A data path based on RDMA or shared memory is established between each target edge device, and inference is executed in an asynchronous pipeline parallel manner. Each target edge device is equipped with a multi-level input buffer queue. When the current device completes the inference of a sub-model fragment and passes the intermediate feature data to the next edge device, it immediately retrieves the next batch of data to be processed from the multi-level input buffer queue for processing, so as to realize the overlapping execution of each stage of the pipeline. When the ratio between the computation latency and the data transmission latency of any stage in the pipeline is detected to deviate from the preset equalization threshold, the batch size of each sub-model fragment is automatically adjusted or the data prefetching depth between each device is dynamically adjusted to eliminate pipeline bubbles. Step S35: Dynamic Adaptive Adjustment and Fault Tolerance Degradation: During the collaborative inference process, continuously monitor the real-time resource status and network status of each target edge device; when any target edge device experiences a sudden increase in load, a decrease in computing power, or network jitter, re-trigger steps S31 to S34 to perform online re-sharding and redeployment, dynamically migrating the affected sub-model fragments to other edge devices with sufficient computing power; when the collaborative inference link is interrupted or the end-to-end latency exceeds the preset alarm threshold, automatically switch to single-device independent inference mode, with the gateway or router with the strongest computing power loading the complete model to perform inference, and returning the degraded inference result to the upper-layer application.
[0015] Furthermore, in step S4, the step of switching the model version of the edge device online when the monitored device load or inference performance triggers a preset threshold condition specifically includes: when the CPU utilization or NPU utilization of the target edge device is detected to continuously exceed a first preset threshold, automatically switching the currently running model version of the target edge device to a low-granularity version with less computational load; when the CPU utilization or NPU utilization of the target edge device is detected to continuously fall below a second preset threshold and the network is idle, automatically switching the currently running model version of the target edge device to a high-granularity version with higher precision. When the network bandwidth is detected to be lower than the third preset threshold or the network latency is detected to be higher than the fourth preset threshold, the collaborative pipeline inference will be automatically switched to single-device independent inference, and the current model version will be switched to a low-granularity version with a smaller amount of intermediate data.
[0016] Furthermore, in step S5, calculating the differential update package between the new optimized model and the lightweight model currently running on the edge device specifically involves: performing a binary differential comparison between the new optimized model and the lightweight model currently running on the edge device, extracting the difference in model weights to generate a differential update package, wherein the amount of data transmitted in the differential update package is less than 10% of the total data volume of the complete new optimized model.
[0017] Furthermore, step S5 also includes: the edge device periodically transmits local inference logs and difficult example samples back to the cloud management platform; the cloud management platform performs incremental learning or retraining of the model based on the transmitted local inference logs and difficult example samples, generates a new version of the optimized model, and forms a continuous iterative closed loop of deployment, monitoring, optimization, and updating.
[0018] Furthermore, before step S3, the following steps are included: when a new edge device is detected to be accessing the local area network, the capability perception module automatically identifies the hardware specifications and real-time resource status of the new edge device, automatically matches or generates an adapted lightweight model based on the device capability profile of the new edge device, and completes the model deployment within a preset time limit.
[0019] Furthermore, in step S3, matching and deploying the corresponding version of the lightweight model for each edge device specifically includes: establishing a model-hardware affinity mapping relationship, prioritizing the deployment of the model layer corresponding to NPU-friendly operators to edge devices equipped with NPUs, and deploying the model layer corresponding to CPU-intensive post-processing operators to the gateway or router with the strongest computing power, thereby achieving affinity scheduling between model computing load and heterogeneous hardware resources. The advantages of this invention are: 1. By constructing device capability profiles and using them as constraints, multi-granularity model versions are generated using hardware-aware neural architecture search and mixed-precision quantization. Combined with model-hardware affinity scheduling, the computational load is accurately matched to the NPU or CPU. At the same time, the complete model is dynamically sharded by operator and matched to the optimal device using a heuristic algorithm based on the latency matrix. With the help of asynchronous pipeline parallelism, bubble elimination mechanism and lightweight transmission of intermediate features, a cross-device collaborative inference and computation offloading path is effectively constructed. This improves the utilization rate of heterogeneous computing power and inference throughput. At the same time, it relies on real-time monitoring of multi-dimensional runtime indicators to trigger elastic version switching to cope with load fluctuations. Combined with OTA hot replacement technology based on differential updates, it ensures that model iteration does not interrupt service. And through cloud incremental learning driven by inference logs and hard example backhaul, a continuous optimization closed loop is formed. It also supports automatic discovery and deployment of new devices. Finally, it comprehensively achieves the dynamic optimal balance between accuracy and efficiency and system-level performance improvement in heterogeneous edge environments from four dimensions: inference efficiency, resource utilization, collaborative efficiency and operation and maintenance manageability.
[0020] 2. By actively detecting hardware specifications and real-time resource status (such as utilization and temperature) such as CPU core count / frequency, NPU model / computing power, and memory capacity through the capability awareness module, a device capability profile is constructed, replacing the traditional static configuration method. Based on this profile, hardware-aware Neural Network Architecture Search (NAS) technology is used to automatically search and generate sub-network structures of different granularities, such as ultra-lightweight, lightweight, and standard versions, to adapt to the differentiated computing power levels of cameras, routers, and gateways. Furthermore, by calculating the sensitivity of each network layer to quantization error and combining it with the hardware architecture characteristics, INT8 or FP16 precision modes are selected for each layer to achieve mixed precision quantization. At the same time, knowledge distillation is introduced, using the original model or a large cloud model as the teacher to guide the training of the NAS-generated sub-networks (student models), maintaining more than 95% of the teacher model accuracy even with a compression ratio of over 70%. This technical solution completely changes the traditional "one-size-fits-all" compression method, realizing "on-demand customization" of model structure and accuracy. It avoids the waste of computing power redundancy on high-end equipment, while ensuring the feasibility and accuracy limit of inference on low-end equipment, thus bridging the adaptation gap between model computing characteristics and heterogeneous hardware from the root.
[0021] 3. For single-model inference tasks, operator-level analysis is performed on the computation graph (extracting computation type, parameter quantity, FLOPs, memory access volume, etc.). With the goal of minimizing end-to-end latency, and constrained by the available computing power, memory, network bandwidth, and latency limits of each device, dynamic programming or graph partitioning algorithms are used to split the complete model into N independently executable sub-model segments, either layer-wise or operator-wise. Subsequently, a "device-segment latency matrix" is constructed. With the optimization objective of minimizing the total pipeline cycle time, combined with load balancing constraints and end-to-end latency constraints, a heuristic search algorithm is used to match the optimal edge device for each segment, achieving affinity scheduling between model computation load and heterogeneous hardware resources (e.g., deploying NPU-friendly operators to devices equipped with NPUs and deploying CPU-intensive operators to the gateway with the strongest computing power). This mechanism breaks through the limitations of traditional static deployment, integrating heterogeneous computing power from "end-edge-network" into a unified and collaborative inference resource pool. This allows front-end cameras to share deep network computation with the computing power of back-end gateways, fully leveraging the complementary advantages of computing power gradients and significantly improving the overall system throughput and resource utilization.
[0022] 4. A feature compression adapter is set at the connection point of the sub-model segments. At the sending end, lightweight processing such as adaptive channel pruning / downsampling based on feature importance assessment, quantization compression based on accuracy sensitivity analysis, and dynamic sparsification transmission based on transmission bandwidth awareness is performed on the intermediate feature map. At the receiving end, a recovery adapter is set up for approximate recovery, thereby significantly reducing the communication overhead of collaborative inference between devices and making cross-device data exchange no longer a bottleneck. A high-speed data path based on RDMA or shared memory is established between each target device, and inference is executed in an asynchronous pipelined parallel manner. Multi-level input buffer queues are set up to achieve overlapping execution of each stage. At the same time, when the ratio of computation latency to data transmission latency in each stage of the pipeline is detected to deviate from the balance threshold, the batch size or data prefetch depth is automatically adjusted to eliminate pipeline bubbles. This solution is the first to build a complete and dynamically optimized collaborative inference pipeline in edge device groups such as cameras, routers, and gateways. It effectively solves the problem of missing vertical task partitioning and horizontal offloading in traditional solutions, enabling the system to maintain high throughput and low latency inference performance in real edge environments with limited communication bandwidth and heterogeneous computing power.
[0023] 5. At the runtime management level, the inference service deployed on each device continuously monitors real-time metrics such as CPU / NPU utilization and memory usage. When the load consistently exceeds the first threshold, it automatically switches to a lower-granularity version with less computation to avoid overload; when the load consistently falls below the second threshold and the network is idle, it automatically switches to a higher-granularity version with higher precision; when the network bandwidth falls below the third threshold or the latency exceeds the fourth threshold, it automatically switches from collaborative pipelined inference to independent inference on a single device and switches to a lower-granularity version with less intermediate data. This dynamic and elastic switching mechanism ensures the robustness and efficiency of the system under resource fluctuation scenarios; at the iterative update level, the cloud... After the edge device generates a new optimized model, it does not distribute it in its entirety. Instead, it calculates a binary differential update package between the current running version and the new model. This package contains less than 10% of the complete model and is pushed via OTA (Over-The-Air) updates, allowing for hot replacement in memory without interruption of inference services. In addition, edge devices periodically send back inference logs and hard example samples, which the cloud uses to perform incremental learning or retraining, forming a continuous iterative closed loop of "deployment-monitoring-optimization-update". This system completely changes the traditional "deployment is fixed" model, enabling the AI model to maintain optimal adaptation to the dynamic environment throughout the long actual operation cycle, and greatly reducing operation and maintenance costs and the risk of update interruption.
[0024] 6. When a new edge device is detected to be connected to the local area network, the capability awareness module automatically identifies its hardware specifications and real-time resource status, automatically matches or generates an adapted lightweight model based on the device capability profile, and completes the deployment within a preset time limit. This mechanism enables the system to have "plug and play" capabilities, which greatly reduces the deployment and management complexity of large-scale edge device groups in dynamic addition and deletion scenarios, and provides good scalability support for the AI system in practical applications such as smart security and smart home where the scale of devices continues to expand.
[0025] 7. The capability awareness module proactively initiates probes using LAN protocols, rather than passively receiving information. This enables the system to capture real-time status of cameras, routers, and gateways across multiple dimensions, including CPU / NPU heterogeneous computing power, memory levels, and power consumption and heat dissipation. By constructing device capability profiles and dynamic resource pools, this solution provides a precise "data foundation" for all subsequent decisions. It ensures that even with fluctuations in computing power (such as frequency throttling due to high temperatures) or sudden load changes, the system still possesses globally optimal scheduling criteria, significantly improving the robustness and timeliness of deployment strategies in real physical environments.
[0026] 8. The single model is split into sub-models at the operator granularity and scheduled to gateways, routers, and cameras for heterogeneous collaborative inference. Its advantages are reflected in three aspects: First, optimal fragmentation is achieved through dynamic programming with the goal of minimizing end-to-end latency, and a device-fragment affinity matrix is established, realizing load balancing of heterogeneous computing power (such as NPU and CPU). Second, a feature compression adapter is introduced between devices to sparsify or prune intermediate feature maps based on transmission bandwidth awareness, greatly alleviating local area network transmission bottlenecks. Third, asynchronous pipelines and bubble elimination mechanisms are adopted to achieve overlapping execution of each stage by dynamically adjusting the batch size. This combination of methods compresses the total cycle time of multi-device collaborative inference to the theoretical lower limit.
[0027] 9. By continuously monitoring CPU / NPU utilization and network bandwidth, the system can automatically switch online between ultra-lightweight, lightweight, and standard models once a preset threshold is triggered. For example, when device overheating and frequency reduction are detected, it seamlessly switches to a low-granularity version to ensure frame rate; when the network is idle and computing power is abundant, it switches back to a high-granularity version to ensure accuracy. This "on-demand" elastic scaling capability ensures that AI services are always in the best energy efficiency state during 24 / 7 operation, greatly extending the lifespan of edge devices and service stability.
[0028] 10. In the face of abnormal situations such as network jitter, sudden increase in device load, or interruption of collaborative links, the system has the ability to actively "self-heal". When an anomaly is triggered, it can automatically perform online re-sharding and dynamically migrate the affected model fragments to spare devices. If the collaborative link is completely interrupted, it will automatically degrade to the single-device independent inference mode, with the gateway with the strongest computing power as a backup. This design ensures that even under extreme conditions of partial node failure, the core AI inference task can still run in a degraded manner, without causing the entire system to collapse, thus meeting the stringent reliability requirements of industrial applications. Attached Figure Description
[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0030] Figure 1 This is a flowchart of a lightweight AI model deployment method for cameras, routers, and gateways according to the present invention. Detailed Implementation
[0031] The overall approach of the technical solution in this application is as follows: A capability awareness module deployed on a gateway or router actively detects the hardware specifications and real-time resource status of heterogeneous edge devices to construct a device capability profile. Then, using this profile as a constraint, lightweight model libraries adapted to different computing power levels are generated using hardware-aware neural architecture search, mixed-precision quantization, and knowledge distillation techniques. Based on this, corresponding model versions are matched and deployed for each device according to the device profile and task QoS requirements. Automatic model sharding and pipeline orchestration are performed for single inference tasks to schedule multi-device collaborative inference. Simultaneously, device load and performance are continuously monitored during runtime, and model versions are switched online when thresholds are triggered to ensure efficiency. Finally, uninterrupted iterative upgrades of the model are achieved through differential update package OTA push and memory hot-swap. Combined with cloud-based incremental learning driven by inference logs and hard example sample feedback, a continuous optimization loop is formed, thereby systematically improving inference efficiency, resource utilization, collaborative performance, and operational manageability in heterogeneous edge environments.
[0032] Please refer to Figure 1 As shown, a preferred embodiment of the lightweight AI model deployment method for cameras, routers, and gateways of the present invention includes the following steps: Step S1: Edge Device Capability Awareness and Modeling: The capability awareness module deployed on the gateway or router actively detects the hardware specifications and real-time resource status of each edge device in the local area network through the local area network protocol, and builds a device capability profile and dynamic resource pool; the edge device is a camera, router or gateway. Preferably, the capability awareness module employs an active probing mechanism combining a lightweight service discovery protocol (mDNS / DNS-SD) and a simple network management protocol (SNMP). Specifically, the capability awareness module periodically sends service discovery multicast messages within the local area network (default probing period is 30 seconds, configurable), requesting each edge device to report its device type, hardware specifications, and current resource status. Upon receiving this request, the lightweight agent program residing on each edge device reads system files (such as / proc / cpuinfo, / sys / class / thermal / thermal_zone* / temp) and calls hardware abstraction layer interfaces (such as the get_utilization() API provided by the NPU driver) to obtain the hardware specifications and real-time resource status, and then encapsulates this information in JSON format before responding to the capability awareness module. For existing camera devices that do not support installing agent programs, the capability awareness module indirectly infers their computing power level and load status by parsing the device information field in their ONVIF protocol and the status code of the RTSP stream. Each time the capability awareness module receives a response, it updates the device capability profile cache in memory using the device's MAC address as a unique identifier, and writes the change record to a local lightweight database (such as SQLite) to support rapid recovery after a system restart.
[0033] Step S2: Lightweight Model Generation and Optimization: Obtain the original AI model and deployment constraints. Based on hardware-aware neural network structure search technology, use the device capability profile as constraints to automatically search for the optimal sub-network structure and generate a basic model library containing multiple compression granularities. Perform mixed-precision quantization and knowledge distillation on the models in the basic model library to output multiple versions of lightweight models adapted to edge devices with different computing power levels. Specifically, the hardware-aware neural network architecture search technology does not rely solely on device capability profiles as static final selection criteria, but rather uses them as core feedback signals throughout the search process. In each search iteration, for each generated candidate sub-network architecture, a hardware performance predictor (which can be pre-trained using actual measurement data from different sub-networks on various edge devices) rapidly estimates its inference latency and peak memory usage on the target device. If the estimation result exceeds the computing power or memory constraints of the device in the device capability profile, the candidate architecture is assigned a very low fitness score, guiding the search algorithm towards lighter architectures. By embedding hardware constraints into the search's reward or loss function, the final searched sub-network architecture ensures that it meets the real-time requirements of the target hardware before deployment, thus achieving true "hardware awareness."
[0034] Preferably, the search space for the neural network structure search adopts a hierarchical design, including three dimensions: (1) Convolution kernel configuration: ordinary convolution, depthwise separable convolution, dilated convolution, or grouped convolution can be selected, and the convolution kernel size is selected from 3×3, 5×5, and 7×7; (2) Channel expansion ratio: in the bottleneck structure, the channel expansion ratio is selected from {1,2,3,4,6} to control the model width; (3) Network depth: in the preset 20-layer backbone search space, the network depth is adjusted by retaining or deleting skip connections. The search algorithm adopts a reinforcement learning-based controller (such as Proximal Policy Optimization) or an evolutionary algorithm (such as Regularized Evolution), with the computing power, memory, and latency constraints in the device capability profile as hard constraints, and the Top-1 accuracy on the validation set as the optimization objective. Specifically, during the search process, for any candidate subnetwork structure, if its estimated inference latency on the target device exceeds the latency limit in the device capability profile, or its peak memory usage exceeds 80% of the available memory, the candidate structure is directly eliminated and does not participate in the subsequent accuracy evaluation, thereby ensuring search efficiency and the deployability of the generated model.
[0035] Step S3: Adaptive Model Deployment and Scheduling: Based on the device capability profile and task QoS requirements of each edge device, match and deploy the corresponding version of the lightweight model for each edge device; for single model inference tasks, perform automatic model sharding and pipeline orchestration, split the complete model into independently executable sub-models by layer or by operator, and schedule them to multiple edge devices for collaborative pipeline inference; Step S4: Dynamic management of model runtime: The inference service deployed on each edge device loads the corresponding version of the lightweight model and performs local inference or collaborative inference; continuously monitors the real-time resource status and inference performance indicators of each edge device, and switches the model version of the edge device online when the device load or inference performance triggers the preset threshold conditions. Step S5: Model lifecycle maintenance and iterative update: After the cloud management platform generates a new version of the optimized model, it calculates the differential update package between the new version of the optimized model and the lightweight model currently running on the edge device. The differential update package is pushed to the target edge device via OTA. The target edge device receives the differential update package and completes the hot replacement of the model in memory without interruption of the inference service.
[0036] In step S1, the hardware specifications include the number and frequency of CPU cores, the model and computing power of GPU or NPU, memory capacity, storage space and power consumption limits; the real-time resource status includes CPU utilization, GPU utilization or NPU utilization, memory usage, temperature and power consumption.
[0037] In step S2, generating a basic model library with multiple compression granularities specifically includes: using neural network structure search technology, with the computing power constraints, memory constraints, and latency constraints of edge devices as search space constraints, automatically searching and generating sub-network structures of three granularities: ultra-lightweight, lightweight, and standard, respectively adapted to the computing power levels of cameras, routers, and gateways.
[0038] In step S2, the mixed precision quantization includes: calculating the sensitivity of each network layer in the model in the basic model library to quantization error; selecting INT8 precision mode or FP16 precision mode for each network layer according to the sensitivity and the hardware architecture characteristics of the target edge device; and performing hardware-aware quantization parameter calibration in conjunction with the calibration dataset. The knowledge distillation includes: using the original AI model or a large cloud model as the teacher model, using the sub-network structure generated by neural network structure search as the student model, and using the mid-layer feature maps and output soft labels of the teacher model to guide the training of the student model, so that the student model maintains more than 95% of the inference accuracy of the teacher model under the condition of compression ratio exceeding 70%.
[0039] The selection of mid-layer feature maps is not random or refers to all hidden layers in general. Preferably, an attention-based feature distillation method is used, which calculates the attention maps of the teacher model and the student model on feature maps of corresponding resolutions and minimizes the mean squared error between them. Specifically, feature maps in the teacher model with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the input image size are selected as guiding targets, and feature maps with corresponding downsampling factors are selected in the student model for alignment. This attention map-based distillation method forces the student to reproduce the teacher's focus in spatial location. Compared with directly fitting all feature maps, its distillation efficiency is higher, and its accuracy recovery effect is particularly significant for extremely small student models with compression ratios exceeding 70%.
[0040] Preferably, the total loss function of knowledge distillation consists of three weighted components: L_total = α·L_hard + β·L_soft + γ·L_feat, where L_hard is the cross-entropy loss between the student model's prediction and the true label, L_soft is the KL divergence loss between the soft labels output by the student model and the teacher model, and L_feat is the mid-layer feature map alignment loss (using mean squared error or mean squared error based on attention mapping). The weight coefficients are set to α=0.1, β=0.3, and γ=0.6 by default. During distillation training, the teacher model has fixed weights, and the student model uses the AdamW optimizer with an initial learning rate of 3×10⁻. 4 A cosine annealing learning rate scheduling strategy was adopted, with a total of 120 training rounds and a batch size of 256. The temperature parameter T was set to 4 to soften the output probability distribution of the teacher model and provide richer dark knowledge. During feature alignment, feature maps from the teacher model with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the input image size were selected and aligned with the corresponding resolution feature maps in the student model.
[0041] In step S3, the automatic fragmentation and pipeline orchestration of the execution model specifically includes: Step S31: Dynamic partitioning at the fragment granularity: Obtain the computation graph of the single model to be deployed, perform operator-level analysis on the computation graph, and extract the computation type, parameter quantity, floating-point operation count, memory access volume, and tensor shape information of each operator; based on the operator-level analysis results, with the goal of minimizing end-to-end inference latency, and constrained by the current available computing power, available memory, network transmission bandwidth, and end-to-end latency limit of each edge device, use dynamic programming or graph partitioning algorithms to divide the computation graph of the single model into N independently executable sub-model fragments, where N≥2 and N does not exceed the total number of edge devices participating in collaborative inference; minimize the sum of the output intermediate feature data of each sub-model fragment; When the graph segmentation algorithm is used to divide the computation graph into N sub-model fragments, it is necessary to ensure that the split points are set at nodes with simple tensor dependency relationships between operators. Preferably, the split point is selected at the tensor transmission position with only a single data dependency, so as to avoid cutting off complex data flow structures such as Residual Connection or cross-layer Concat. When the split point is located after a convolutional layer (Conv) or a fully connected layer (FC) and before an activation layer (ReLU), the activation layer and its subsequent operators need to be divided into the next sub-model fragment together, so as to ensure that the semantics of the input and output tensors of each sub-model fragment are clear, and the fragment can be executed independently without additional data stitching or state synchronization. Meanwhile, for each sub-model fragment after splitting, a data format (e.g., NCHW / NHWC) conversion adaptation layer is automatically generated at its input and output nodes, to ensure that the tensor format output by the previous fragment can be correctly received and processed by the next fragment.
[0042] Preferably, when the graph segmentation algorithm is adopted, the computation graph is modeled as a directed acyclic graph (DAG), where nodes represent operators and directed edges represent tensor dependency relationships. The specific process of the graph segmentation algorithm is as follows: (1) Perform topological sorting on the DAG to obtain the execution order of all operators; (2) Adopt a recursive bisection strategy, calculate the Cut Cost of all splittable edges (that is, edges that do not affect the independent execution of the subgraph after splitting and do not cut off residual connections or cross-layer splicing) in each currently divided subgraph, the cutting cost is defined as the ratio of the data volume (in bytes) of the output tensor of the edge to the sum of the calculation amounts of the operators at both ends of the edge; (3) With the goal of minimizing the total sum of the output intermediate feature data volume between sub-fragments, and taking the balancing constraint that the estimated execution time of each sub-fragment on the target device does not exceed 1.2 times the Pipelining Cycle Time, recursively perform bisection until the number of sub-fragments equals the number N of devices participating in collaborative inference. If the number N of devices participating in collaborative inference is greater than the number of valid splittable edges, the excess devices participate in a Standby mode and do not undertake inference tasks. When the dynamic programming method is adopted, the DAG is divided by layers, and the state dp[i][j] is defined as the minimum end-to-end inference latency when the first i layers are divided into j fragments. During state transition, the starting layer k (k < i) of the j-th fragment is enumerated, the sum of the inference latency of the fragment (k, i] on the matching device and the intermediate data transmission latency is calculated, and finally the segmentation scheme corresponding to dp[L][N] is taken as the optimal division, where L is the total number of layers of the model.
[0043] Step S32: Device-Segment Affinity Matching and Optimal Deployment Decision: In real time, acquire the device capability profiles of each edge device participating in collaborative inference. The device capability profiles include the CPU computing power margin, NPU dedicated computing power margin, memory margin, communication bandwidth, and current task queue depth of each edge device. Calculate the estimated inference latency and intermediate feature data transmission latency of each sub-model segment on different edge devices to construct a device-segment latency matrix. With minimizing the total pipeline cycle time as the optimization objective and the constraints of load balancing and end-to-end latency of each device as conditions, use a heuristic search algorithm to match the optimal target edge device for each sub-model segment and generate the optimal deployment mapping scheme. Preferably, the heuristic search algorithm adopts an improved version of the Hungarian algorithm - a minimum cycle time allocation algorithm with load balancing constraints. The specific process is as follows: (1) Based on the device-segment delay matrix, the optimization objective is to minimize the total pipeline cycle time (i.e., the maximum completion time among all devices); (2) The bottleneck-aware Earliest Finish Time (BA-EFT) strategy is adopted, and the sub-model segments are allocated sequentially according to the execution order in the computation graph. Each time, the current sub-model segment is allocated to the edge device that minimizes its estimated completion time and whose total load after allocation does not exceed 90% of its available computing power; (3) If the load balancing degree of the initial allocation scheme (defined as the load ratio of the busiest device to the idlest device) exceeds the preset threshold (default 1.5), local search optimization is triggered: under the premise of keeping the total cycle time from increasing, one or more sub-model segments on the device with the highest load are migrated to the available device with the lowest load, and the overall delay is recalculated. The iteration continues until the load balancing degree meets the threshold requirement or reaches the maximum number of iterations (default 100 times). For scenarios with a large search space (number of devices × number of sub-model fragments > 100), a genetic algorithm is used to replace BA-EFT. The population size is 50, the crossover probability is 0.8, the mutation probability is 0.1, the number of generations is 200, the fitness function is defined as the reciprocal of the total pipeline cycle time, and load balancing is introduced as a penalty term.
[0044] Step S33: Lightweight Transmission of Intermediate Feature Data: A feature compression adapter is set at the connection point of each of the sub-model segments. The feature compression adapter performs at least one of the following lightweight processing on the intermediate feature map output by the current sub-model segment at the sending end before transmitting it to the next sub-model segment: (i) adaptive channel pruning or downsampling based on feature importance evaluation; (ii) feature map quantization compression based on the accuracy sensitivity analysis of the next sub-model segment; (iii) dynamic sparsity transmission based on transmission bandwidth awareness. A feature recovery adapter is set at the receiving end, and performs recovery or approximate recovery operation on the received lightweight intermediate feature data before inputting it into the corresponding sub-model segment. Preferably, the adaptive channel pruning based on feature importance assessment specifically involves: calculating the mean and variance of activation values for each channel in the intermediate feature map; determining channels with variances greater than 50% of the global mean variance as important channels and retaining these channels; pruning the remaining channels; and dynamically adjusting the pruning ratio between 20% and 60%, specifically determined based on the ratio of the currently available transmission bandwidth to the original feature map data volume: when bandwidth is ample, the pruning ratio tends towards the lower limit (20%), and when bandwidth is tight, it tends towards the upper limit (60%). The quantization compression based on precision sensitivity analysis specifically involves: performing sensitivity analysis on each feature map channel in advance on the receiving terminal model segment—quantizing each channel sequentially to INT8 and measuring the end-to-end precision decrease; quantizing channels with a precision decrease of less than 1% to INT8, quantizing channels with a precision decrease between 1% and 3% to FP16, and maintaining the original FP32 precision for channels with a precision decrease exceeding 3%. The dynamic sparsity transmission based on transmission bandwidth awareness specifically involves: dynamically setting a sparsity threshold based on the currently measured available network bandwidth—feature values with absolute values less than the threshold are set to zero and transmitted in sparse tensor format (COO or CSR format); the threshold setting rule is to match the amount of data after sparsification with the maximum amount of data that can be transmitted within the expected transmission time window (default 50ms) under the current available bandwidth, while ensuring that the decrease in inference accuracy after sparsification does not exceed 2%. The feature recovery adapter at the receiving end restores the original number of channels to the feature map after channel clipping with zero padding; performs dequantization on the quantized and compressed feature map; and reconstructs the sparsely transmitted feature map into a dense tensor format before inputting it into the next sub-model segment.
[0045] Step S34: Asynchronous pipeline parallel execution and bubble elimination: Each sub-model fragment is deployed to the corresponding matched target edge device. A data path based on RDMA or shared memory is established between each target edge device, and inference is executed in an asynchronous pipeline parallel manner. Each target edge device is equipped with a multi-level input buffer queue. When the current device completes the inference of a sub-model fragment and passes the intermediate feature data to the next edge device, it immediately retrieves the next batch of data to be processed from the multi-level input buffer queue for processing, so as to realize the overlapping execution of each stage of the pipeline. When the ratio between the computation latency and the data transmission latency of any stage in the pipeline is detected to deviate from the preset equalization threshold, the batch size of each sub-model fragment is automatically adjusted or the data prefetching depth between each device is dynamically adjusted to eliminate pipeline bubbles. Data paths based on RDMA or shared memory are not limited to using the RDMA protocol between all devices. More flexibly, the establishment method of the data path adaptively selects based on the hardware capabilities of the devices participating in collaborative inference: if both communicating devices have RDMA capabilities (such as a RoCEv2-supporting smart gateway and high-end router), an RDMA data path is established first to reduce CPU overhead and transmission latency; if either device lacks RDMA capabilities, it automatically degrades to high-performance socket communication based on the TCP / IP protocol, and the aforementioned feature compression adapter is enabled for extreme compression to ensure acceptable transmission performance in a standard Ethernet environment. Furthermore, for communication between sub-model fragments of different processes or containers within the same gateway or router, shared memory is used directly to achieve zero-copy data exchange.
[0046] Preferably, the strategy for automatically adjusting the batch size is as follows: Assume the pipeline has N stages (i.e., N sub-model segments), the computation latency of the i-th stage is C_i, the data transmission latency is T_i (where T_N=0, meaning the last stage does not require transmission), the total pipeline cycle time is max_i(C_i+T_i), and the bubble rate is (Σ(C_i+T_i)-max_i(C_i+T_i)) / Σ(C_i+T_i). Adjustment is triggered when the bubble rate exceeds 15%. The adjustment strategy is as follows: If the computation latency is dominant in a stage (C_i / (C_i+T_i)>0.8), the batch size of that stage will be reduced by 20% to lower the computation latency; if the data transmission latency is dominant in a stage (T_i / (C_i+T_i)>0.4), the compression ratio of the output feature compression adapter in that stage will be increased (with a maximum compression ratio of 70%) to reduce the amount of data transmitted. After adjustment, the bubble rate will be re-evaluated. If it still does not meet the target, the data prefetch depth between devices will be further adjusted—gradually increasing the prefetch depth from the default value of 2 to 4 or 8, so that the upstream stage prepares more batches of intermediate feature data in advance to mask the startup delay of the downstream stage. The above adjustments are evaluated every 30 seconds, and the adjustment range adopts an annealing strategy, with a single adjustment range not exceeding 25% of the current value to avoid system oscillation.
[0047] Step S35: Dynamic Adaptive Adjustment and Fault Tolerance Degradation: During the collaborative inference process, continuously monitor the real-time resource status and network status of each target edge device; when any target edge device experiences a sudden increase in load, a decrease in computing power, or network jitter, re-trigger steps S31 to S34 to perform online re-sharding and redeployment, dynamically migrating the affected sub-model fragments to other edge devices with sufficient computing power; when the collaborative inference link is interrupted or the end-to-end latency exceeds the preset alarm threshold, automatically switch to single-device independent inference mode, with the gateway or router with the strongest computing power loading the complete model to perform inference, and returning the degraded inference result to the upper-layer application.
[0048] Preferably, the specific determination condition for re-triggering is triggered when any of the following conditions are met: (i) the CPU utilization or NPU utilization of any target edge device exceeds 95% for 15 seconds; (ii) the available memory of any target edge device is less than 15% of its total memory; (iii) the chip temperature of any target edge device exceeds its thermal control threshold (default 85°C, subject to the device datasheet); (iv) the network round-trip time (RTT) increases by more than 100% relative to the baseline value (the average RTT when initially establishing collaborative inference) within 5 consecutive measurement cycles (1 second per cycle); (v) the actual inference latency of any sub-model segment exceeds 1.5 times the estimated inference latency in step S32. After re-triggering, the migration of the affected sub-model fragments adopts the "Make-Before-Break" strategy: first, the corresponding sub-model fragments are loaded and warmed up on the target surplus edge device. After the new instance is ready, the data stream is atomically switched from the original device to the target device. After the switch is completed, the sub-model fragment resources on the original device are released. The entire migration process is transparent to the upper layer application, and the switching time is controlled within 100ms.
[0049] In step S4, the step of switching the model version of the edge device online when the monitored device load or inference performance triggers a preset threshold condition specifically includes: when the CPU utilization or NPU utilization of the target edge device continuously exceeds a first preset threshold, automatically switching the currently running model version of the target edge device to a low-granularity version with less computational load; when the CPU utilization or NPU utilization of the target edge device continuously falls below a second preset threshold and the network is idle, automatically switching the currently running model version of the target edge device to a high-granularity version with higher precision. When the network bandwidth is detected to be lower than the third preset threshold or the network latency is detected to be higher than the fourth preset threshold, the collaborative pipeline inference will be automatically switched to single-device independent inference, and the current model version will be switched to a low-granularity version with a smaller amount of intermediate data.
[0050] Online model version switching employs a double-buffer mechanism to ensure service continuity. Specifically, when a version switching condition is triggered, the edge device simultaneously loads the currently running version (old version) and the target version to switch to (new version) into its memory. The new version model completes initialization and warm-up in the background (e.g., performing several empty inferences to load the cache), during which all inference requests are still handled by the old version. Once the new version has warmed up, the request pointer for the inference service is atomically switched from the old version to the new version, and then the memory resources occupied by the old version are released. For inference requests that are being executed at the time of switching, the old version continues processing the current batch and then exits, with the new version taking over from the next inference batch. The entire process is completely transparent to the upper-layer application, ensuring zero request loss and zero service interruption.
[0051] In step S5, calculating the differential update package between the new optimized model and the lightweight model currently running on the edge device specifically involves: performing a binary differential comparison between the new optimized model and the lightweight model currently running on the edge device, extracting the difference in model weights to generate a differential update package, wherein the amount of data transmitted in the differential update package is less than 10% of the total data of the complete new optimized model.
[0052] The new optimized model is compared with the lightweight model currently running on the edge device using binary differential comparison, which is applicable only if the two versions of the model have identical structures. If the new optimized model generated by the cloud management platform involves changes to the network structure (e.g., adding or deleting layers, replacing operators), the cloud management platform first automatically identifies the structural mapping relationship between the old and new models using a structural mapping algorithm before calculating the differential update package. For layers with the same structure, incremental weights are generated using the aforementioned binary differential method; for newly added or deleted layers, their full weights or zero-set information are packaged into the differential update package, along with a structural change descriptor. After receiving the differential update package, the edge device first reconstructs the computation graph of the new model in memory based on the structural change descriptor, then applies the binary incremental weights, and finally completes the hot replacement of the model. This method ensures lightweight OTA updates in most scenarios while also being compatible with the needs of structural evolution.
[0053] Step S5 further includes: the edge device periodically transmits local inference logs and difficult example samples back to the cloud management platform; the cloud management platform performs incremental learning or retraining of the model based on the transmitted local inference logs and difficult example samples, generates a new version of the optimized model, and forms a continuous iterative closed loop of deployment, monitoring, optimization and updating.
[0054] Before step S3, the method further includes: when a new edge device is detected to be connected to the local area network, the capability perception module automatically identifies the hardware specifications and real-time resource status of the new edge device, automatically matches or generates an adapted lightweight model based on the device capability profile of the new edge device, and completes the model deployment within a preset time limit.
[0055] Preferably, when a new edge device is detected accessing the local area network, the automatic deployment process is as follows: The capability awareness module continuously monitors DHCP lease update events or mDNS service announcement messages within the local area network. Upon detecting the IP address and MAC address of the new device, it immediately sends a capability detection request to the device. When the lightweight agent program on the new device starts for the first time, it proactively sends a device registration message to the gateway's capability awareness module. This message contains static information such as device type and hardware specifications. After receiving the response, the capability awareness module completes the construction of the device capability profile within 100ms and matches the corresponding version from the generated lightweight model library according to the computing power level in the profile (divided into three levels: L0 ultra-lightweight, L1 lightweight, and L2 standard): L0 matches the ultra-lightweight model (suitable for low-power devices such as cameras), L1 matches the lightweight model (suitable for routers), and L2 matches the standard model (suitable for gateways). If there is no corresponding adapted version in the model library, the cloud management platform completes the generation and distribution of the model for that computing power level within 5 minutes. Model delivery uses a chunked transmission method, with each chunk being 512KB in size, and supports resuming interrupted downloads. After deployment, the capability awareness module adds the new device to the dynamic resource pool, updates the global device list, and synchronously updates the available device list for each collaborative inference task, ensuring that the new device can be included in the scheduling scope in the next inference batch. The entire process, from device access to model deployment completion, takes no more than 3 minutes by default, ensuring that new devices can be quickly deployed.
[0056] In step S3, matching and deploying the corresponding version of the lightweight model for each edge device specifically includes: establishing a model-hardware affinity mapping relationship, prioritizing the deployment of the model layer corresponding to NPU-friendly operators to edge devices equipped with NPUs, and deploying the model layer corresponding to CPU-intensive post-processing operators to the gateway or router with the strongest computing power, thereby achieving affinity scheduling between model computing load and heterogeneous hardware resources.
[0057] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A lightweight AI model deployment method for cameras, routers, and gateways, characterized in that, Includes the following steps: Step S1: Edge Device Capability Awareness and Modeling: The capability awareness module deployed on the gateway or router actively detects the hardware specifications and real-time resource status of each edge device in the local area network through the local area network protocol, and builds a device capability profile and dynamic resource pool; the edge device is a camera, router or gateway. Step S2: Lightweight Model Generation and Optimization: Obtain the original AI model and deployment constraints. Based on hardware-aware neural network structure search technology, use the device capability profile as constraints to automatically search for the optimal sub-network structure and generate a basic model library containing multiple compression granularities. Perform mixed-precision quantization and knowledge distillation on the models in the basic model library to output multiple versions of lightweight models adapted to edge devices with different computing power levels. Step S3: Adaptive Model Deployment and Scheduling: Based on the device capability profile and task QoS requirements of each edge device, match and deploy the corresponding version of the lightweight model for each edge device; for single model inference tasks, perform automatic model sharding and pipeline orchestration, split the complete model into independently executable sub-models by layer or by operator, and schedule them to multiple edge devices for collaborative pipeline inference; Step S4: Dynamic Management of Model Runtime: The inference service deployed on each edge device loads the corresponding version of the lightweight model and performs local inference or collaborative inference; Continuously monitor the real-time resource status and inference performance indicators of each edge device. When the device load or inference performance triggers a preset threshold condition, switch the model version of the edge device online. Step S5: Model lifecycle maintenance and iterative update: After the cloud management platform generates a new version of the optimized model, it calculates the differential update package between the new version of the optimized model and the lightweight model currently running on the edge device. The differential update package is pushed to the target edge device via OTA. The target edge device receives the differential update package and completes the hot replacement of the model in memory without interruption of the inference service.
2. The lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S1, the hardware specifications include the number and frequency of CPU cores, the model and computing power of GPU or NPU, memory capacity, storage space and power consumption limits; the real-time resource status includes CPU utilization, GPU utilization or NPU utilization, memory usage, temperature and power consumption.
3. The lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S2, generating a basic model library with multiple compression granularities specifically includes: using neural network structure search technology, with the computing power constraints, memory constraints, and latency constraints of edge devices as search space constraints, automatically searching and generating sub-network structures of three granularities: ultra-lightweight, lightweight, and standard, respectively adapted to the computing power levels of cameras, routers, and gateways.
4. The lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S2, the mixed precision quantization includes: calculating the sensitivity of each network layer in the model in the basic model library to quantization error; selecting INT8 precision mode or FP16 precision mode for each network layer according to the sensitivity and the hardware architecture characteristics of the target edge device; and performing hardware-aware quantization parameter calibration in conjunction with the calibration dataset. The knowledge distillation includes: using the original AI model or a large cloud model as the teacher model, using the sub-network structure generated by neural network structure search as the student model, and using the mid-layer feature maps and output soft labels of the teacher model to guide the training of the student model, so that the student model maintains more than 95% of the inference accuracy of the teacher model under the condition of compression ratio exceeding 70%.
5. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S3, the automatic fragmentation and pipeline orchestration of the execution model specifically includes: Step S31: Dynamic partitioning at the fragment granularity: Obtain the computation graph of the single model to be deployed, perform operator-level analysis on the computation graph, and extract the computation type, parameter quantity, floating-point operation count, memory access volume, and tensor shape information of each operator; based on the operator-level analysis results, with the goal of minimizing end-to-end inference latency, and constrained by the current available computing power, available memory, network transmission bandwidth, and end-to-end latency limit of each edge device, use dynamic programming or graph partitioning algorithms to divide the computation graph of the single model into N independently executable sub-model fragments, where N≥2 and N does not exceed the total number of edge devices participating in collaborative inference; minimize the sum of the output intermediate feature data of each sub-model fragment; Step S32: Device-Segment Affinity Matching and Optimal Deployment Decision: In real time, acquire the device capability profiles of each edge device participating in collaborative inference. The device capability profiles include the CPU computing power margin, NPU dedicated computing power margin, memory margin, communication bandwidth, and current task queue depth of each edge device. Calculate the estimated inference latency and intermediate feature data transmission latency of each sub-model segment on different edge devices to construct a device-segment latency matrix. With minimizing the total pipeline cycle time as the optimization objective and the constraints of load balancing and end-to-end latency of each device as conditions, use a heuristic search algorithm to match the optimal target edge device for each sub-model segment and generate the optimal deployment mapping scheme. Step S33: Lightweight Transmission of Intermediate Feature Data: A feature compression adapter is set at the connection point of each of the sub-model segments. The feature compression adapter performs at least one of the following lightweight processing on the intermediate feature map output by the current sub-model segment at the sending end before transmitting it to the next sub-model segment: (i) adaptive channel pruning or downsampling based on feature importance evaluation; (ii) feature map quantization compression based on the accuracy sensitivity analysis of the next sub-model segment; (iii) dynamic sparsity transmission based on transmission bandwidth awareness. A feature recovery adapter is set at the receiving end, and performs recovery or approximate recovery operation on the received lightweight intermediate feature data before inputting it into the corresponding sub-model segment. Step S34: Asynchronous pipeline parallel execution and bubble elimination: Each sub-model fragment is deployed to the corresponding matched target edge device. A data path based on RDMA or shared memory is established between each target edge device, and inference is executed in an asynchronous pipeline parallel manner. Each target edge device is equipped with a multi-level input buffer queue. When the current device completes the inference of a sub-model fragment and passes the intermediate feature data to the next edge device, it immediately retrieves the next batch of data to be processed from the multi-level input buffer queue for processing, so as to realize the overlapping execution of each stage of the pipeline. When the ratio between the computation latency and the data transmission latency of any stage in the pipeline is detected to deviate from the preset equalization threshold, the batch size of each sub-model fragment is automatically adjusted or the data prefetching depth between each device is dynamically adjusted to eliminate pipeline bubbles. Step S35: Dynamic Adaptive Adjustment and Fault Tolerance Degradation: During the collaborative inference process, continuously monitor the real-time resource status and network status of each target edge device; when any target edge device experiences a sudden increase in load, a decrease in computing power, or network jitter, re-trigger steps S31 to S34 to perform online re-sharding and redeployment, dynamically migrating the affected sub-model fragments to other edge devices with sufficient computing power; when the collaborative inference link is interrupted or the end-to-end latency exceeds the preset alarm threshold, automatically switch to single-device independent inference mode, with the gateway or router with the strongest computing power loading the complete model to perform inference, and returning the degraded inference result to the upper-layer application.
6. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S4, the step of switching the model version of the edge device online when the monitored device load or inference performance triggers a preset threshold condition specifically includes: when the CPU utilization or NPU utilization of the target edge device continuously exceeds a first preset threshold, automatically switching the currently running model version of the target edge device to a low-granularity version with less computational load; when the CPU utilization or NPU utilization of the target edge device continuously falls below a second preset threshold and the network is idle, automatically switching the currently running model version of the target edge device to a high-granularity version with higher precision. When the network bandwidth is detected to be lower than the third preset threshold or the network latency is detected to be higher than the fourth preset threshold, the collaborative pipeline inference will be automatically switched to single-device independent inference, and the current model version will be switched to a low-granularity version with a smaller amount of intermediate data.
7. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S5, calculating the differential update package between the new optimized model and the lightweight model currently running on the edge device specifically involves: performing a binary differential comparison between the new optimized model and the lightweight model currently running on the edge device, extracting the difference in model weights to generate a differential update package, wherein the amount of data transmitted in the differential update package is less than 10% of the total data of the complete new optimized model.
8. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, Step S5 further includes: the edge device periodically transmits local inference logs and difficult example samples back to the cloud management platform; the cloud management platform performs incremental learning or retraining of the model based on the transmitted local inference logs and difficult example samples, generates a new version of the optimized model, and forms a continuous iterative closed loop of deployment, monitoring, optimization and updating.
9. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, Before step S3, the method further includes: when a new edge device is detected to be connected to the local area network, the capability perception module automatically identifies the hardware specifications and real-time resource status of the new edge device, automatically matches or generates an adapted lightweight model based on the device capability profile of the new edge device, and completes the model deployment within a preset time limit.
10. A lightweight AI model deployment method for cameras, routers, and gateways as described in claim 1, characterized in that, In step S3, matching and deploying the corresponding version of the lightweight model for each edge device specifically includes: establishing a model-hardware affinity mapping relationship, prioritizing the deployment of the model layer corresponding to NPU-friendly operators to edge devices equipped with NPUs, and deploying the model layer corresponding to CPU-intensive post-processing operators to the gateway or router with the strongest computing power, thereby achieving affinity scheduling between model computing load and heterogeneous hardware resources.