AI model adaptive compression method and system oriented to maneuvering application scene

By tracking the computing power and memory changes in the maneuvering environment in real time, dynamically calculate the priority weight rate of the compression strategy, and selecting the optimal compression strategy, the problem of the surge in accuracy loss and delay in the maneuvering edge environment in the existing technology is solved, and the dynamic balance between model accuracy and resource consumption is achieved.

CN120179421AActive Publication Date: 2025-06-20ZHONGKE EDGE SMART INFORMATION TECH (SUZHOU) CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510653036.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Existing AI model compression technology is difficult to cope with dynamic resource changes and sudden task loads in maneuverable edge environments, resulting in a surge in model accuracy loss and delay.

Method used

An adaptive compression method for AI model oriented to maneuver application scenarios is proposed. By tracking computing power and memory changes in real time, the priority weight rate of the compression strategy is calculated, resource constraints and compression strategies are dynamically matched, and optimal compression strategies are selected to achieve dynamic balance between model accuracy and resource consumption.

Benefits of technology

In a maneuverable environment with limited resources, model compression is realized that ensures the normal operation of AI recognition tasks and the minimum accuracy loss is minimized, and the problem of delay surge in edge equipment burst tasks is solved, which improves task completion rate and model stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179421A_ABST
    Figure CN120179421A_ABST
Patent Text Reader

Abstract

The invention provides a maneuvering application scene-oriented AI model adaptive compression method and system, and the method comprises the steps: respectively calculating the priority weight rate of each compression strategy in a compression strategy set through tracking calculation power fluctuation and memory change in real time, so as to determine an optimal compression strategy; compared with a traditional fixed threshold value based on static compression, dynamic matching of the resource constraint and the compression strategy is achieved, and the problem of delay surge under the sudden task of the edge device is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of AI models, and in particular, relates to an AI model adaptive compression method and system for mobile application scenarios. Background Art

[0002] As an important optimization method in the field of artificial intelligence, AI model compression technology mainly reduces model complexity through algorithms such as parameter pruning and quantization reconstruction. The current mainstream compression methods are mainly designed for cloud servers or fixed computing nodes, and their optimization goals are focused on improving computing efficiency in static scenarios. In typical mobile edge scenarios such as drone cluster collaboration and on-board real-time decision-making, the dynamic network environment and heterogeneous device resources form unique constraints: on the one hand, there is an essential difference between the volatility of computing power of mobile devices in mobile environments and the stable resource configuration of fixed nodes; on the other hand, the real-time switching requirements of tasks and the parallel operation of multiple models put forward more complex dynamic adaptation requirements for memory bandwidth and computing units.

[0003] Existing compression technologies generally adopt global static compression strategies, which neither consider the dynamic characteristics of edge node resources nor cope with the elastic scaling requirements under sudden task loads, resulting in a large loss of accuracy of the compression model in mobile environments. This technical misalignment causes traditional compression solutions to fail when dealing with scenarios such as drone formation real-time environmental perception, often resulting in a surge in model response delays or a cliff-like drop in recognition accuracy. In response to this core contradiction, it is urgent to establish a dynamic compression mechanism that is deeply adapted to the mobile edge environment, so as to achieve a dynamic balance between model accuracy and resource consumption. Summary of the invention

[0004] The main problem that the present invention solves is how to compress the model in a mobile environment with limited resources while ensuring the normal operation of the AI ​​recognition task and minimizing the loss of accuracy, and provides an AI model adaptive compression method and system for mobile application scenarios.

[0005] In order to solve the above technical problems, the technical solution adopted is: An AI model adaptive compression method for mobile application scenarios includes the following steps: Step 1: Obtain a task set and a compression strategy set, wherein each task in the task set has a task resource requirement, wherein the task resource requirement includes a task CPU resource requirement and a memory resource requirement, and the compression strategy set includes multiple compression strategies; Step 2: Calculate the total CPU resource usage and total memory resource usage at time t based on the task resource requirements of each task; Step 3: Calculate the priority weight rate of each compression strategy in the compression strategy set according to the total CPU resource usage and total memory resource usage at time t; Step 4: Calculate the optimal compression strategy value according to the priority weight rate and precision loss rate of each compression strategy, where the precision loss rate is an attribute of each compression strategy; Step 5: Obtain the optimal compression strategy corresponding to the optimal compression strategy value according to the optimal compression strategy value; Step 6: Compress the model according to the optimal compression strategy and execute the task using the compressed model.

[0006] Furthermore, the method for calculating the total CPU resource usage and total memory resource usage at time t according to the task resource requirements of each task is as follows: ; where is the total CPU resource usage of the mobile environment at time is the total memory resource usage of the mobile environment at time and is the moving average coefficient used to smooth resource fluctuations, is the th real-time CPU computing resource requirement of the is the th memory requirement of the task, and m is the number of tasks in the task set.

[0007] Furthermore, the method for calculating the priority weight rate of each compression strategy in the compression strategy set is as follows: ; where is the th priority weight rate of the compression strategy; and are the total CPU resource usage threshold and total memory resource usage threshold respectively, is the indicator function, and the priority weight rate only takes effect when the total memory resource usage does not exceed the threshold, is the th precision loss rate of the compression strategy.

[0008] Furthermore, the optimal compression strategy selection value: , , where N is the total number of strategies.

[0009] The present invention also provides an AI model adaptive compression system for mobile application scenarios, which is implemented by each step of an AI model adaptive compression method for mobile application scenarios.

[0010] By adopting the above technical solution, the present invention has the following beneficial effects: The present invention provides an AI model adaptive compression method and system for mobile application scenarios. By real-time tracking of computing power fluctuations and memory changes, the priority weight rate of each compression strategy in the compression strategy set is calculated to determine the optimal compression strategy. Compared with the traditional fixed threshold based on static compression, it realizes dynamic matching of resource constraints and compression strategies, and solves the problem of delay surge under sudden tasks of edge devices.

[0011] Through the priority weight rate and optimal selection mechanism, the accuracy loss rate is decoupled from the resource threshold, making the accuracy loss rate independent of the resource threshold. Compared with the traditional method that couples the two and results in forced reduction of precision when exceeding the limit, the present invention allows more strategies to be selected when the memory constraint is activated. The optimal strategy maximizes the set of available strategies under the same resources. Compared with the traditional global compression, it increases the feasible solution space, improves the task completion rate under the same resources, and avoids the defect of the traditional method of drastically reducing the precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flow chart of the system of the present invention. DETAILED DESCRIPTION

[0013] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0014] Figure 1 The present invention shows an AI model adaptive compression method for mobile application scenarios, comprising the following steps: Step 1: Obtain a task set and a compression strategy set, wherein each task in the task set has a task resource requirement, wherein the task resource requirement includes a task CPU resource requirement and a memory resource requirement, and the compression strategy set includes a plurality of compression strategies.

[0015] Step 2: Calculate the total CPU resource usage and total memory resource usage at time t based on the task resource requirements of each task.

[0016] In this embodiment, the method for calculating the total CPU resource usage and total memory resource usage at time t according to the task resource requirements of each task is: ; in, for Total CPU resource usage in the momentary mobile environment, is Total memory resource usage in the momentary mobile environment, and is the moving average coefficient used to smooth resource fluctuations, is the real-time CPU computing resource requirement for the th task, is the th task's memory requirement, and m is the number of tasks in the task set.

[0017] By calculating the total CPU resource usage and total memory resource usage at time t, the computing power fluctuations and memory changes can be tracked in real time.

[0018] Step 3: Calculate the priority weight rate of each compression strategy in the compression strategy set according to the total CPU resource usage and total memory resource usage at time t.

[0019] In this embodiment, the method for calculating the priority weight rate of each compression strategy in the compression strategy set is: ; where is the priority weight rate of the th compression strategy; and are the total CPU resource usage threshold and total memory resource usage threshold respectively, is the indicator function, and the priority weight rate only takes effect when the total memory resource usage does not exceed the threshold, is the accuracy loss rate of the th compression strategy.

[0020] Calculate the priority weight rate according to the accuracy loss rate. When calculating the priority weight rate, decouple the accuracy loss rate from the resource threshold, and only allow accuracy-resource trade-off when the memory does not exceed the threshold, i.e., . While the traditional method couples the two, resulting in forced reduction of accuracy when exceeding the limit. By using the optimal strategy to maximize the available strategy set under the same resources (more strategies can be selected when the memory constraint is activated), compared with the traditional global compression, the feasible solution space is increased, and the task completion rate is improved under the same resources, avoiding the defect of the traditional method's cliff-like drop in accuracy.

[0021] Step 4: Calculate the optimal compression strategy value according to the priority weight rate and accuracy loss rate of each compression strategy. The accuracy loss rate is an attribute of each compression strategy.

[0022] In this embodiment, the optimal compression strategy selection value: , , where N is the total number of strategies.

[0023] Based on the priority weight ratio and real-time computing power and memory , when the memory does not exceed the threshold, calculate the compression strategy values respectively, and take the minimum value as the optimal compression strategy value. Use the compression strategy corresponding to this minimum value as the optimal compression strategy, and when encountering a sudden task switch in an environment with limited mobile resources, use the optimal compression strategy to achieve stable inference of the model with low latency and low energy consumption.

[0024] Step 5: Obtain the optimal compression strategy corresponding to the optimal compression strategy value according to the optimal compression strategy value.

[0025] Step 6: Compress the model according to the optimal compression strategy and execute the task using the compressed model.

[0026] The following is a specific illustration through an example: <1> Case Background When a certain UAV cluster is performing a field search and rescue mission, it needs to process high-resolution infrared images in real time to identify trapped persons. However, limited by on-board computing resources (CPU computing power fluctuates within a range of ±20%, and the memory capacity is only 4GB) and sudden task switches (such as temporarily adding multi-target tracking requirements), traditional static compression models cannot meet the requirements of real-time and low energy consumption.

[0027] <2> Implementation Process a. Dynamic resource modeling and compression strategy triggering.

[0028] Input conditions: ① Initial resource thresholds: CPU computing power threshold ( = 85%), memory threshold ( = 12GB) ② Sudden tasks: The number of tasks increases from 10 to 15, and the additional CPU requirements for the new tasks ( ), memory requirements ( ). Dynamic modeling results: Calculate the real-time resource status through the formula: ① CPU resources: Jumps from 60% to 65.67%, and the moving average coefficient = 0.8; ② Memory resources: Increases from 11GB to 11.85GB, and the moving average coefficient = 0.9, and the memory does not exceed the threshold ; Since the memory does not exceed the threshold ( ), the indicator function Π( ) = 1, and all compression strategies participate in the weight calculation.

[0029] b. Priority weight rate and policy selection policy pool and parameter configuration. In this embodiment, there are five compression policies in the compression policy set, as shown in Table 1.

[0030] Table 1 Five Compression Policies

[0031] Optimal policy calculation: According to the formula: , , where N is the total number of policies.

[0032] Substitute real-time data ( =11.85 / 65.67≈0.180), and calculate the values of each policy: KD: (0.180 + 0.08) / 0.92 = 0.2826; LR: (0.180 + 0.15) / 0.85 = 0.3882; LS: (0.180 + 0.12) / 0.88 = 0.3409; NAS: (0.180 + 0.18) / 0.82 = 0.4390; MP: (0.180 + 0.10) / 0.90 = 0.3111.

[0033] Select the policy corresponding to the minimum value: KD policy (κ = 0.2826).

[0034] <3> Comparison of Execution Effects

[0035] It can be seen that the inference latency of the model is reduced from 420 ms to 285 ms, and the energy consumption, task completion rate, and memory occupancy error are all improved.

[0036] This case verifies the effectiveness of the dynamic compression policy in resource-constrained edge scenarios. The core is to push the compression technology from "offline optimization" to "online decision-making" through real-time modeling and weight rate mechanism, providing a new paradigm for the elastic deployment of AI models.

[0037] The present invention breaks through the fixed threshold limit of traditional static compression by real-time tracking of computing power fluctuations and memory changes, and dynamically scales the policy weights only when the memory does not exceed the threshold, realizing the dynamic matching of resource constraints and compression policies, and solving the problem of sudden increase in latency under burst tasks of edge devices.

[0038] The present invention also provides an AI model adaptive compression system for mobile application scenarios, which is implemented by each step of an AI model adaptive compression method for mobile application scenarios.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI model adaptive compression method for mobile application scenarios, characterized in that: The following steps are involved: Step 1: Obtain a task set and a compression strategy set, wherein each task in the task set has a task resource requirement, wherein the task resource requirement includes a task CPU resource requirement and a memory resource requirement, and the compression strategy set includes multiple compression strategies; Step 2: Calculate the total CPU resource usage and total memory resource usage at time t based on the task resource requirements of each task; Step 3: Calculate the priority weight rate of each compression strategy in the compression strategy set according to the total CPU resource usage and total memory resource usage at time t; Step 4: Calculate the optimal compression strategy value according to the priority weight rate and precision loss rate of each compression strategy, where the precision loss rate is a property of each compression strategy; Step 5: Obtain the optimal compression strategy corresponding to the optimal compression strategy value according to the optimal compression strategy value; Step 6: Compress the model according to the optimal compression strategy and use the compressed model to perform the task.

2. The AI ​​model adaptive compression method for mobile application scenarios according to claim 1 is characterized in that: The method to calculate the total CPU resource usage and total memory resource usage at time t based on the task resource requirements of each task is: ; in, for Total CPU resource usage of the mobile environment at all times, for The total memory resource usage of the mobile environment at all times, and is the sliding average coefficient used to smooth resource fluctuations, For the The real-time CPU computing resource requirements of each task, For the The memory requirement of a task, m is the number of tasks in the task set.

3. The AI ​​model adaptive compression method for mobile application scenarios according to claim 2 is characterized in that: The method for calculating the priority weight ratio of each compression strategy in the compression strategy set is: ; in, For the The priority weight ratio of the compression strategy; and They are the total CPU resource usage threshold and the total memory resource usage threshold. is an indicative function, and the priority weight ratio is only It takes effect when the threshold is not exceeded. For the The precision loss rate of the compression strategy.

4. The AI ​​model adaptive compression method for mobile application scenarios according to claim 3 is characterized in that: Optimal compression strategy selection value: , , N is the total number of strategies.

5. An AI model adaptive compression system for mobile application scenarios, characterized in that: The steps of the AI ​​model adaptive compression method for mobile application scenarios described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Knowledge distillation and quantification technology for power scene edge calculation large model compression

    CN115223049A

  • Mobile terminal depth model compression method with context self-adaption and runtime self-evolution

    CN116029361A

  • Deep learning model compression method and system for micro unmanned aerial vehicle platform

    CN117744742A

  • BTD compressed deep convolutional network construction method and device, and storage medium

    CN118194953A

  • Multi-head multi-target computing power scheduling method for distribution micro-grid

    CN118799117A