An AI model distillation migration edge deployment method across a chip architecture
Patent Information
- Application Number
- CN202511215140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-08-28
AI Technical Summary
相关技术手段包括适配特定芯片的格式转换工具,可将模型转换为边缘设备兼容格式,但此类工具多局限于单一芯片架构;剪枝、量化、知识蒸馏等轻量化优化技术,通过去除冗余参数、降低数据精度、压缩模型规模,以适配边缘设备的算力和内存;还有针对特定场景开发的专用模型,不过这类模型通用性较差,跨场景迁移时需重新训练
(1)本发明通过筛选适配场景的模型及构建反映边缘设备实际运行情况的校准数据集,使模型从训练阶段就贴合边缘部署场景,再经蒸馏统一为YOLO结构,结合剪枝、量化等优化,显著提升模型在不同边缘设备上的部署效率与场景适配能力。
Smart Images

Figure CN121029189B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI model technology, specifically to a method for distilling and migrating AI models across chip architectures for edge deployment. Background Technology
[0002] Currently, edge deployment of AI models in product development and technology application mainly focuses on two directions: model lightweighting and cross-scenario adaptation. Related technical means include format conversion tools adapted to specific chips, which can convert models to edge-device-compatible formats, but these tools are mostly limited to a single chip architecture; lightweight optimization techniques such as pruning, quantization, and knowledge distillation, which remove redundant parameters, reduce data precision, and compress model size to adapt to the computing power and memory of edge devices; and dedicated models developed for specific scenarios, but these models have poor versatility and require retraining when migrating across scenarios. However, existing methods have many shortcomings and cannot meet practical needs. First, poor cross-architecture adaptability. Different edge chips (such as NPUs and GPUs) differ significantly in computing power architecture, supported operators, and optimization strategies, while heterogeneous model structures are diverse. The update speed of edge inference frameworks cannot keep up with the iteration of model architectures, resulting in targeted adjustments required when migrating models across chips, high adaptation costs, and difficulties in feature alignment and knowledge representation consistency of heterogeneous models. Second, excessive resource consumption. Deep learning models have a large number of parameters, which consumes the limited memory of edge devices, and their inference computation is complex. Due to the size and power consumption limitations of edge devices, their computing power is limited, resulting in high model inference latency, making it difficult to meet real-time requirements (such as ≤500ms response time for sudden event detection). Traditional compression methods are prone to loss of model accuracy due to excessive parameter pruning. Thirdly, their transfer and generalization capabilities are weak. Models for specific scenarios have poor generality and require retraining when transferring to other scenarios. Furthermore, in dynamic environments, due to the lack of a unified knowledge representation method and imperfect prior knowledge embedding mechanism, the model cannot effectively capture nonlinear output differences, making it difficult to adapt and resulting in insufficient generalization ability. Summary of the Invention
[0003] The purpose of this invention is to overcome the above-mentioned shortcomings of current edge deployment of AI models and to provide a method for distillation migration of AI models across chip architectures for edge deployment.
[0004] The objective of this invention is achieved through the following technical solution: a method for edge deployment of AI model distillation migration across chip architectures, comprising the following steps: S1. Based on the current actual scenario requirements, select models suitable for the current application scenario; S2. Based on the actual scenario requirements, construct a calibration dataset that can reflect the actual operating conditions faced by edge devices; S3. First, unify the heterogeneous models into the YOLO model structure by distillation. S4. Prune the unified YOLO model to reduce the parameter size; S5. Quantize the pruned YOLO model to improve edge deployment efficiency.
[0005] Furthermore, the distillation method described in step S3 includes the following steps: S31. Select the pre-trained heterogeneous model as the teacher model, freeze its parameters and extract the key feature layer, and unify the feature scale through feature standardization. S32. Construct a lightweight YOLO architecture as the student model, retain the core YOLO detection head, and adjust the backbone depth and width to match the feature dimensions of the teacher model. S33. Design a multi-objective loss function: (L_{total} =λ_1L_{cls}+λ_2L_{feat}+λ_3L_{arch}); where L_{cls} is the classification loss, aligning the class probability distribution of the teacher-student model; L_{feat} is the feature loss, aligning the intermediate layer features; L_{arch} is the architecture adaptation loss, λ_1=0.4, λ_2=0.3, λ_3=0.3; S34. Use Kendall's correlation coefficient to measure the consistency of feature distribution between the above teacher model and student model, and ensure the consistency of feature distribution by maximizing Kendall's correlation coefficient. S35. Conduct distillation training; S36. Test the performance of the student model on the validation set and compare it with the teacher model. The target detection mAP loss should be ≤2% and the number of parameters should be reduced by ≥50%.
[0006] Step S4, "pruning the unified YOLO model and compressing the parameter scale," specifically includes the following steps: S41. Select validation set samples, calculate the pruning sensitivity of each convolutional layer in YOLO, and mark low-sensitivity layers as priority pruning targets; wherein, the formula for "calculating the pruning sensitivity of each convolutional layer in YOLO" is: S_l = Delta mAP / DeltaParam_l, where Delta mAP is the accuracy loss after pruning 50% of the channels of the layer, and Delta Param_l is the parameter reduction ratio; the low-sensitivity layer is defined as S_l < 0.01.
[0007] S42. For the marked low-sensitivity layers, calculate the importance of each channel using the L2 norm, and adjust the score by combining prior knowledge of the specific scenario; wherein, the channel importance measurement formula for “calculating the importance of each channel using the L2 norm” is S_c=frac{1}{N} sum_{i=1}^N|W_{c,i}|_2, where W_{c,i} is the convolution kernel parameter of the c-th channel, and N is the number of convolution kernels; S43. Set a basic threshold and use dynamically adjusted redundant channels for different levels to ensure the integrity of the model structure after pruning. S44. Fine-tune the pruning model for 20-30 rounds using the Adam optimizer, freeze the detection head parameters and only update the backbone, and restore accuracy through fine-tuning methods so that the mAP loss after fine-tuning is ≤1%; S45. Test the number of parameters, computational load, and inference speed of the pruned model to ensure it meets the memory constraints of edge devices.
[0008] Step S5, "Quantizing the pruned YOLO model," includes the following steps: S51. Based on the requirements for model accuracy and inference speed in the task book, a hybrid strategy of post-training quantization and quantization-aware training is adopted. First, post-training quantization is used to perform preliminary quantization of the model, and then quantization-aware training is used to fine-tune the model. S52. Use KL divergence to calibrate the weight quantization threshold, and determine the optimal quantization interval by minimizing D_{KL}, where D_{KL}(P|Q) = sum P(x) logfrac{P(x)}{Q(x)}, where P is the original weight distribution and Q is the quantized distribution; S53. Freeze the weight quantization parameters, fine-tune the quantization range of activation values, and train for 10-15 rounds using the learning rate. Optimize the quantization accuracy by increasing the proportion of the targeted target in the training samples. S54. For different edge devices, optimize according to their instruction set architecture, use tools such as TensorRT to perform operator fusion, split operators such as 3×3 convolution into forms that are more in line with the computing characteristics of the device, and optimize the memory layout of the model.
[0009] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) This invention enables the model to fit the edge deployment scenario from the training stage by screening the model that is suitable for the scenario and constructing a calibration dataset that reflects the actual operation of the edge device. The model is then distilled and unified into the YOLO structure. Combined with optimizations such as pruning and quantization, the deployment efficiency and scenario adaptability of the model on different edge devices are significantly improved.
[0010] (2) When the heterogeneous model is unified into the YOLO structure by distillation, the present invention designs a multi-objective loss function and combines it with Kendall correlation coefficient to ensure the consistency of feature distribution. While the number of parameters is greatly reduced, the target detection mAP loss is guaranteed to be ≤2%, thus realizing efficient transfer of knowledge from heterogeneous models and effective maintenance of model performance.
[0011] (3) In the pruning process, the present invention accurately locates the low-sensitivity layer by calculating the pruning sensitivity of the convolutional layer, evaluates the importance of the channel by combining the L2 norm and the prior knowledge of the scene, and dynamically adjusts the redundant channels. After fine-tuning, the mAP loss is ≤1%. While greatly compressing the parameter scale, it balances the model performance and the resource consumption of edge devices such as memory.
[0012] (4) The quantization process of the present invention adopts a hybrid strategy, uses KL divergence to calibrate the quantization threshold to determine the optimal quantization range, combines targeted fine-tuning to optimize quantization accuracy, and optimizes the instruction set architecture of different edge devices to improve the model inference speed and effectively control the accuracy loss, thus taking into account both speed and accuracy requirements.
[0013] (5) The present invention revolves around the needs of edge deployment, from model selection, knowledge distillation, pruning to quantization. The optimized model can adapt to the hardware characteristics of different edge devices and has strong applicability and scalability, providing strong support for the widespread application of AI models in edge scenarios. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0015] Figure 2 This is a schematic diagram of the distillation process of the present invention.
[0016] Figure 3 This is a schematic diagram of the pruning process of the present invention.
[0017] Figure 4 This is a schematic diagram illustrating the quantization process of the pruned YOLO model according to the present invention. Detailed Implementation
[0018] The present invention will be further described in detail below with reference to the embodiments, but the implementation of the present invention is not limited thereto.
[0019] Example
[0020] Figure 1 This is a schematic diagram of the overall process of this embodiment, which includes the following 5 steps: S1. Based on the current actual scenario requirements, select models suitable for the current application scenario. Users need to select the most suitable model for the current scenario based on the actual scenario requirements, which will facilitate subsequent optimization.
[0021] S2. Based on this practical scenario requirement, construct a calibration dataset that reflects the actual operating conditions faced by edge devices. This calibration dataset needs to reflect the complex situations faced by edge devices during actual operation to avoid data bias during the quantization process.
[0022] The complexities faced by edge devices in actual operation are determined by the diversity of their deployment scenarios, the limitations of hardware resources, and the dynamic nature of the environment. These generally include environmental dynamic interference, hardware resource constraints, data input instability, and multi-device collaboration and compatibility issues. Regarding environmental dynamic interference, the focus is on edge devices deployed in open or semi-open environments, where real-time changes in environmental factors directly affect the quality of input data, thus interfering with model inference. Hardware resource constraints mainly include limitations in computing power, insufficient storage and memory, and power consumption and heat dissipation limitations. Data input instability mainly includes sensor errors, data transmission fluctuations, and data distribution offsets. Multi-device collaboration and compatibility issues mainly include differences in hardware architecture and bottlenecks in collaborative inference.
[0023] S3. First, unify the heterogeneous models into the YOLO model structure by distillation.
[0024] The distillation method refers to the use of knowledge distillation technology to transfer the "knowledge" of multiple heterogeneous models (i.e., models with different structures and types) into a unified YOLO model structure. This allows the final YOLO model to retain the advantages of heterogeneous models while maintaining the efficiency of the YOLO architecture. The YOLO model is a real-time object detection model based on deep learning. Its core feature is that it transforms the object detection task into a single regression problem. Through a single forward propagation of the neural network, object localization (boundary box prediction) and classification can be completed simultaneously, achieving efficient real-time detection.
[0025] The YOLO model typically consists of two parts: a backbone (i.e., a feature extraction network) and a detection head (also known as a "prediction layer"). The backbone is responsible for extracting multi-scale feature information from the input image; the detection head, based on the features extracted by the backbone, directly predicts the bounding box coordinates, confidence score (the probability of the target's presence), and class probability, achieving end-to-end object detection. Compared with traditional object detection methods (such as the two-stage detection of the R-CNN series, which first generates candidate regions and then classifies them), the YOLO model has advantages not only in its speed, meeting real-time requirements (such as in video surveillance and autonomous driving scenarios), but also in its very high accuracy. The YOLO model has become one of the most widely used object detection frameworks in industry and academia.
[0026] The distillation process in this embodiment is as follows: Figure 2 As shown, it includes the following steps: S31. Select pre-trained heterogeneous models as teacher models, freeze their parameters, and extract key feature layers. Standardize the features to unify the feature scale. "Pre-trained" here means that these teacher models have already been trained on large-scale datasets (such as COCO, VOC, etc.) and possess mature object detection capabilities, requiring no further parameter adjustments. "Heterogeneous models" refers to teacher models with different structures and design approaches, but all capable of performing the same object detection task.
[0027] The parameter freezing mentioned in this step refers to fixing all weight parameters of the teacher model and excluding them from subsequent distillation training. This is because the teacher model is already well-trained, and its parameter distribution is a carrier of "high-quality knowledge." Freezing the parameters can prevent the performance of the teacher model itself from degrading during distillation. Extracting key feature layers refers to selecting intermediate layer (non-output layer) features from the teacher model that are crucial to the detection task. These features include at least low-level and high-level features. The low-level features include edge and texture information, while the high-level features include semantic information.
[0028] The aforementioned feature standardization and unified feature scale are used to eliminate the "format incompatibility" problem caused by differences in model structure, allowing student models to learn features from multiple teacher models simultaneously and avoiding ineffective knowledge alignment due to scale confusion. Since the feature layer outputs of different heterogeneous models may have different numerical ranges (e.g., model A's feature values are in [0,1], while model B's are in [-10,10]) or different feature dimensions (e.g., model A outputs 64 channels, while model B outputs 128 channels), it is necessary to use methods such as normalization (e.g., subtracting the mean, dividing by the standard deviation) and dimension mapping (e.g., adjusting the number of channels through 1x1 convolution) to convert the key feature layer outputs of different teacher models into a unified numerical range and dimension.
[0029] S32. Construct a lightweight YOLO architecture as the student model, retaining the core YOLO detection head, and adjusting the backbone depth and width to match the feature dimensions of the teacher model. This step optimizes the feature extraction network structure for the student model. Since the key feature layers of the teacher model (heterogeneous model) have fixed output dimensions, and the student model's backbone needs to output features with the same dimensions as the teacher model, otherwise, a "feature format incompatibility" problem will occur. Therefore, this step must ensure that its output feature dimensions are consistent with the key feature layer dimensions of the teacher model to achieve effective knowledge transfer.
[0030] The depth mentioned here refers to the number of network layers or the number of times they are stacked in the backbone. For example, the YOLO backbone may contain multiple convolutional blocks (such as CSPBlock, Bottleneck, etc.). Increasing the depth means increasing the number of convolutional blocks or the number of times they are stacked, allowing the network to learn more complex features (such as from simple edges to complex semantics).
[0031] The width refers to the number of channels in the feature maps of each layer in the backbone. For example, if the feature map output by a certain convolutional layer has 64 channels, increasing the width increases the number of channels (e.g., by adjusting it to 128), allowing the network to capture more dimensional feature information (such as different attributes like color, texture, and shape) simultaneously.
[0032] S33. Design a multi-objective loss function: L_{total} =λ_1L_{cls}+λ_2L_{feat}+λ_3L_{arch}; where L_{cls} is the classification loss, aligning the class probability distribution of the teacher-student model; L_{feat} is the feature loss, aligning the intermediate layer features; and L_{arch} is the architecture adaptation loss, with λ_1=0.4, λ_2=0.3, and λ_3=0.3.
[0033] Aligning the class probability distributions of the teacher and student models is one of the core operations in model distillation that allows the student model to learn from the teacher model's knowledge. The goal is to ensure that the student model's prediction logic for the target class, especially its uncertainty and class association relationships, remains consistent with the teacher model. The "class probability distribution" refers to the likelihood that the target belongs to each class.
[0034] S34. The Kendall correlation coefficient is used to measure the consistency of feature distribution between the teacher model and the student model, and the consistency of feature distribution is ensured by maximizing the Kendall correlation coefficient.
[0035] This embodiment uses the Kendall correlation coefficient (τ) to measure the consistency of feature distribution between teacher and student models, ensuring feature distribution consistency by maximizing (τ). During training, (τ) is calculated every 10 epochs, and the feature mapping parameters are fixed when τ ≥ 0.85. An epoch refers to the process by which the model completes one full traversal of the entire training dataset, and it is the basic unit for measuring the model's training progress.
[0036] S35. Perform distillation training. In this embodiment, the SGD (Stochastic Gradient Descent) optimizer is used. Its function is to minimize the loss function during training by continuously adjusting the model parameters, thereby allowing the model to learn the patterns in the data. The initial learning rate is 0.001, decaying by 10% every 30 epochs, with a total of 100-150 training epochs; the input data uses data augmentation strategies such as random flipping and scaling.
[0037] S36. Test the performance of the student model on the validation set and compare it with the teacher model. The target detection mAP loss should be ≤2% and the number of parameters should be reduced by ≥50%. This step is a post-distillation validation step. Its main purpose is to test the performance of the student model on the validation set, that is, to compare the performance of the student model with that of the teacher model. If its target detection mAP loss is ≤2% and the number of parameters is reduced by ≥50%, then the student model meets the requirements for lightweight edge deployment.
[0038] S4. Prune the unified YOLO model to compress the parameter size.
[0039] See the appendix for details. Figure 3 As shown, it specifically includes the following steps: S41. Select validation set samples, calculate the pruning sensitivity of each convolutional layer in YOLO, and mark low-sensitivity layers as priority pruning targets. The main purpose of this step is to identify those network layers that have a smaller impact on model performance during model optimization, and to prioritize pruning them to reduce model complexity while minimizing performance degradation.
[0040] The sensitivity mentioned refers to the degree of influence of a network layer on the overall performance of the model, including low-sensitivity layers and high-sensitivity layers. Low-sensitivity layers are those whose removal or simplification results in only a small decrease in model performance, indicating a weak impact on the core functionality of the model. Conversely, high-sensitivity layers have a significant impact on model performance; pruning them could lead to a substantial drop in accuracy and requires careful handling.
[0041] In this embodiment, the formula for "calculating the pruning sensitivity of each YOLO convolutional layer" is: S_l = Delta mAP / DeltaParam_l, where Delta mAP is the accuracy loss after pruning 50% of the channels of the layer, and Delta Param_l is the parameter reduction ratio. When S_l < 0.01, the layer is considered to have low sensitivity.
[0042] S42. For the marked low-sensitivity layers, calculate the importance of each channel using the L2 norm, and adjust the scores based on prior knowledge of the specific scenario. This step is mainly used to achieve intelligent pruning of the low-sensitivity layers.
[0043] The L2 norm (the square root of the sum of squares of vector elements) is a fundamental metric for quantifying channel importance, widely used in channel pruning to measure the importance of channel parameters. A low L2 norm indicates a weak contribution to feature extraction, meaning removal has little impact on model performance. Conversely, a high L2 norm indicates a strong contribution to feature extraction, meaning removal has a significant impact on model performance. This embodiment calculates this value for all channels in the low-sensitivity layer to obtain an initial importance score.
[0044] In this embodiment, the channel importance metric formula for calculating the importance of each channel using the L2 norm is S_c=frac{1}{N} sum_{i=1}^N|W_{c,i}|_2, where W_{c,i} is the convolution kernel parameter of the c-th channel, and N is the number of convolution kernels.
[0045] The aforementioned scenario-specific prior knowledge refers to the manual intervention in determining the importance of channels based on the characteristics of the application scenario (such as data distribution and task objectives). It is primarily used to compensate for data bias, incorporate structural information, and introduce domain-specific indicators. Specifically, prior knowledge is combined with the L2 norm in the following ways: firstly, through weighted fusion: score = L2 norm × prior weight; secondly, through threshold correction, where channels with an L2 norm below a threshold are forcibly retained if they correspond to key features in the prior knowledge; and thirdly, through regularization constraints, adjusting the sparsity of channel parameters using L2 regularization terms.
[0046] S43. Set a basic threshold and use dynamically adjusted redundant channels for different levels to ensure the structural integrity of the model after pruning.
[0047] The basic threshold mentioned here is the initial standard for determining whether a channel is "redundant." In this embodiment, redundancy is determined by setting the channel importance score based on the L2 norm. The main reason for using dynamically adjusted redundant channels at different levels is that different network layers have vastly different functions, complexities, and impacts on overall performance in the model. Using the same threshold for pruning could lead to over-pruning of critical layers (performance crash) or under-pruning of redundant layers (poor compression effect).
[0048] S44. Use the Adam optimizer to fine-tune the pruning model for 20-30 rounds, freeze the detection head parameters and only update the backbone, and restore the accuracy through fine-tuning so that the mAP loss after fine-tuning is ≤1%.
[0049] S45. Test the number of parameters, computational load, and inference speed of the pruned model to ensure it meets the memory constraints of edge devices.
[0050] The direct goal of the above pruning is to reduce model redundancy. In this embodiment, after pruning, the optimization effect needs to be verified through quantitative indicators. Among them, the number of parameters is the total number of learnable parameters (such as weights and biases) in the model. After pruning, the number of parameters is reduced, which means that the storage space occupied by the model (such as memory and flash memory) is reduced. The computational cost is the number of operations required for the model to complete one forward inference. The smaller the computational cost, the less hardware computing power is consumed during inference, and the lower the energy consumption. The inference speed is the time it takes for the model to make a prediction on the input data (such as images and text). The faster the speed, the better it can meet the real-time requirements of edge devices.
[0051] The aforementioned compliance with edge device memory constraints means that the pruned model must meet the following conditions: first, the ROM space occupied during storage does not exceed the device's storage limit; second, the parameters and intermediate feature maps loaded into RAM during runtime do not exceed the device's memory limit. If the model exceeds these constraints, it may cause the device to fail to load the model, crash during runtime, or experience frequent stuttering.
[0052] S5. Quantize the pruned YOLO model to improve edge deployment efficiency.
[0053] The specific process for this step is as follows: Figure 4 As shown, it includes the following steps: S51. Based on the requirements for model accuracy and inference speed in the task description, a hybrid strategy combining post-training quantization (PTQ) and quantization-aware training (QAT) is adopted. First, PTQ is used to perform preliminary quantization of the model, and then QAT is used to fine-tune the model. This step is the selection of quantization strategy, which essentially adopts a hybrid strategy combining post-training quantization (PTQ) and quantization-aware training (QAT) based on the requirements for model accuracy and inference speed in the task description. PTQ is first used to perform preliminary quantization of the model, quickly reducing the accuracy of the model data, and then QAT is used to fine-tune the model to compensate for the accuracy loss caused by quantization.
[0054] S52. The weight quantization threshold is calibrated using KL divergence, and the optimal quantization interval is determined by minimizing D_{KL}. The main purpose of this step is to make the quantized weight distribution as close as possible to the original distribution, reducing accuracy loss. The formula for calculating D_{KL}(P|Q) is: D_{KL}(P|Q)=sum P(x) logfrac{P(x)}{Q(x)}, where P is the original weight distribution and Q is the quantized distribution.
[0055] S53. Freeze the weight quantization parameters, fine-tune the quantization range of activation values, and train for 10-15 rounds using the learning rate. Optimize the quantization accuracy by increasing the proportion of the targeted target in the training samples.
[0056] S54. For different edge devices, optimize according to their instruction set architecture, use tools such as TensorRT to perform operator fusion, split operators such as 3×3 convolution into forms that are more in line with the computing characteristics of the device, and optimize the memory layout of the model.
[0057] The instruction set architecture refers to the set of machine language instructions that hardware (such as CPU, GPU, NPU) can directly execute. The instruction sets of different edge devices vary significantly. The essence of instruction set optimization is to "align" the model's computational logic with the device's native instructions, reduce instruction translation overhead, and fully utilize the hardware's computing power.
[0058] The operator is the basic computational unit in the model. The forward inference of the model requires the execution of multiple operators in sequence. The execution of each operator involves the process of "reading input data from memory → calculating → writing output to memory". Frequent memory read and write will become an efficiency bottleneck.
[0059] Operator fusion refers to combining multiple consecutive operators into a single composite operator using optimization tools (such as TensorRT and ONNX Runtime), reducing the number of memory read / write operations on intermediate data. In this step, the TensorRT tool automatically analyzes the model's computation graph, identifies fusionable operator combinations, and generates an optimized computational flow. The 3×3 convolution is one of the most commonly used operators in deep learning. This step breaks down the operator into forms more suited to the device's computational characteristics, aiming to match the operator's computational granularity with the hardware's parallel processing capabilities and reduce computational complexity.
[0060] In this step, the memory layout refers to the storage format of the model's input data, intermediate feature maps, and weight parameters in memory. Edge devices typically have low memory bandwidth (data read / write speed), and CPU / GPU cache sizes are limited. Optimizing the memory layout can reduce memory access conflicts, improve cache hit rate, and accelerate data read / write.
[0061] The optimizations described in this embodiment include at least adjusting the dimension order; data alignment, which refers to storing data in alignment according to the hardware's memory bus width (such as 64 bytes or 128 bytes) to avoid read / write efficiency loss due to misaligned memory addresses; and compression and rearrangement, which refers to rearranging the weight parameters or intermediate feature maps in memory so that consecutive calculations access consecutive memory addresses, making full use of the cache's preloading mechanism.
[0062] In summary, this embodiment achieves a comprehensive effect of "accuracy loss ≤3% + parameter reduction ≥85% + inference speed improvement ≥200%" on edge devices through three-step collaborative optimization of "distillation-pruning-quantization", thus adapting to the deployment requirements of resource-constrained devices.
[0063] As described above, the present invention can be well implemented.
Claims
1. A method for distilling and migrating AI models to the edge for deployment across chip architectures, characterized in that, Includes the following steps: S1. Based on the current actual scenario requirements, select models suitable for the current application scenario; S2. Based on the actual scenario requirements, construct a calibration dataset that can reflect the actual operating conditions faced by edge devices; S3. First, unify the heterogeneous models into the YOLO model structure through distillation, which includes the following steps: S31. Select a pre-trained heterogeneous model as the teacher model, freeze the parameters and extract key feature layers, and unify the feature scale through feature standardization. The frozen parameters refer to fixing all weight parameters of the teacher model and not participating in subsequent distillation training. The extraction of key feature layers refers to selecting intermediate layer features that are crucial to the detection task from the teacher model. These features include at least low-level features and high-level features. The low-level features include edge and texture information, and the high-level features include semantic information. The unification of feature scale through feature standardization involves adjusting the number of channels through normalization and 1x1 convolution to convert the key feature layer outputs of different teacher models into a unified numerical range and dimension. S32. Construct a lightweight YOLO architecture as the student model, retain the core YOLO detection head, and adjust the backbone depth and width to match the feature dimensions of the teacher model. S33. Design a multi-objective loss function: ;in, For classification loss, align the class probability distributions of the teacher-student model; For feature loss, align the intermediate layer features; For architectural adaptation losses, =0.4、 =0.3、 =0.3; S34. The Kendall correlation coefficient is used to measure the consistency of feature distributions between the teacher model and the student model, and the consistency of feature distributions is ensured by maximizing the Kendall correlation coefficient; the Kendall correlation coefficient is... During training, the calculation is performed every 10 epochs. ,when Fixed feature mapping parameters when ≥0.85; S35. Conduct distillation training; S36. Test the performance of the student model on the validation set and compare it with the teacher model. The mAP loss for object detection is ≤2% and the number of parameters is reduced by ≥50%, which meets the lightweight requirements for edge deployment. S4. Prune the unified YOLO model to reduce the parameter size, including: S41. Select validation set samples, calculate the pruning sensitivity of each YOLO convolutional layer, and mark low-sensitivity layers as priority pruning targets; the formula for calculating the pruning sensitivity of each YOLO convolutional layer is: ,in, This represents the accuracy loss after pruning 50% of the channels in this layer. The parameter reduction ratio; the labeled low-sensitivity layer refers to this <0.01; S42. For the marked low-sensitivity layers, calculate the importance of each channel using the L2 norm, and adjust the score by combining prior knowledge of specific scenarios; the prior knowledge of specific scenarios refers to the manual intervention on the importance of channels based on the data distribution and task objectives of the application scenario. S43. Set a basic threshold and use dynamically adjusted redundant channels for different levels to ensure the integrity of the model structure after pruning. S44. Fine-tune the pruning model for 20-30 rounds using the Adam optimizer, freeze the detection head parameters and only update the backbone, and restore accuracy through fine-tuning methods so that the mAP loss after fine-tuning is ≤1%; S45. Test the number of parameters, computational load and inference speed of the pruned model to meet the memory constraints of edge devices; S5. Quantize the pruned YOLO model to improve edge deployment efficiency.
2. The method for edge deployment of AI model distillation migration across chip architectures according to claim 1, characterized in that, Step S5 involves quantizing the pruned YOLO model, including the following steps: S51. Based on the requirements for model accuracy and inference speed in the task book, a hybrid strategy of post-training quantization and quantization-aware training is adopted. First, post-training quantization is used to perform preliminary quantization of the model, and then quantization-aware training is used to fine-tune the model. S52. Use KL divergence calibration weight quantization threshold, and minimize... To determine the optimal quantization interval, the Where P is the original weight distribution and Q is the quantized distribution; S53. Freeze the weight quantization parameters, fine-tune the quantization range of activation values, and train for 10-15 rounds using the learning rate. Optimize the quantization accuracy by increasing the proportion of the targeted target in the training samples. S54. For different edge devices, optimize according to their instruction set architecture, use TensorRT tool to perform operator fusion, split the 3×3 convolution operator into a form that is more in line with the device's computing characteristics, and optimize the model's memory layout.
3. The method for cross-chip architecture AI model distillation migration edge deployment according to claim 2, characterized in that, The channel importance metric formula for calculating the importance of each channel using the L2 norm in step S42 is as follows: 2, of which Here, N represents the parameters of the i-th convolutional kernel in the c-th channel, and N is the number of convolutional kernels. It is an L2 norm.
Citation Information
Patent Citations
Neural network reasoning acceleration method, target detection method, equipment and storage medium
CN116702835A
Enhanced retrieval-based agent rapid construction method and system
CN120407751A
Portable deep learning model lightweight method for rapeseed quality detection
CN120408132A