Lightweight ViT based on image feature cutting and cloud edge knowledge distillation

Optimizing the ViT model through image feature cropping and cloud edge knowledge distillation, the problems of high computing cost and low inference efficiency on edge devices are solved, and efficient and accurate real-time image inference and multitasking processing are achieved on edge devices.

CN120375050APending Publication Date: 2025-07-25NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510440087.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing ViT model has problems such as high computing cost, large communication latency and low inference efficiency when deploying on edge devices. The existing lightweight model still has a large gap in accuracy and inference delay, which is difficult to meet the practical application needs.

Method used

By designing the image feature cropping module, full quantization compression of ViT attention layer and cloud-edge knowledge distillation, the ViT model framework is optimized, edge devices and cloud-side training can be achieved, dynamically adjust the pruning ratio and parameter updates can be updated, and calculation and communication costs can be reduced.

Benefits of technology

Real-time image inference is implemented on edge devices with resource-constrained resources, effectively reducing computing overhead and communication delays, while ensuring the accuracy and efficiency of model inference, and supporting multi-task migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375050A_ABST
    Figure CN120375050A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight ViT based on image feature cutting and cloud edge knowledge distillation, belongs to the technical field of computer vision and edge computing, is used for efficiently deploying a ViT model on an edge device with limited computing power, and comprises the steps of designing a lightweight ViT Block and a sub-sampling Block. The method comprises the following steps: designing an image feature cutting module, collecting images by edge equipment, screening highly uncertain samples in the images as a knowledge distillation data set, and transmitting the knowledge distillation data set and local model parameters to a cloud end. A ViT model is deployed for the edge equipment with limited resources and weaker performance to carry out an image processing task, and the ViT is subjected to lightweight processing; therefore, the Flops of the model is reduced while the accuracy of the model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and edge computing technology, and specifically relates to a lightweight ViT based on image feature clipping and cloud-edge knowledge distillation. Background Art

[0002] Current neural network models have been applied to all aspects of life. Visual models such as CNN and ViT have been applied to many fields such as face object recognition, motion trajectory tracking, traffic violation identification and traffic management. Among them, Vision Transformer has excellent performance in various tasks due to its powerful model reasoning ability.

[0003] However, ViT's powerful performance also brings high computing costs. Most Transformer-based visual models require about 10B parameters to achieve optimal performance on large datasets such as ImageNet-1k. Model parameters of tens of billions will bring model sizes and computing costs that most edge devices cannot afford. This is also the core bottleneck for deploying ViT and LLM to edge devices. Existing traditional optimization methods still face the following problems:

[0004] Most image recognition devices currently used in real life do not support local real-time data processing, but transmit the captured images to the cloud, where powerful GPU clusters process massive images and transmit recognition results to edge devices. However, the disadvantages of relying entirely on cloud reasoning are that data privacy cannot be met, the communication cost is high due to the large amount of data transmitted, and the communication delay is high. If the cloud equipment is damaged, a large number of devices will be unusable.

[0005] Existing lightweight ViT models such as MCUFomer and MCUViT are fully deployed on edge devices. Although they have successfully compressed larger models such as ViT and can even run on MCU, the core idea of the above models is to design a tool that can achieve inference efficiency and model accuracy, that is, to ensure the accuracy of inference by compressing the model size and Flops as much as possible. However, the throughput of the model itself has not been effectively improved. The higher accuracy in the experimental data is due to the inference latency that is unacceptable for most actual application devices. TinyViT improves the generalization ability and accuracy of the model on the smaller ViT model by distilling the knowledge of the large model to the small model, but the inference efficiency and accuracy are still far behind the high-performance model. Summary of the invention

[0006] In view of the deficiencies in the above background, the present invention proposes a lightweight Vision Transformer (ViT) method based on image feature cropping, attention quantization, and cloud-edge-end knowledge distillation. This method realizes the deployment of cloud-edge collaborative training and inference models on resource-constrained edge devices (such as embedded sensors and mobile terminals) through a dynamic feature screening mechanism for images, full quantization compression technology for ViT attention layers and linear layers, an improved ViT Block architecture, and dynamic weight update of the knowledge distillation of the cloud teacher model for the edge device student model, completing real-time image inference tasks, and ensuring the accuracy of model inference while effectively reducing computational overhead and communication latency.

[0007] Step S1: Optimize the ViT model framework and design a lightweight ViT Block: In view of problems such as insufficient computing power and storage space when deploying complex models on edge devices, design a lightweight Vision Transformer, and while compressing the model to a size deployable on the device as much as possible, ensure the model's ability to process tasks.

[0008] Step S2: Design an image feature cropping module: Deploy an image feature cropping module before the ViT Block to crop tokens with less information, and realize dynamic adjustment of the pruning ratio according to the activation degree of real-time input features during the inference stage, balancing computational efficiency and model accuracy.

[0009] Step S3: The edge device collects images, screens high-uncertainty samples in the images as the knowledge distillation dataset and transmits them to the cloud together with the local model parameters; the cloud performs ThreeAugment data augmentation on the dataset images; uses adaptive dynamic knowledge distillation to update the model parameters, generates soft labels and intermediate feature maps, guides the optimization of the edge-end student model through a dynamic weighted loss function, and transmits the weighted updated parameters back to the edge end to optimize the model.

[0010] Further, in step S1, implementing a complete ViT Block includes a quantization cascaded group attention module and a subsampling module, and the specific steps are as follows:

[0011] Step S1-1: Construct the first forward propagation layer:

[0012] The model contains 3 ViT Blocks and 2 subsampling modules. Each ViT Block processes the processed input feature map F in the following order ∈ R B×C×H×W , where B is the batch size, C is the number of channels, and H×W is the spatial resolution. Apply a 3×3 depthwise separable convolution DW Conv to extract local spatial features F DW1 = DW Conv3×3 (F), and the specific calculation formula is:

[0013]

[0014] O(i, j) is an element on the output feature map, where I(i, j) is an element on the input feature map, and K(k, l) is an element in the depth convolution kernel. Then, through a 1×1 convolution expansion layer, a GELU activation function, and another 1×1 convolution contraction layer, the computational cost is reduced compared to a 3×3 convolution layer, and the model's expressive power is improved:

[0015] F FFB1 = W2·ELU(W1·F DW1 )

[0016] where W1 ∈ R 2C×C , W2 ∈ R C×2C .

[0017] Step S1-2: Construct a quantized cascaded group attention layer:

[0018] The input feature map F FFB1 ∈ R B×C×H×W after the first forward propagation is evenly divided into G groups along the channel dimension (where G is the number of attention heads in the selected ViT BackBone), and the number of channels in each group is C / G:

[0019] {F1, F2, …, F G} = Split(F, dim = 1, num_groups =)

[0020] Process the i-th group of features in order. If i > 1, add to the input feature map: Add the current group input F i to the output of the previous group : For , generate the query matrix Q i , key matrix K i , and value matrix V i through 1×1 convolution for qkv. Q i , K i , V i The sequence length is l, and the dimension is d Q = l×C / (4), d Q / K = l×C / , where:

[0021]

[0022] At the same time, apply a 5×5 depthwise separable convolution to Q i to enhance local information:

[0023] Q i ← DWConv5×5 (Q i ),

[0024] Calculation formula for the attention matrix of the i-th attention head:

[0025] O i = PV. Since Q i , K i , V i The sequence length l is much larger than the dimension d Q (l >> d Q ), so the attention calculation formula is optimized. On this basis, by smoothing the matrix K, reduce Q i , K i , V i Precision loss of Int8 quantization: Define the quantization function

[0026]

[0027] Solve for the smoothing matrix K i :

[0028] γ(K i ) = K i - max(K i ),

[0029] Attention score matrix after Int8 quantization

[0030]

[0031] Retain the output of the current group for use by the next group.

[0032] Concatenate the outputs of each group along the channel dimension and unify the dimensions through 1×1 convolution:

[0033]

[0034] Repeat the operation in step S1-1 to further fuse the local feature F res2 = F CA + DWConv 3×3 (F CA ), complete the final feature mapping: F FFB2 = W4·ELU(W3·F res2 ), obtain the output feature map F out = F FFB2 .

[0035] Step S1-3: Construct the subsampling module:

[0036] The design of the subsampling module is similar to the ViT Block, but the attention layer is replaced by an inverted residual block. The specific formula is as follows: F out2 = F FFB2 ·Residual·F FFB1 (F out1 ).

[0037] Furthermore, in step S2, an image token pruning module is designed to prune the tokens with less information, thereby reducing the model's Flops. The specific steps are as follows:

[0038] Step S2-1: During the training process, generate training signals for training the scoring module:

[0039] The formula for the attention matrix of the feature tokens matrix X is Use the attention matrix A to perform weighted averaging on the input tokens matrix X to generate an aggregated representation X' of X, and enhance it using a single layer of MLP:

[0040]

[0041] where LN is layer normalization and MLP is a multi-layer perceptron.

[0042] During the model training stage, through the Classification Head Block classification head module, generate an intermediate prediction signal Logits = Linear(LN(X')) based on X', and the loss function F loss = Cross-Entropy Loss to calculate the error between the predicted Logit and the true label, and generate a gradient signal:

[0043]

[0044] where z c is the output of the c-th class in Logit, and y c is the one-hot encoding 0 / 1 of the true label.

[0045] The model calculates the task relevance score of each token through cross-attention based on the prediction results The tokens with high scores Score have high task relevance. Low-scoring tokens are usually background or noise tokens. Prune the redundant tokens to effectively reduce Flops and memory usage:

[0046] X k = Top-K(X|a), X p = X\X k

[0047] The number of token pruning k decreases gradually. As the network goes deeper, the number of tokens in subsequent layers gradually decreases, reducing the amount of calculation in subsequent layers.

[0048] Step S3-1: Dataset sample sampling based on uncertainty:

[0049] Since most edge devices have limited data transmission bandwidth, it is unrealistic to transmit all samples collected by edge devices to the cloud for processing. In order to maximize the training efficiency of transmitted samples, the dataset images are sampled based on uncertainty. i For the same image x i Perform n independent forward propagations to obtain n groups of category probability distributions {p1(y|x i ),p2(y|x i ),…,p n (y|x i )}, calculate the mean probability, and measure the uncertainty of the model prediction through the standard deviation:

[0050]

[0051] V unc Larger values indicate that the model is less certain about its predictions for that image.

[0052] Set a predefined threshold θ and filter the unc >θ high uncertainty image x i , only the selected high uncertainty image set {x i |V unc >θ} is transmitted to the cloud server for subsequent model optimization.

[0053] Step S3-2: Preprocessing of data images. The model uses the ThreeAugment data enhancement method for the images in the data set: Let the input image be I∈R H×W×3 The preprocessing process can be expressed as a composite function The operations at each stage are as follows: Perform primary spatial transformation on image I, and randomly scale, crop, and flip during training:

[0054]

[0055] Where s is the crop area ratio, adjusted to the target size S target =224×224,

[0056] When testing the model, normalize the images to a fixed size:

[0057] F resize-center (I) = CenterCrop 224· Resize 256 (I)

[0058] Perform secondary image enhancement operations on image I:

[0059]

[0060] Grayscale conversion:

[0061]

[0062] Exposure inversion:

[0063]

[0064] Gaussian blur:

[0065]

[0066] Finally, perform normalization and sampling operations:

[0067]

[0068] Let the total number of samples be N and the number of GPUs be K. The image index assignment satisfies:

[0069]

[0070] Step 3-3: The dynamic parameter update strategy for transferring the edge model parameters to cloud knowledge distillation is as follows:

[0071] While the edge device transfers high-uncertainty samples to the cloud for processing, it uploads the device model parameters to the cloud as the baseline model parameters θ for parameter update Base . Let the edge student model parameters be θ edge , select the pre-trained ViT-Large / 16 as the teacher model, load the model parameters from the pre-trained weights, freeze all layers, and only use them for forward inference to generate soft label probabilities Guide the optimization of the lightweight student model through the soft labels and intermediate feature maps of the teacher model.

[0072] Define the knowledge distillation loss function L total , comprehensively consider the cross-entropy loss and KL divergence loss and perform dynamic weighting operations:

[0073] Assume that the true label is one-hot encoded y ∈ {0, 1} C , and the prediction probability of the student model is p s = softmax(z s ), then the cross-entropy loss is:

[0074]

[0075] The gradient with respect to the parameter θ is as follows:

[0076]

[0077] Let the soft label probability of the teacher model be p t = softmax(z t / T), and the predicted probability of the student model be p s = softmax(z s / T). Then the KL divergence loss is:

[0078]

[0079] where T is used to soften the distribution, and the gradient with respect to the parameter θ is:

[0080]

[0081] The total loss function of knowledge distillation is the weighted sum of cross-entropy and KL divergence:

[0082]

[0083] The combined gradient formula is:

[0084]

[0085] The gradient with respect to the parameter θ after combination is:

[0086]

[0087] By minimizing the updated model parameters θ of knowledge distillation are generated new .

[0088] Step S3-4: Cloud teacher model adaptive dynamic parameter quantization update strategy:

[0089] To solve the problem of low parameter transmission efficiency when the cloud teacher model updates parameters during knowledge distillation, and at the same time ensure the accuracy of the important parameters transmitted, an adaptive dynamic quantization is performed on the difference value Δθ between the updated parameter set θ new of the knowledge distillation optimization and the device-side reference parameter set θ base :

[0090] Δθ = θ new - θ base , θ q = Quantize(Δθ, Q)

[0091] where Q is a dynamically calibrated quantization function that preferentially retains high-significance parameter updates during parameter quantization and uses a higher compression ratio for low-significance parameters.

[0092] Transmit the compression parameter θ q to the edge device through the communication link, significantly reducing the bandwidth occupancy and transmission delay. The device receives the compression parameter θ q and directly superimposes it with the reference parameter θ base to generate the updated edge model parameter θ edge :

[0093] θ edge = θ base + θ q

[0094] Adaptive knowledge distillation enables the student model deployed on the edge device to fully learn the generalization ability and representation ability of the teacher model while ensuring the agility and efficiency of the model in processing visual tasks. At the same time, it reduces the communication cost of optimizing the model through knowledge distillation by the device model, which is crucial for deployment in a typical device computing environment with limited resources.

[0095] Compared with the prior art, the beneficial effects of the present invention are:

[0096] Effectively reduce the Flops of the model, the data transmission cost and parameter update cost of cloud knowledge distillation, improve the GPU throughput rate, and greatly reduce the inference delay of the edge device.

[0097] In the imagenet-1k classification task, the use of adaptive knowledge distillation, data augmentation and other operations significantly improves the Top-1 accuracy on devices with the same computing power.

[0098] Compared with the traditional ViT model, the peak memory occupancy is significantly reduced, the number of model parameters is greatly reduced, and it is suitable for use on edge devices.

[0099] Optimized in the knowledge distillation strategy, reducing the cloud-device two-way communication cost, and enabling the student model deployed on the edge device to fully learn the generalization ability and representation ability of the teacher model while ensuring the agility and efficiency of the model in processing visual tasks.

[0100] Supports multi-task migration, and the feature map output by the model can be used for visual tasks such as object detection and semantic segmentation by adding corresponding task heads. Brief Description of the Drawings

[0101] Figure 1 It is the design framework diagram of the device ViT of the present invention;

[0102] Figure 2 It is the schematic diagram of the quantization cascade group attention of the ViT Block of the present invention;

[0103] Figure 3Schematic diagram of the image cropping algorithm in the present invention;

[0104] Figure 4 Schematic diagram of the lightweight ViT model implemented based on image feature cropping, attention quantization, and cloud-edge-end knowledge distillation in the present invention;

[0105] Figure 5 Experimental result diagram of the present invention. Specific implementation method

[0107] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.

[0108] Step S1: Optimize the ViT model framework and design a lightweight ViT Block: In view of problems such as computational examples and insufficient storage space occurring in the deployment of complex models on edge devices, design a lightweight Vision Transformer, while compressing the model to a size deployable on the device as much as possible and ensuring the model's ability to process tasks.

[0109] In step S1, for the lightweight ViT Block, the designed quantized cascade group attention block is optimized in terms of device memory overhead, model inference efficiency, and model inference accuracy compared with the ordinary ViTBlock. The overall architecture schematic diagram of ViT is as Figure 1 shown:

[0110] The present invention uses 3 ViT Blocks and 2 subsampling modules. Each ViT module consists of two feed-forward propagation modules (FeedForward Block) and a quantized cascade group attention block (Quantized Cascade Group AttentionBlock). The input ViT Block feature map is F ∈ R B×C×H×W , B is the batch size, C is the number of channels, H×W is the spatial resolution. The first FFB layer contains an applied 3×3 depthwise separable convolution DW Conv , extracting local spatial features F DW1 = DW Conv3×3 (F). The specific calculation formula is: O(i,j) is an element on the output feature map, where I(i,j)I(i,j) is an element on the input feature map, and K(k,l)K(k,l) is an element in the depth convolution kernel. Then, through a 1×1 convolution expansion layer, GELU activation function, and another 1×1 convolution contraction layer, the computational amount is reduced compared with the 3×3 convolution layer and the model expressiveness is improved: F FFB1 = W2·ELU(W1·F DW1 ) where W1 ∈ R 2C×C,W2∈R C×2C 。

[0111] The design of the subsampling module is similar to that of the ViT Block, but the attention layer is replaced by an inverted residual block. The specific formula is as follows: F out2 = F FFB2 ·Residual·F FFB1 (F out1 ). Immediately afterwards, the first forward propagation layer is transmitted to the Quantized Cascade Group Attention Block. The framework diagram of the attention module is as shown in Figure 2 : The input feature map F FFB1 ∈R B×C×H×W after passing through the FF layer forward propagation is evenly divided into G groups along the channel dimension, where G is the number of attention heads in the selected ViT BackBone), and the number of channels in each group is C / G:

[0112] {F1,F2,…,F G}= Split(F, dim = 1, num_groups = )

[0113] Process the features of the i-th group in order. If i>1, superimpose on the input feature map: Add the current group input F i to the output of the previous group : Then, the input feature qkv is input into the quantized attention block for attention matrix quantization. The qkv of generates the query matrix Q i , the key matrix K i , and the value matrix V i through a 1×1 convolution. Q i , K i , V i The sequence length is l, and the dimension is d Q = l×C / (4), d Q / K = l×C / , where At the same time, apply a 5×5 depthwise separable convolution to Q i to enhance local information:

[0114] Q i ← DWConv 5×5 (Q i )

[0115] The calculation formula for the attention matrix of the i-th attention head: P i = softmax(S i ), O i = PV. Since Qi , K i , V i The sequence length l is much larger than the dimension d Q (l >> d Q ), so the attention calculation formula is optimized. On this basis, the smooth matrix K is used to reduce Q i , K i , V i Precision loss of Int8 quantization: Define the quantization function

[0116] Solve the smooth matrix K i :

[0117]

[0118] Attention score matrix after Int8 quantization

[0119]

[0120] Retain the output of the current group for use in the next group.

[0121] Concatenate the outputs of each group along the channel dimension and unify the dimensions through a 1×1 convolution:

[0122]

[0123] Pass the output F of the attention layer CA through the forward propagation module again to further fuse the local feature F res2 = F CA + DWConv 3×3 (F CA ), to obtain the final feature map: F FFN2 = W4·ELU(W3·F res2 ), to obtain the final output feature map F out1 = F FFN2 . The output feature map F out1 is output through the subsampling model as:

[0124] F out2 = F FFB2 ·Residual·F FFB1 (F out1 )

[0125] The final input feature map can be connected to different task heads to achieve tasks such as image classification and object detection.

[0126] Step S2: Design an image feature cropping module: As Figure 1As shown in the figure, an image feature cropping module is deployed before the ViT Block, and tokens with less cropped information are input before the ViT Block, so as to dynamically adjust the pruning ratio according to the activation degree of real-time input features during the inference stage, and balance the computational efficiency and model accuracy.

[0127] The algorithm flow of image feature cropping is as Figure 3 shown: The formula for the attention matrix of the feature tokens matrix X is Use the attention matrix A to perform weighted averaging on the input tokens matrix X to generate an aggregated representation X′ of X, and use a layer of MLP for enhancement:

[0128]

[0129] where LN is layer normalization and MLP is a multi-layer perceptron.

[0130] During the model training stage, through the Classification Head Block classification head module, an intermediate prediction signal Logits = Linear(LN(X′)) is generated based on X′, and the loss function F loss = Cross-Entropy Loss is used to calculate the error between the predicted Logit and the true label, and a gradient signal is generated:

[0131]

[0132] where z c is the output of the c-th class in Logit, and y c is the one-hot encoding 0 / 1 of the true label.

[0133] The model calculates the task relevance score of each token through cross-attention according to the prediction result Tokens with a high score Score have high task relevance. Tokens with a low score are usually background or noise tokens. Pruning redundant tokens can effectively reduce Flops and memory occupancy:

[0134] X k = Top-K(X|a), X p = X\X k

[0135] where the number k of token cropping decreases. As the network deepens, the number of tokens in the subsequent layers gradually decreases, reducing the computational amount of the subsequent layers.

[0136] Step S3: The edge device collects images, filters out high-uncertainty samples in the images as knowledge distillation datasets and transmits them to the cloud together with local model parameters; the cloud performs ThreeAugment data enhancement on the dataset images; uses cloud-edge adaptive knowledge distillation to update model parameters, generate soft labels and intermediate feature maps, guide edge student model optimization through a dynamic weighted loss function, and transmit the weighted update parameters back to the edge optimization model.

[0137] The schematic diagram of the lightweight ViT model based on image feature clipping, attention quantization and cloud-edge knowledge distillation implemented in the present invention is as follows: Figure 4 As shown, the specific steps are as follows:

[0138] The dataset image is sampled based on uncertainty, and the input image x i For the same image x i Perform n independent forward propagations to obtain n groups of category probability distributions {p1(y|x i ),p2(y|x i ),…,p n (y|x i )}, calculate the mean probability, and measure the uncertainty of the model prediction through the standard deviation:

[0139]

[0140] V unc Larger values indicate that the model is less certain about its predictions for that image.

[0141] Set a predefined threshold θ and filter the unc >θ high uncertainty image x i , only the selected high uncertainty image set {x i |V unc >θ} is transmitted to the cloud server for subsequent model optimization.

[0142] After the high uncertainty image samples are transmitted to the cloud via the uplink, the data images are preprocessed. The present invention uses the ThreeAugment data enhancement method for the images in the data set: let the input image be I∈R H×W×3 The preprocessing process can be expressed as a composite function Perform primary spatial transformation on image I, randomly scaling, cropping and flipping during training:

[0143] F crop-flip (I)=Flip p=0.5 (Resize(Crop(I,s~U(0.08,1.0))))

[0144] Where s is the crop area ratio, adjusted to the target size Starget = 224 × 224,

[0145] During model testing, the image is normalized to a fixed size:

[0146] F resize-center (I) = CenterCrop 224 · Resize 256 (I)

[0147] Perform a secondary image enhancement operation on image I:

[0148]

[0149] Grayscale conversion:

[0150]

[0151] Exposure inversion:

[0152]

[0153] Gaussian blur:

[0154]

[0155] Finally, perform normalization and sampling work:

[0156]

[0157] Let the total number of samples be N and the number of GPUs be K. The image index assignment satisfies:

[0158]

[0159] Step 3 - 3: The dynamic parameter update strategy for transmitting the edge model parameters to cloud knowledge distillation is as follows:

[0160] While the edge device transmits high - uncertainty samples to the cloud for processing, it uploads the device model parameters to the cloud as the benchmark model parameters θ Base . Let the edge - side student model parameters be θ edge , select the pre - trained ViT - Large / 16 as the teacher model, load the model parameters from the pre - trained weights, freeze all layers, and only use them for forward inference to generate soft label probabilities Guide the optimization of the lightweight student model through the soft labels and intermediate feature maps of the teacher model.

[0161] Define the knowledge distillation loss function L total , comprehensively consider the cross - entropy loss and KL - divergence loss and perform dynamic weighting operations:

[0162] Assume the true label is the one-hot encoded \(y\in\{0,1\}\). C , and the predicted probability of the student model is \(p\). s =\(\text{softmax}(z\). s ), then the cross-entropy loss is:

[0163]

[0164] The gradient with respect to the parameter \(\theta\) is:

[0165]

[0166] Let the soft label probability of the teacher model be \(p\). t =\(\text{softmax}(z\). t / T), and the predicted probability of the student model is \(p\). s =\(\text{softmax}(z\). s / T), then the KL divergence loss is:

[0167]

[0168] where \(T\) is used to soften the distribution, and the gradient with respect to the parameter \(\theta\) is:

[0169]

[0170] The total loss function of knowledge distillation is the weighted sum of cross-entropy and KL divergence:

[0171]

[0172] The combined gradient formula is:

[0173]

[0174] The gradient with respect to the parameter \(\theta\) after combination is:

[0175]

[0176] By minimizing generate the updated model parameter \(\theta\) of knowledge distillation new .

[0177] Perform adaptive dynamic quantization on the difference value \(\Delta\theta\) between the updated parameter set \(\theta\) optimized by knowledge distillation new and the device-side reference parameter set \(\theta\). base :

[0178] \(\Delta\theta=\theta\). new -\(\theta\). base ,\(\theta\). q =\(\text{Quantize}(\Delta\theta,Q)\).

[0179] Among them, Q is a dynamically calibrated quantization function, which preferentially retains high-significance parameter updates during the parameter quantization process and adopts a higher compression ratio for low-significance parameters.

[0180] Transmit the compressed parameter θ q to the edge device through the communication link, significantly reducing the bandwidth occupancy and transmission delay. The device receives the compressed parameter θ q and directly superimposes it with the reference parameter θ base to generate the updated edge model parameter θ edge :

[0181] θ edge = θ base + θ q

[0182] To prove the effectiveness of the present invention, preliminary experiments were conducted. The method proposed in the present invention was compared with EfficientViT. Among them, the training method uses ImageNet as the sample data set and is constructed using Pytorch 1.11.0 and Timm 0.5.4. It is trained from scratch for 300 epochs on 8 NVIDIA V100 GPUs using the ADAMW optimizer and the cosine learning rate scheduler, and the total batch size is set to 2,048. The model adjusts the input image and randomly crops it to 224×224, the initial learning rate is 1×10-3, and the weight decay is 2.5×10-2. The two models were trained with six different parameter dimensions (including the number of attention heads, embedding layer dimensions, etc.) for six different sizes of models M0-M5, and the comparison results are as Figure 5 shown. Since this model adopts operations such as image cropping, attention quantization, and knowledge distillation, the throughput rate is increased by more than 40% compared with the EfficientViT model, and the loss of model accuracy is only within 1%. The ViT model trained using the method of the present invention can not only quickly recognize images but also ensure the accuracy of the model. Therefore, the present invention is effective.

[0183] The above description is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those of ordinary skill in the art according to the content disclosed in the present invention shall be included in the protection scope recorded in the claims.

Claims

1. A lightweight ViT based on image feature cropping and cloud-edge-end knowledge distillation, characterized in that, It includes the following steps: Step 1, design lightweight ViT Block and subsampling Block: In response to the problems that occur when deploying complex models on edge devices, design a lightweight Vision Transformer to ensure the model's ability to process tasks while compressing the model to a size deployable on the device as much as possible; Step 2, design an image feature cropping module: Deploy an image feature cropping module before the ViT Block, crop tokens during subsampling, and dynamically adjust the pruning ratio according to the activation degree of real-time input features during the inference stage to balance computational efficiency and model accuracy; Step 3, the edge device collects images, screens high-uncertainty samples in the images as a knowledge distillation dataset, and transmits them to the cloud together with the local model parameters; The cloud performs ThreeAugment data augmentation on the dataset images; uses adaptive dynamic knowledge distillation to update the model parameters, generates soft labels and intermediate feature maps, guides the optimization of the edge-side student model through a dynamic weighted loss function, and transmits the weighted updated parameters back to the edge side to optimize the model.

2. The lightweight ViT based on image feature cropping and cloud-edge-end knowledge distillation according to claim 1, wherein In Step 1, the process of the ViT Block and the subsampling Block for processing images to generate feature maps is as follows: Design the ViT Block to process the processed input feature map F in the following order ∈ R B×C×H×W , where B is the batch size, C is the number of channels, and H×W is the spatial resolution. Apply the depthwise separable convolution DW Conv , and extract the local spatial feature F DW1 = DW Conv3×3 (F), and the specific calculation formula is: Where O(i,j) is an element on the output feature map, I(i,j) is an element on the input feature map, K(k,l) is an element in the depth convolution kernel, and then through a 1×1 convolution expansion layer, a GELU activation function, and another 1×1 convolution contraction layer. Compared with directly using a 3×3 convolution layer, using two 1×1 convolutions reduces the computational amount and improves the model's expressiveness: F FFN1 = W2 · ELU(W1 · F DW1 ) where W1 ∈ R 2C×C , W2 ∈ R C×2C ; The input feature map F after the first - layer forward propagation FFN1 ∈R B×C×H×W is block - processed and evenly divided into G groups along the channel dimension, where G is the number of Vision Transformer Attention Heads, and the number of channels in each group is C / G, denoted as: {F1,F2,…,F G} = Split(F, dim=1, num_groups=) Process the i-th group of features sequentially. If i > 1, superimpose on the input feature map: the current group input F i with the output of the previous group and add them together: For the qkv, generate the query matrix Q through a 1×1 convolution i , the key matrix K i , and the value matrix V i , Q i , K i , V i The sequence length is l and the dimension is d Q = l×C / (4), d Q / K = l×C / : Meanwhile, for Q i Apply a 5×5 depthwise separable convolution to enhance local information: Q i ←DWConv 5×5 (Q i ) To reduce Q i , K i , V i For the calculation amount of matrix multiplication of the attention matrix, perform block and quantization operations on the calculation of the attention matrix to reduce the floating-point calculation amount, and reduce the precision loss through a smoothing matrix; The calculation formula for the attention matrix of the i-th attention head: O i = PV, On this basis, reduce Q through the smoothing matrix K i ,K i ,V i Precision loss of Int8 quantization: Define the quantization function: Solve the smoothing matrix K i : The attention score matrix after Int8 quantization: Reserve the current group output For use by the next group; Concatenate the outputs of each group along the channel dimension and unify the dimensions through a 1×1 convolution: Repeat and stack depthwise separable convolutions to further fuse local features: F res2 = F CGA + DWConv 3×3 (F CGA ) Complete the final feature mapping: F FFN2 = W4·ELU(W3·F res2 ) Obtain the final output feature map F out1 = F FFN2 ; The design of the subsampling module also uses two forward propagations like the ViT Block, but the attention layer is replaced by an inverted residual block, and the specific formula is as follows: F out2 = F FFB2 ·Residual·F FFB1 (F out1 )。 3. The lightweight ViT based on image feature cropping and cloud-edge-terminal knowledge distillation according to claim 1, wherein: In Step 2, the token cropping module includes: The formula for the feature tokens matrix X and the attention matrix is The input tokens matrix X is weighted and averaged using the attention matrix A to generate an aggregated representation X' of X, and enhanced using a single layer of MLP: aggregator(X∣A) = MLP(LN(X′)) + X′ Where LN is layer normalization and MLP is a multi-layer perceptron; During the model training phase, through the Classification Head Block, an intermediate prediction signal Logits = Linear(LN(X′)) is generated based on X′, and the loss function F loss is the cross-entropy loss Calculate the error between the predicted Logit and the true label to generate a gradient signal: where z c is the output of the c-th class in Logit, and y c is the one-hot encoding 0 / 1 of the true label; Based on the prediction results, the model calculates the task relevance scores of each token through cross-attention Tokens with high scores have high task relevance, while tokens with low scores are background or noise tokens. Pruning redundant tokens can effectively reduce Flops and memory usage X k = Top-K(X|a), X p = X\X k Where the number k of token cropping decreases. As the network deepens, the number of tokens in subsequent layers gradually decreases, reducing the computational amount of subsequent layers. Dynamically select tokens while keeping the remaining number P of tokens as a multiple of 8 to adapt to GPU memory alignment and improve throughput.

4. The lightweight ViT based on image feature cropping and cloud-edge-end knowledge distillation according to claim 1, wherein: In Step 3, the methods of using ThreeAugment and screening ImageNet-1k based on uncertainty include: Preprocessing of data images. The model uses the Threeaugment data augmentation method for the images in the dataset: Let the input image be I ∈ R H×W×3 The preprocessing process can be expressed as a composite function The operations at each stage are as follows: Perform primary spatial transformation on the image I, randomly scale, crop, and flip during training: F crop-flip (I) = Flip p=0.5 (Resize(Crop(I, s ~ U(0.08, 1.0)))) where s is the cropping area ratio, adjusted to the target size S target = 224 × 224, During model testing, standardize the image to a fixed size: F resize-center (I) = CenterCrop 224 · Resize 256 (I) Perform secondary image enhancement operations on the image I: Grayscale: Exposure inversion: Gaussian blur: Finally, perform standardization and sampling work: Let the total number of samples be N and the number of GPUs be K. The index assignment satisfies:

5. The lightweight ViT based on image feature cropping and cloud-edge-end knowledge distillation according to claim 1, wherein, In Step 3, the knowledge distillation strategy of the cloud teacher model is as follows: While the edge device transmits highly uncertain samples to the cloud for processing, it uploads the device model parameters to the cloud as the benchmark model parameters θ for parameter update. Base , let the edge-side student model parameters be θ edge , select the pre-trained ViT-Large / 16 as the teacher model, load the model parameters from the pre-trained weights, freeze all layers, and only use them for forward inference to generate soft label probabilities. Guide the optimization of the lightweight student model through the soft labels and intermediate feature maps of the teacher model; Define the knowledge distillation loss function L total , comprehensively consider the cross-entropy loss and the KL divergence loss and perform a dynamic weighting operation: Assume that the true label is one-hot encoded \(y\in\{0,1\}\) C , and the predicted probability of the student model is \(p\) s =\text{softmax}(z s ). Then the cross-entropy loss is: The gradient of the parameter θ is: Let the soft label probability of the teacher model be p t = softmax(z t / T), and the predicted probability of the student model be p s = softmax(z s / T). Then the KL divergence loss is as follows: where T is used for softening the distribution, and the gradient with respect to the parameter θ is: The total loss function of knowledge distillation is the weighted sum of cross-entropy and KL divergence: The combined gradient formula is: The gradient with respect to the parameter θ after combination is: By minimizing generate the model parameters θ after knowledge distillation update new .

6. The lightweight ViT based on image feature cropping and cloud-edge-end knowledge distillation according to claim 1, wherein In step 3, the edge device transmits the model parameters and high-uncertainty images to the cloud. After knowledge distillation by the cloud teacher model, the parameters are dynamically compressed and updated to the edge device; Due to the limitations of data transmission bandwidth in most edge devices, it is unrealistic to transmit all samples collected by edge devices to the cloud for processing. To improve the training efficiency of transmitted samples as much as possible, uncertainty-based dataset sample sampling is performed on the dataset images for the input image x i For the same image x i Perform n independent forward propagations to obtain n sets of class probability distributions {p1(y|x i ), p2(y|x i ), …, p n (y|x i )}, calculate the probability mean, and measure the uncertainty of model prediction through the standard deviation: V unc The larger the value, the more uncertain the model's prediction of the image is; Set a predefined threshold θ to screen for high-uncertainty images x that satisfy V unc > θ, and only transmit the set of screened high-uncertainty images {x i |V i > θ} to the cloud server for subsequent model optimization; unc ​ To solve the problem of low parameter transmission efficiency when the cloud teacher model updates parameters during the knowledge distillation process and ensure the accuracy of the important parameters transmitted, the difference value Δθ between the updated parameter set θ new after knowledge distillation optimization and the device-side reference parameter set θ base is adaptively dynamically quantized: Δθ = θ new -θ base , θ q = Quantize(Δθ, Q), where Q is a dynamically calibrated quantization function that preferentially retains high-significance parameter updates during parameter quantization and uses a higher compression ratio for low-significance parameters; Transmit the compression parameter θ q to the edge device through the communication link, significantly reducing the bandwidth occupancy and transmission delay. The device receives the compression parameter θ q and directly superimposes it with the reference parameter θ base to generate the updated edge model parameter θ edge : θ edge = θ base + θ q .

Citation Information

Cited By

  • Knowledge distillation method and device, computer equipment and readable storage medium

    CN120745748A

  • Knowledge distillation method and device, computer device and readable storage medium

    CN120745748B

  • Knowledge distillation-based lightweight insect identification method and customs real-time universal equipment

    CN120997882A

  • Display device and server

    CN121262427A

  • ECA lightweight facial expression recognition method based on edge cloud collaboration

    CN121459408A