Multi-modal model lightweight method based on dual-channel sparse distillation

Through the hybrid expert architecture and the dual-channel sparse knowledge distillation mechanism, the computing resource limitation and semantic consistency problems of multimodal models when deploying at the edge end are solved, and efficient and accurate model lightweighting is achieved, with a reduced parameter volume of about 60%, and an improved accuracy of more than 5%.

CN120197666APending Publication Date: 2025-06-24HUNAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510366527.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When deploying the existing multimodal models at the edge end, it is difficult to meet the availability and real-time requirements due to the large number of parameters and high computing requirements. In addition, traditional model lightweight technology is difficult to maintain semantic consistency and accuracy in multimodal scenarios.

Method used

Using a hybrid expert architecture and a dual-channel sparse knowledge distillation mechanism, the cross-modal semantic integrity of the model compression process is achieved through a scalable hybrid expert architecture and a dual-channel sparse knowledge distillation framework.

Benefits of technology

The parameter quantity and calculation overhead of the model are significantly reduced, while maintaining efficient inference performance and accuracy. The parameter quantity is reduced by about 60%, and the accuracy rate is increased by more than 5% on cross-modal understanding tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197666A_ABST
    Figure CN120197666A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-channel sparse distillation-based multi-modal model lightweight method, which comprises the following five core steps of: connecting a pre-training visual encoder and a language model, and constructing a hybrid expert architecture; vision-language feature preliminary alignment is realized through feature mapping and cross-modal attention; two-channel knowledge migration is designed, an explicit channel uses adaptive KL divergence to align teacher and student model output distribution, and an implicit channel migrates feature knowledge through a cross-modal attention adapter; reasoning optimization training is carried out on the training set by constructing positive and negative samples, and student model learning is guided to distinguish high-quality output and low-quality output; top-k experts are dynamically selected during reasoning deployment, and expert output is weighted and aggregated through routing weights. According to the method, the calculation overhead is remarkably reduced while the expression ability of the model is maintained, the parameter quantity is reduced by about 60%, the accuracy rate of a cross-modal understanding task is improved by more than 5%, and the reasoning delay on edge equipment is controlled within 300ms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a multimodal model lightweight method based on architecture evolution and sparse distillation, which is particularly suitable for deploying high-efficiency multimodal artificial intelligence systems on terminal devices with limited computing resources. Background Art

[0002] With the widespread application of large multimodal models in tasks such as image and text understanding and cross-modal reasoning, their huge number of parameters and computing requirements have become the main obstacles to edge deployment. Taking a typical multimodal model as an example, the number of parameters generally exceeds 7 billion, and a single inference requires tens of GB of memory, which is difficult to meet the availability and real-time requirements of edge device deployment.

[0003] The current mainstream model lightweighting technologies mainly include the following three categories. The first is model pruning, which reduces the complexity of the model by removing redundant neurons or connections in the neural network. Typical methods include structured pruning (such as channel pruning) and unstructured pruning (such as weight pruning). However, pruning technology is easy to destroy cross-modal association paths in multimodal scenarios, resulting in decreased semantic consistency; the second is parameter quantization, which converts high-precision floating-point parameters into low-bit fixed-point representations, such as 8-integer quantization. Although it can effectively compress the model volume, the loss of accuracy in the quantization process will amplify the modal alignment error, especially in the visual-language feature interaction layer; the third is knowledge distillation, which transfers large model knowledge to small models through a teacher-student framework. Traditional methods mainly align the final output distribution, but ignore the cross-modal attention mode of the intermediate layer unique to multimodal models, which makes it difficult for student models to accurately learn the modal fusion mechanism. Therefore, an efficient and accurate multimodal model lightweight method is needed, which needs to have the following functions: Maintain cross-modal feature alignment: The lightweight process needs to maintain the semantic consistency of visual-language features to ensure that the model can still accurately understand multimodal inputs after compression; Adapt to the dynamic needs of different scenarios: The compression strategy needs to be able to flexibly adjust resource allocation according to the characteristics of different tasks, including but not limited to scenarios such as image-text question and answer, cross-modal retrieval, etc.; Efficiency and usability: The compressed model needs to be able to run efficiently on edge devices while maintaining high inference accuracy. Summary of the invention

[0004] Based on the multi-expert sparsity of the multimodal model architecture and the multi-layer coupling principle of knowledge transfer, this paper proposes a multimodal model lightweight method based on dual-channel sparse distillation. This method achieves the preservation of cross-modal semantic integrity during model compression by constructing a scalable hybrid expert architecture and a dual-channel sparse knowledge distillation mechanism.

[0005] The technical principles of lightweight multimodal models in the present invention include the following:

[0006] (1) The present invention adopts a mixture-of-experts architecture as the basic structure of the student model. This architecture consists of three main parts: a visual encoder, a vision-language adapter, and a language model based on mixture of experts. Among them, the visual encoder uses a pre-trained vision model to extract image features; the vision-language adapter is used to achieve feature space alignment; the language model adopts a sparse MoE feed-forward neural network structure, constructs expert modules by replicating N feed-forward networks, and introduces a linear routing layer to calculate the expert assignment probability. The routing probability is calculated by the Softmax function: r = Softmax(x·Wr), where x is the input feature and Wr is the parameter of the routing layer. Subsequently, the Top-k strategy is adopted to select the activated experts to ensure computational efficiency. This architecture design not only maintains the expressive ability of the model but also realizes the efficient utilization of computing resources.

[0007] (2) The present invention designs a two-channel sparse knowledge distillation framework, including a basic semantic alignment stage and a two-channel knowledge transfer stage. In the basic semantic alignment stage, a cross-modal feature adapter is constructed to achieve the alignment of visual and language feature spaces through a learnable feature mapping network, and a modal interaction attention mechanism is introduced to enhance cross-modal feature interaction. The optimization objective function of this stage is: where is the feature alignment loss, is the attention consistency loss, and α1 and α2 are weight coefficients for balancing the two loss terms. In the two-channel knowledge transfer stage, two knowledge transfer channels, explicit and implicit, are respectively constructed. The explicit channel realizes knowledge transfer by using the output distribution generated by the teacher model and adopting temperature-adaptive KL divergence. The optimization objective is: where, p T is the output probability distribution of the teacher model, pi is the output probability distribution of the i-th expert, is the routing weight of the i-th expert, τ(t) is the temperature coefficient, which is dynamically adjusted during training, and D KL calculates the KL divergence of the two probability distributions. The implicit channel establishes the mapping relationship of the cross-modal attention layer between the teacher model and the student model through a cross-modal attention adapter to ensure knowledge transfer at the feature level. The constructed attention mapping loss is: where, A T is the attention matrix of the teacher model, A i is the attention matrix of the i-th expert, W map is the learnable mapping matrix, |·| Fdenotes the Frobenius norm. The total loss function of dual-channel distillation is as follows:

[0008] (3) The present invention introduces an inference optimization mechanism. By constructing a training set of positive and negative sample pairs and using the teacher model as a reference benchmark, the student model is guided to learn to distinguish high-quality and low-quality outputs. The optimization objective function is: where x is the input, y+ is the positive sample (high-quality output), y- is the negative sample (low-quality output), π S and π T are the conditional probabilities of the student model and the teacher model respectively, β is the temperature coefficient, σ is the sigmoid function, and the goal is to make the probability ratio of the student model for positive samples higher and the probability ratio for negative samples lower in the inference stage. By dynamically selecting Top-k experts and activating the most relevant experts according to the input features, and finally aggregating the expert outputs by weighted routing weights, efficient and reliable inference is achieved. This strategy effectively improves the discriminative ability of the model and reduces the risk of the model generating hallucinations.

[0009] The technical solution for lightweighting the multi-modal model in the present invention includes the following steps:

[0010] Step 1: Construct a basic mixture-of-experts architecture. Connect the pre-trained visual encoder and the language model through a vision-language adapter, and build a MoE structure in the language model, including copying N feed-forward networks as expert modules and adding a linear routing layer for expert selection.

[0011] Step 2: Perform basic semantic alignment. Through a learnable feature mapping network and a modality interaction attention mechanism, optimize the objective function to achieve the initial alignment of the visual and language feature spaces.

[0012] Step 3: Implement dual-channel knowledge transfer. In the explicit channel, align the output distributions of the teacher model and the student model through temperature-adaptive KL divergence ; in the implicit channel, construct a cross-modal attention adapter to achieve knowledge transfer at the feature level.

[0012] Step 4: Perform inference optimization training. Construct a training set of positive and negative sample pairs, and optimize the objective function to guide the student model to learn to distinguish high-quality and low-quality outputs and improve the discriminative ability of the model.

[0013] Step 5: Deploy inference optimization. Dynamically select Top-k experts in the inference stage, and aggregate the expert outputs by weighted routing weights to achieve efficient inference. ​

[0014] In the above solution, further, the number of experts N in the Mixture of Experts (MoE) architecture is set to 4, and the number of experts k activated during each inference is set to 2, to achieve efficient utilization of computing resources.

[0015] Further, the feature mapping network in the basic semantic alignment stage adopts a multi-layer perceptron structure, and the multi-modal interaction attention mechanism adopts a multi-head attention structure, with the number of heads set to 4 - 8.

[0016] Further, the initial value τ0 of the temperature coefficient τ(t) is set to 4.0, and it decays exponentially with the number of training rounds t: τ(t) = τ0·exp(-λt), where λ is the decay rate, set to 0.001.

[0017] Further, the cross-modal attention adapter realizes bidirectional feature mapping through a reversible neural network, ensuring the reversibility and consistency of feature transfer.

[0018] Further, the positive and negative sample pairs in the inference optimization training are constructed by comparing the confidence scores output by the teacher model, and the confidence thresholds are set to 0.8 and 0.3 respectively.

[0019] The beneficial effects of the present invention are:

[0020] The present invention realizes the efficient lightweight of the model through the Mixture of Experts (MoE) architecture, significantly reducing the computational overhead while maintaining the model's expressive ability. Compared with traditional model compression methods, the MoE architecture of the present invention can dynamically activate relevant experts according to the input, achieving on-demand allocation of computing resources, and reducing the number of parameters by about 60% while maintaining similar performance.

[0021] The dual-channel distillation framework designed by the present invention ensures the effective transfer of knowledge. The explicit channel ensures the alignment of the output distribution through temperature-adaptive KL divergence, and the implicit channel maintains feature consistency through attention mapping. Experiments show that this framework improves the accuracy by more than 5% in cross-modal understanding tasks compared with traditional distillation methods.

[0022] The inference optimization strategy of the present invention significantly improves the discriminative ability of the model. By constructing a training set of positive and negative sample pairs, the student model can better identify high-quality outputs. It maintains a high inference efficiency, and at the same time, the inference latency on the edge device RK3588 is controlled within 300 ms. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is the distillation flow chart of the dual-channel sparse architecture for the multi-modal large model.

[0024] Figure 2 It is the design diagram of the MoE sparse architecture for the multi-modal large model.

[0025] Figure 3 It is a schematic diagram of a dual-channel distillation process.

[0026] Figure 4 It is a schematic diagram of inference optimization training. Specific implementation manners

[0027] Example 1: Model architecture construction Step 1: Visual encoder configuration. Use the pre-trained CLIP-ViT-B / 16 model as the visual encoder, adjust the input image size to 224×224 pixels, and the output feature dimension is 512 dimensions. Freeze the model parameters and only retain the last 3 Transformer blocks for fine-tuning. Step 2: Language model construction. Use the Transformer architecture with a parameter scale of 1.3B as the basic language model, and expand its feed-forward network into a MoE structure containing 4 experts. Each expert network contains two fully connected layers, the hidden layer dimension is 2048, and the activation function uses GeLU. Step 3: Routing mechanism implementation. The routing layer uses a linear transformation structure, the input dimension is the same as the hidden layer dimension of the language model (1024 dimensions), and the output dimension is the same as the number of experts (4 dimensions). Set k = 2 in the Top-k strategy, and the routing weight calculation method is: where is a learnable parameter matrix.

[0028] Example 2: Knowledge distillation training Step 1: Basic semantic alignment training. Use the image-text pair dataset (COCOCaptions) to perform preliminary alignment training on the student model, and optimize the objective function: where Adopt the cosine similarity loss: Adopt the attention matrix L2 loss: Training parameters: learning rate 1e-4, batch size 128, number of training epochs 5. Step 2: Dual-channel knowledge distillation. Use the Qwen-VL-Chat-7B model as the teacher model, keep the parameters frozen throughout the process, perform forward inference on the mixed dataset (COCO+VGQA) based on the teacher model, and generate the output probability distribution p T and the cross-modal attention matrix A T as the supervision signal to guide the training of the student model. The explicit channel uses temperature-adaptive KL divergence: Decay exponentially with the training round t: τ(t) = τ0·exp(-λt), where the initial temperature τ0 = 4.0 and the decay rate λ = 0.001. The implicit channel adopts the attention alignment loss: where W map is a learnable mapping matrix, and |·| F represents the Frobenius norm. The total loss function: Training parameters: learning rate 5e-5, batch size 64, training rounds 10.

[0029] Example 3: Inference Optimization Step 1: Construction of positive and negative samples. Use the teacher model to perform forward inference on the training set to generate candidate outputs. Select high-quality outputs with the teacher model confidence score > 0.8 as positive samples, and select low-quality outputs with the confidence score < 0.3 as negative samples. Construct 1 positive sample and 3 negative samples for each input sample to form a triple (x, y + , y - ). Step 2: Optimization objective setting. Construct the inference optimization loss function: where β ∈ [0.1, 1.0] is an adjustable temperature coefficient, and β = 0.5 is taken in the implementation, and σ is the sigmoid function. Step 3: Dynamically select the Top-k experts in the inference stage, and aggregate the expert outputs by weighted routing weights :

[0030] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multimodal model lightweight method based on dual-channel sparse distillation, characterized in that: The following steps are involved: Step 1: Build a basic hybrid expert architecture, connect the pre-trained visual encoder and the language model through the vision-language adapter, and build the MoE structure in the language model, including replicating N feedforward networks as expert modules and adding a linear routing layer for expert selection; Step 2: Perform basic semantic alignment and optimize the objective function through a learnable feature mapping network and modality interaction attention mechanism Achieve initial alignment of visual and language feature spaces, where is the feature alignment loss, is the attention consistency loss, α1 and α2 are the weight coefficients for balancing the two loss terms; Step 3: Implement dual-channel knowledge transfer. In the explicit channel, use temperature adaptive KL divergence Align the output distribution of the teacher model and the student model; in the implicit channel, build a cross-modal attention adapter to achieve feature-level knowledge transfer; Step 4: Perform inference optimization training, build a training set of positive and negative sample pairs, and optimize the objective function Guide student models to learn to distinguish between high-quality and low-quality outputs and improve the model's discrimination ability; Step 5: Deploy inference optimization, dynamically select Top-k experts in the inference stage, and use routing weights Weighted Aggregate Expert Output Achieve efficient reasoning.

2. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The number of experts N in the hybrid expert architecture is set to 4, and the number of experts k activated for each reasoning is set to 2, so as to achieve efficient utilization of computing resources.

3. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The feature mapping network of the basic semantic alignment stage adopts a multi-layer perceptron structure, and the modal interaction attention mechanism adopts a multi-head attention structure, with the number of heads set to 4-8.

4. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The initial value τ0 of the temperature coefficient τ(t) is set to 4.0, and decays exponentially with the training round t: τ(t)=τ0·exp(-λt), where λ is the decay rate, which is set to 0.

001.

5. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The cross-modal attention adapter realizes bidirectional feature mapping through a reversible neural network, ensuring the reversibility and consistency of feature migration.

6. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The positive and negative sample pairs in the inference optimization training are constructed by comparing the confidence scores output by the teacher model, and the confidence thresholds are set to 0.8 and 0.3 respectively.

7. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The implicit channel knowledge transfer in step 3 adopts attention alignment loss: Among them A T is the attention matrix of the teacher model, A i is the attention matrix of the i-th expert, Wmap is the learnable mapping matrix, and |·|F represents the Frobenius norm.

8. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The inference optimization objective function in step 4 is: Where x is the input, y+ is the positive sample, y- is the negative sample, π s and π T are the conditional probabilities of the student model and the teacher model, β is the temperature coefficient, and σ is the sigmoid function.

9. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The routing layer in step 1 adopts a linear transformation structure, and the routing probability is calculated by the Softmax function: r=Softmax(x·Wr), where x is the input feature and Wr is the routing layer parameter.

10. The multimodal model lightweight method based on dual-channel sparse distillation according to claim 1, characterized in that: The total loss function of the dual-channel knowledge migration in step 3 is: Where |θ|2 is the parameter regularization term.

Citation Information

Cited By

  • Text common sense reasoning method based on dynamic top-k selection expert model

    CN120875042A

  • A Textual Commonsense Reasoning Method Based on Dynamic Top-k Selection Expert Model

    CN120875042B