Multi-modal large-model heterogeneous multi-teacher distillation deployment method for low-resource equipment

Through the multimodal large model heterogeneous multi-teacher distillation deployment method, the problem of difficulty in deploying CLIP models on low-resource devices is solved, and the accuracy of efficient cross-modal knowledge transfer and image recognition classification is improved, which enhances the applicability and flexibility of the model.

CN120411997AActive Publication Date: 2025-08-01SHANDONG UNIV

Patent Information

Application Number
CN202510905603.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deploy multimodal large models, especially CLIP models, on low-resource devices, mainly due to its high computing and memory requirements, and the failure of existing methods to effectively utilize multi-teacher collaborative guidance and cross-modal semantic patterns, resulting in difficulty in deploying in clinical environments.

Method used

The multimodal large model heterogeneous multi-teacher distillation deployment method is adopted. Through image-text pairs as input, multimodal feature embedding is extracted, image modal and text modal weights are calculated, feature distillation and contrasting relationship distillation are performed, and multi-teacher knowledge is integrated into the student model using the adaptive knowledge fusion module.

Benefits of technology

It realizes efficient cross-modal knowledge transfer on low-resource devices, enhances the accuracy of image recognition classification and cross-domain generalization capabilities of models, and provides greater model design applicability and deployment flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411997A_ABST
    Figure CN120411997A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large-model heterogeneous multi-teacher distillation deployment method for low-resource equipment, and relates to the technical field of multi-modal knowledge distillation, and the method comprises the steps: taking an image-text pair as input, extracting image embedding and text embedding, and determining image-to-text and text-to-image comparison distribution; calculating and determining an image modal weight and a text modal weight; weighting image embedding and text embedding of the teacher model to obtain fusion image embedding and fusion text embedding, and weighting contrast distribution of the teacher model to obtain fusion distribution from image to text and from text to image; and carrying out characteristic distillation and comparative relationship distillation. According to the method, complementary multi-modal information is extracted from aligned images and texts, and the multi-modal information is distilled into a student model, so that efficient cross-modal knowledge transmission is realized, the ability of understanding cross-modal semantics is enhanced, and the accuracy of image recognition and classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal knowledge distillation, and particularly to a heterogeneous multi-teacher distillation deployment method for multimodal large models for low-resource devices. Background Art

[0002] Compared with traditional visual representation models that only rely on visual annotations, Contrastive Language-Image Pre-Training (CLIP) introduces text supervision and effectively injects rich semantic knowledge in natural language into the visual channel. By jointly training on large-scale image-text pairs, CLIP performs excellently in tasks such as zero-shot classification, cross-modal retrieval, and dense prediction.

[0003] Currently, there are studies applying contrastive learning methods to medical scenarios. By using global-local objectives to improve image-text alignment, but relying on large-scale paired datasets, and in the clinical environment, such datasets are usually very scarce. To solve this problem, MedCLIP (a multimodal deep learning framework based on model improvement) introduces pre-training on unpaired image-text data to improve the performance of downstream tasks such as disease classification. IMITATE (a machine learning method for training agents by imitating expert behavior) incorporates clinical priors and hierarchical supervision to enhance semantic representation. UniChest (a model based on the Transformer model architecture) decouples global and local semantics to better achieve cross-dataset generalization.

[0004] The above methods aim to learn robust feature representations in the pre-training stage through large-scale multimodal alignment. However, it is still challenging to deploy the CLIP model in the actual clinical environment. The full-size CLIP model using ViT-B or ViT-L as the backbone network contains hundreds of millions of parameters, requiring huge memory and computing resources, which limits its application in resource-constrained edge devices.

[0005] To solve this problem, there are studies introducing Knowledge Distillation (KD) to transfer knowledge from large teacher models to smaller student models. The multi-teacher distillation method is introduced to improve the generalization ability of the student model by aggregating the knowledge of multiple experts. DistillVLM (a distillation model based on visual language models) and contrastive relation distillation extend knowledge distillation to multimodal tasks by aligning cross-modal feature or relation structures.

[0006] Meanwhile, research has explored distilling the CLIP model for efficient deployment. TinyCLIP (a lightweight version of CLIP compressed based on knowledge distillation) introduced affinity imitation, which can preserve the relational structure but requires architecture alignment between the teacher and the student. CLIP-KD overcomes this limitation through a flexible multi-teacher framework, making it possible to distill between heterogeneous models. MedCLIP applies distillation in a cross-modal setting, jointly learning images and text to improve efficiency and semantic alignment.

[0007] However, current methods mainly focus on distilling knowledge from a single teacher model and are dedicated to unimodal tasks. Distilling multi-modal large models such as CLIP is still not effectively achievable, and the potential of multi-teacher collaborative guidance in enhancing the cross-domain generalization ability of student models is overlooked. In addition, these methods fail to utilize the cross-modal complementarity in the CLIP dual-encoder architecture and do not consider the text-image semantic patterns as the synchronization of different knowledge streams. Therefore, CLIP distillation for multi-teacher heterogeneous models remains a less studied topic. Summary of the Invention

[0008] To solve the above problems, the present invention proposes a heterogeneous multi-teacher distillation deployment method for multi-modal large models on low-resource devices, which extracts complementary multi-modal information from aligned images and text, distills the multi-modal information into the student model, realizes efficient cross-modal knowledge transfer, enhances the ability to understand cross-modal semantics, and improves the accuracy of image recognition and classification.

[0009] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a heterogeneous multi-teacher distillation deployment method for multi-modal large models on low-resource devices, including: In the computing cluster of the target cloud, obtain image-text pairs from a preset database based on the obtained problem domain, and determine the student model according to the target device, with a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination condition; Use the obtained image-text pairs as the input of the student model and multiple teacher models to obtain the image embeddings and text embeddings of each model. Respectively, taking the image embeddings and text embeddings as anchor points, determine the image-to-text contrast distribution and text-to-image contrast distribution of each model; According to the image embeddings and text embeddings of the student model, calculate the similarity scores between the student model and each teacher model respectively, and thereby determine the image modality weight and text modality weight; According to the image modality weight and text modality weight, the image embeddings and text embeddings of all teacher models are weighted summed to obtain the fused image embedding and fused text embedding. The contrast distributions of all teacher models are weighted aggregated to obtain the fused distribution of image to text and the fused distribution of text to image. Feature distillation is performed based on the image embedding and text embedding, as well as the fused image embedding and fused text embedding of the student model. Contrastive relationship distillation is performed based on the contrast distribution of the student model, as well as the fused distribution of image to text and the fused distribution of text to image, to complete the training of the student model and deploy the trained student model to the target device.

[0010] As an optional embodiment, Image embedding As anchor point, Image-to-text contrastive distribution of teacher models Comparative distribution of image to text with the student model They are: ; ; First Text embedding As anchor point, Text-to-image contrastive distribution of teacher models Comparative distribution of text to image with the student model They are: ; ; in, represents the dot product; and All are Teacher Model and student models Temperature scaling factor in ; For the Teacher Model No. text embeddings; For the Teacher Model No. text embeddings; Model for students No. text embeddings; Model for students No. text embeddings; For the Teacher Model No. Image embedding; For the Teacher Model No. Image embedding; Model for students No. Image embedding; Model for students No. Image embedding.

[0011] As an optional implementation, an independent projection layer is set in each teacher model to pass the projection function and Map image embedding and text embedding to shared dimensions respectively The feature space of the student model is normalized to a unit vector after projection, and the similarity score is calculated after the projection layer is processed.

[0012] As an alternative implementation, image embedding based on the student model or text embedding , calculate the student model and the Similarity scores between teacher models for: ; Determined modal weights for: ; in, For the first image, the image embedding or text embedding generated by the student model is consistent with the Similarity scores between image embeddings or text embeddings generated by the teacher models; For the first image, the image embedding or text embedding generated by the student model is consistent with the The similarity scores between the image embeddings or text embeddings generated by the teacher models, is the number of teacher models; is the sigmoid activation function; is the intermediate parameter; is the similarity weight matrix; is a scaling vector; m represents image modality or text modality; M represents image embedding or text embedding.

[0013] As an optional implementation, in the process of determining the modal weight, a weight regularization loss is introduced for: ;in, is an ideal uniform weight; is the size of the dataset; The weighted sum of the weight regularization loss of the image model and the weight regularization loss of the text modality gives the total weight regularization loss.

[0014] As an alternative implementation, during the training of the student model, the total training loss function includes the CLIP loss, the contrastive relation distillation loss, the feature distillation loss, and the total weight regularization loss; Among them, the contrastive relation distillation loss is: ; ; ; Among them, is the fusion distribution from image to text; is the fusion distribution from text to image; and are the contrastive distributions from image to text and from text to image of the student model; The feature distillation loss is: ; Among them, is the fused image embedding; is the fused text embedding; and are the image embedding and text embedding of the student model.

[0015] In a second aspect, the present invention provides a heterogeneous multi-teacher distillation deployment system for a multi-modal large model for low-resource devices, including: A deployment module, configured to obtain image-text pairs from a preset database based on the acquired problem domain in a computing cluster of a target cloud, and determine a student model according to the target device, with a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination condition; A feature extraction module, configured to use the acquired image-text pairs as the input of the student model and multiple teacher models, obtain the image embedding and text embedding of each model, and determine the contrastive distribution from image to text and the contrastive distribution from text to image of each model with the image embedding and text embedding as the anchor points respectively; A weight calculation module, configured to calculate the similarity scores between the student model and each teacher model respectively according to the image embedding and text embedding of the student model, and thereby determine the image modality weight and the text modality weight; A fusion module, configured to perform weighted summation on the image embeddings and text embeddings of all teacher models respectively according to the image modality weight and the text modality weight, to obtain a fused image embedding and a fused text embedding, and perform weighted aggregation on the contrast distributions of all teacher models, to obtain a fused distribution from image to text and a fused distribution from text to image; A distillation module, configured to perform feature distillation according to the image embedding and text embedding of the student model and the fused image embedding and the fused text embedding; perform contrast relationship distillation according to the contrast distribution of the student model and the fused distribution from image to text and the fused distribution from text to image, so as to complete the training of the student model, and deploy the trained student model to a target device.

[0016] In a third aspect, the present invention provides an electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0018] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a heterogeneous multi-teacher distillation and deployment method for multi-modal large models for low-resource devices, using image-text pairs as the input of the student model and multiple teacher models, so as to extract multi-modal feature embeddings and contrast distributions from image to text and from text to image; according to the multi-modal feature embeddings of the student model, determine the image modality weight and the text modality weight respectively, so as to perform weighted summation on the image embeddings and text embeddings of all teacher models respectively, to obtain a fused image embedding and a fused text embedding, and perform weighted aggregation on the contrast distributions of all teacher models, to obtain a fused distribution from image to text and a fused distribution from text to image, and then perform feature distillation and contrast relationship distillation to complete the training of the student model. The present invention extracts complementary multi-modal information from aligned images and texts, distills the multi-modal information into the student model, realizes efficient cross-modal knowledge transfer, enhances its ability to understand cross-modal semantics, and improves the accuracy of image recognition and classification.

[0020] The present invention designs an adaptive heterogeneous multi-teacher knowledge distillation framework, which utilizes multi-modal knowledge distillation for heterogeneous multi-teacher models. By instance-based weighted integration of knowledge from multiple teacher models, lightweight deployment is ensured. Compared with existing methods, the present invention can adaptively learn the importance of each teacher model for each instance, thereby generating integrated soft targets; decouples the student model from the teacher architecture, enabling the student model to operate independently of the teacher model architecture without being restricted by the structure of the teacher model, providing greater model design applicability and deployment flexibility; outperforms the baselines of single-teacher and multi-teacher average distillation in multiple classification benchmark tests, especially being more prominent in resource-constrained environments.

[0021] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.

[0023] Figure 1 Flowchart of the multi-modal large model heterogeneous multi-teacher distillation deployment method for low-resource devices provided in Embodiment 1 of the present invention; Figure 2 Schematic diagram of the multi-modal large model heterogeneous multi-teacher distillation deployment method for low-resource devices provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following further describes the present invention in conjunction with the drawings and embodiments.

[0025] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0027] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0028] Embodiment 1 As Figure 1 shown, this embodiment proposes a heterogeneous multi-teacher distillation deployment method for multi-modal large models for low-resource devices, including: In the computing cluster of the target cloud, obtain image-text pairs from a preset database based on the obtained problem domain, and determine a student model according to the target device, with a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination conditions; Use the obtained image-text pairs as the input of the student model and multiple teacher models to obtain the image embeddings and text embeddings of each model. Respectively, using the image embeddings and text embeddings as anchors, determine the image-to-text contrast distribution and text-to-image contrast distribution of each model; According to the image embeddings and text embeddings of the student model, calculate the similarity scores between the student model and each teacher model respectively, and thereby determine the image modality weight and text modality weight; According to the image modality weight and text modality weight, perform weighted summation on the image embeddings and text embeddings of all teacher models respectively to obtain fused image embeddings and fused text embeddings, and perform weighted aggregation on the contrast distributions of all teacher models to obtain the fused image-to-text distribution and text-to-image fused distribution; Perform feature distillation according to the image embeddings and text embeddings of the student model and the fused image embeddings and fused text embeddings; perform contrast relationship distillation according to the contrast distribution of the student model and the fused image-to-text distribution and text-to-image fused distribution, thereby completing the training of the student model, and deploying the trained student model to the target device.

[0029] The method proposed in this embodiment includes three core parts: (1) Feature extraction of the multi-teacher model to capture complementary cross-modal patterns from different expert models; (2) Adaptive feature fusion for heterogeneous knowledge integration through dynamic weight calibration; (3) Dual-objective knowledge distillation to optimize both image-text alignment and domain-specific semantic retention.

[0030] The method in this embodiment realizes efficient cross-modal knowledge transfer through a formalized heterogeneous feature extraction, instance-based teacher fusion, and dual-objective knowledge distillation mechanism. Using multiple pre-trained CLIP models as teacher models, complementary multimodal information is extracted from aligned images and texts and distilled into the student model to enhance its ability to understand cross-modal semantics. At the same time, to address the heterogeneity of structures and representations among teacher models, an Adapter module for adaptive knowledge fusion is proposed. This module evaluates the similarity between the student model and each teacher model in the image and text spaces, dynamically assigns importance weights to different teacher models, and the feature embeddings of the teacher models are fused in a weighted manner after being projected into a unified representation space to form the final distillation signal.

[0031] In this embodiment, in the computing cluster of the target cloud, based on the obtained problem domain, a target dataset including image-text pairs is determined from a preset database. At the same time, a student model is determined from a preset student model library using the configuration information of the target device, and a distillation model training framework is deployed; where the target device can be a device whose own computing power meets the preset low computing power (low resource) determination condition.

[0032] It should be noted that for the computing cluster of the target cloud, the computing cluster consists of multiple nodes, and each node is configured with multiple GPU (i.e., Graphics Processing Unit) cards and single or multiple high-performance CPUs. Before starting the distillation deployment task, the system will automatically collect the computing resource information of the cluster, including key parameters such as the CPU usage rate of each node, the GPU video memory usage, and the network bandwidth. This information provides certain information for the subsequent model distillation process.

[0033] When the target cloud receives the distillation deployment instruction, it first determines the problem domain for this distillation task. The determination of the problem domain can be achieved in various ways. For example, when the user initiates a distillation deployment request, they directly input the specific problem domain label in the user interface. After receiving the label, it is matched with the domain classification in the preset database to locate the target dataset. At the same time, it is necessary to obtain the configuration information of the target device. The target device refers to a device whose computing power meets the preset low-computing-power determination conditions. The collection of device configuration information can be achieved by installing a lightweight information collection tool on the device side. This information collection tool can automatically collect hardware parameters such as the processor model, memory size, and storage capacity of the target device and upload this information to the cloud in real time. After the target cloud receives the configuration information, it is matched with the model configuration requirements in the preset student model library to determine the target student model. Among them, the preset student model library can store various small-sized large models with different sizes and complexities, and these small-sized large models are optimized for different low-computing-power device configurations to ensure that they can run on the target device.

[0034] During the process of framework deployment, the computing resources can be reasonably allocated to the distillation model training framework according to the computing resource status of the current cluster. For example, for nodes with relatively tight computing resources, the system will give priority to deploying lightweight distillation model training modules to avoid training task failures due to insufficient resources. At the same time, the distillation model training framework can be initialized and configured, including setting training parameters, loading necessary dependency libraries, etc., to ensure that the framework can run normally.

[0035] The following combines Figure 2 the schematic diagram of, taking the power work scenario image and the corresponding description text of safety helmet wearing as an example, to introduce the method of this embodiment in detail.

[0036] In this embodiment, paired power work scenario images and the corresponding description texts of safety helmet wearing are extracted from the existing dataset, and preliminary data cleaning and preprocessing are performed to uniformly crop the photo size. The corresponding description texts are extracted, irrelevant information is removed, and format errors are corrected to ensure the quality and consistency of the input data. Finally, the dataset is randomly divided into a training set, a validation set, and a test set according to a certain ratio (such as 8:1:1). The validation set is used for hyperparameter tuning, and the test set is used to evaluate the generalization performance of the model.

[0037] In this embodiment, multiple pre-trained CLIP models are used as teacher models to establish a network architecture based on the CLIP model. Among them, the student model and each teacher model core are composed of a dual-modal encoder, that is, including an image encoder and a text encoder. The image encoder is used to extract visual features from the image, and the parallel text encoder is used to extract semantic information embedded in the corresponding text.

[0038] Specifically, given a set of images including paired image-text pairs as input, they are respectively input into the image encoder and text encoder of the student model and multiple teacher models. All models independently generate modality-specific feature embeddings, that is, the image embeddings and text embeddings generated by the th teacher model and the student model are and respectively; where is the th image; is the text corresponding to the th th teacher model for the th image generated image embedding; is the th teacher model for the th image corresponding text generated text embedding; is the th student model

[0039] To establish semantic alignment between images and texts, all models adopt the contrastive learning strategy introduced in CLIP to optimize the similarity of matching pairs while pulling apart the distance of non-matching pairs. Its basic principle is to maximize the similarity between matching image-text pairs while minimizing the similarity between non-matching pairs, thereby effectively synchronizing multi-modal representations into a unified representation domain.

[0040] Taking the th image embedding as the anchor point, the image-to-text contrastive distribution of the th teacher model and the image-to-text contrastive distribution of the student model are respectively defined as: where represents the dot product, which is used to measure similarity; and are the A teacher model and a student model with a learnable temperature scaling factor; used to index the contrast distribution; in practice, negative samples are drawn from ; For the th teacher model the th text embedding; For the th teacher model the th text embedding; For the student model the th text embedding; For the student model the th text embedding.

[0041] Similarly, taking the th text embedding as the anchor point, the text-to-image contrast distribution of the th teacher model and the text-to-image contrast distribution of the student model are respectively defined as: ; ; where, For the th teacher model the th image embedding; For the th teacher model the th image embedding; For the student model the th image embedding; For the student model the th image embedding.

[0042] Multi-modal representations and contrast patterns from multiple sources drive two complementary distillation paths, namely: Feature Distillation (FD) for embedding alignment and Contrastive Representation Distillation (CRD) for distribution matching. The implementation of dual-object knowledge distillation requires strategic fusion through the Adapter adapter module for adaptive knowledge fusion.

[0043] To effectively integrate image and text knowledge from multiple heterogeneous teacher models, an Adapter module is proposed. The main goal of this module is to dynamically aggregate multimodal features from multiple teacher models on a per-sample basis while maintaining semantic consistency with the student model's representation. The finally fused features serve as a unified supervision signal for knowledge distillation into the student model.

[0044] Specifically, it includes the following: (1) Feature projection and normalization.

[0045] Given the differences in feature dimensions and distributions among different teacher models, the Adapter sets independent projection layers for each teacher model to map its image embeddings and text embeddings into the feature space of the student model.

[0046] Specifically: Let and respectively represent the image embedding and text embedding generated by the -th teacher model, where is the output dimension of the -th teacher model. Through the corresponding projection functions and in the projection layer, the image embedding and text embedding are mapped into a shared space with dimension to obtain the projected image embedding and the projected text embedding : ; ; where ; if , the projection function degenerates to the identity mapping.

[0047] All projected features are normalized to unit vectors for scale-invariant similarity calculation.

[0048] (2) Adaptive weight calculation.

[0049] To evaluate the contribution of each teacher model to a given sample, a weighting mechanism based on the similarity between the output features of the teacher model and the student model is designed.

[0050] For the image modality, a set of learnable parameters is introduced: the similarity weight matrix and the scaling vector , where is the number of teacher models, Denotes the dimension of the student model's feature space.

[0051] Given the image embedding of the normalized student model , the similarity score between the student model and the -th teacher model is calculated as: ; ; where denotes the sigmoid activation function, which is used to stabilize the scaling factor; is an intermediate parameter.

[0052] Finally, the image modality weights are obtained by applying the softmax activation function to all teacher models: ; where is the similarity score between the image embedding generated by the student model and the image embedding generated by the -th teacher model for the -th image; is the similarity score between the image embedding generated by the student model and the image embedding generated by the -th teacher model for the -th image.

[0053] Similarly, for the text modality, a set of learnable parameters is introduced: the similarity weight matrix and the scaling vector , where is the number of teacher models, denotes the dimension of the student model's feature space.

[0054] Given the text embedding of the normalized student model , the similarity score between the student model and the -th teacher model is calculated as: ; ; where denotes the sigmoid activation function, which is used to stabilize the scaling factor; is an intermediate parameter.

[0055] Finally, the text modality weights are obtained by applying the softmax activation function to all teacher models: ; Among them, for the th image, the similarity score between the text embedding generated by the student model and the text embedding generated by the th teacher model; for the th image, the similarity score between the text embedding generated by the student model and the text embedding generated by the th teacher model.

[0056] Thus, the contribution of the teacher model is dynamically adjusted according to the similarity between the student model and the teacher model, realizing sample-based adaptive fusion.

[0057] (3) Weighted feature fusion.

[0058] According to the calculated image modality weight and text modality weight , the Adapter adapter performs weighted summation on the image embeddings and text embeddings generated by all teacher models to generate unified fused image embeddings and fused text embeddings : ; .

[0059] Normalize and , and align them with the and generated by the student model, and then use them for the feature distillation process.

[0060] Secondly, in addition to the feature embeddings, the Adapter adapter also performs weighted aggregation on the contrast distributions (logits) generated by each teacher model, thereby obtaining the image-to-text fused distribution and the text-to-image fused distribution for contrastive relation distillation.

[0061] (4) Regularized weight calculation.

[0062] To prevent the student model from overly relying on a specific teacher model during the feature fusion process, a weight regularization constraint is introduced to encourage a uniform weight distribution among all teacher models, promote comprehensive knowledge utilization, and reduce the risk of biased fusion dominated by some teacher models.

[0063] For the image modality, the weight regularization loss is calculated through the mean square error between the learned weight and the uniform target: ; Among them, is the number of teacher models, is the ideal uniform weight.

[0064] Similarly, the weight regularization loss of the text modality is obtained .

[0065] Thus, the weight regularization loss of the Adapter adapter is defined as: ; Among them, and are the weighting coefficients of the weight regularization loss of the image modality and the weight regularization loss of the text modality respectively, controlling the intensity of regularization.

[0066] By introducing weight regularization, the weight generation process of the Adapter adapter module becomes more stable, providing a more balanced guidance signal for the student model, thus improving the overall effect of multi-modal knowledge distillation.

[0067] In this embodiment, two distillation strategies, feature distillation and contrastive relation distillation, are introduced, and combined with the proposed Adapter adapter module, multi-teacher knowledge fusion and distillation are carried out using weighted teacher embeddings and contrastive distributions. The design details of each loss function are described below.

[0068] (1) Contrastive relation distillation.

[0069] The basis of the CLIP model is to maximize the similarity between paired image-text embeddings by comparing similarity distributions. For contrastive relation distillation, the main type of knowledge is the output-oriented contrastive distribution, which effectively captures the structured relationships between feature embeddings. Using contrastive relation distillation in multi-teacher distillation enables the student model to better replicate the structured semantic relationships from multiple teacher models, thus improving the quality of feature representations.

[0070] In this embodiment, the proposed contrastive relation distillation method uses the sample adaptive fusion weights generated by the Adapter adapter module to align the contrastive distribution of the student model with the weighted fusion image-to-text fusion distribution and text-to-image fusion distribution. By using the Kullback-Leibler (KL) divergence, it is ensured that the student model can learn robust cross-modal representations from the aggregated teacher knowledge.

[0071] Specifically: The image-to-text and text-to-image contrastive distributions generated by the th teacher model are respectively denoted as and , where It is a mini - batch dataset of paired image - text samples. The image modality weights of the Adapter module and the text modality weights , the calculated image - to - text fusion distribution and the text - to - image fusion distribution are respectively: ; .

[0072] By calculating the KL divergence between the contrastive distributions and of the student model and the and obtained after fusion, contrastive relation distillation is performed to align the contrastive distribution of the student model with the contrastive distribution of the fused teacher model. Thus, the contrastive relation distillation loss is constructed through the KL divergence, and the contrastive relation distillation loss is calculated as the average of the two - direction losses, and the expression is: ; ; ; where is the KL divergence between the image - to - text contrastive distribution of the student model and the image - to - text fusion distribution; is the KL divergence between the text - to - image contrastive distribution of the student model and the text - to - image fusion distribution.

[0073] (2)Feature distillation.

[0074] Feature distillation FD is introduced to further align the feature embeddings of the student model with those of the fused teacher model, thereby enhancing the fine - grained alignment ability of the student model in the representation space. FD complements CRD: while CRD focuses on aligning cross - modal semantic relations, FD focuses on the direct alignment of image and text feature embeddings. Intuitively, if the features of the student model can be closely aligned with those of the teacher model, the performance gap between the student model and the teacher model can be significantly reduced.

[0075] Specifically: the image embedding and text embedding of the student model are respectively represented as and , and have been normalized. The fused image embedding and fused text embedding generated by the Adapter module are represented as and , and have also been normalized accordingly.

[0076] Note that since the feature dimensions of different teacher models may not match those of the student model, the features of the teacher model have been mapped to the feature space of the student model through a projection layer to ensure dimensional consistency.

[0077] By calculating the image embeddings of the student model and text embeddings with the mean squared error loss of the fused image embeddings and fused text embeddings feature distillation is performed; thus the feature distillation loss is defined as: .

[0078] By minimizing the feature distillation loss , the student model can directly learn fine-grained information from the fused teacher model features, thereby enhancing its multimodal representation ability.

[0079] (3) Total loss function.

[0080] To jointly train the student model, the CLIP loss is combined with the knowledge distillation loss to form the total loss function. The CLIP loss is used to train the student model to learn the alignment between image and text representations.

[0081] Taking the contrastive loss of image embeddings to text embeddings as an example, it is defined as: ; where is a learnable temperature parameter for scaling the similarity; represents the mini-batch dataset size.

[0082] Similarly, the contrastive loss of text embeddings to image embeddings is obtained.

[0083] Thus, the CLIP loss combines the contrastive loss of image embeddings to text embeddings and text embeddings to image embeddings, expressed as follows: .

[0084] The knowledge distillation loss includes contrastive relation distillation, feature distillation, and Adapter adapter weight regularization loss, which are respectively: the contrastive relation distillation loss for aligning cross-modal semantic relations, the feature distillation loss for aligning the feature embeddings of the student model and the fused teacher model, and the Adapter adapter weight regularization loss for regularizing the weight generation process. The knowledge distillation loss Constructed by weighted combination of the above losses, defined as: ; Where, and are hyperparameters that respectively adjust the weights of the contrastive relation distillation loss and the feature distillation loss, balancing the contributions of different loss terms to the training of the student model.

[0085] By jointly optimizing the above losses, the student model can effectively learn multimodal knowledge while maintaining the stability of feature fusion.

[0086] Use the Adam optimizer to update the parameters of the student model, and continuously iterate until the model converges or reaches the preset number of epochs. At the end of each epoch, evaluate the model performance on the validation set and save the model parameters with the best performance for subsequent testing.

[0087] In this embodiment, the process of deploying the trained student model includes: (1) Model encapsulation: Encapsulate the trained student model into a standardized prediction interface and provide necessary documentation.

[0088] (2) System integration: Integrate the model into the power system, interface with other modules such as the database and Web front-end display to ensure the smooth flow of data streams.

[0089] (3) Online testing: Test the actual performance of the model in a small-scale real scenario and fine-tune and optimize the model according to the feedback.

[0090] (4) Model go live: Officially deploy the model online, which can assist operation and maintenance personnel in image recognition and improve the analysis efficiency. At the same time, continuously monitor the online recognition results of the model to ensure its stability and rationality.

[0091] (5) Model update: Regularly retrain the model using newly collected power work scenario data to adapt to new requirements and changes in data distribution.

[0092] It should be noted that the acquisition of all data is based on compliance with laws and regulations and user consent, and the data is legally applied.

[0093] Embodiment 2 This embodiment provides a heterogeneous multi-teacher distillation deployment system for multimodal large models for low-resource devices, including: A deployment module, configured to obtain image-text pairs from a preset database based on the acquired problem domain in a computing cluster of a target cloud, determine a student model according to the target device, and use a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination conditions; A feature extraction module, configured to use the obtained image-text pairs as the input of a student model and multiple teacher models, obtain the image embeddings and text embeddings of each model, and respectively use the image embeddings and text embeddings as anchors to determine the image-to-text contrast distribution and text-to-image contrast distribution of each model; A weight calculation module, configured to calculate the similarity scores between the student model and each teacher model according to the image embeddings and text embeddings of the student model, so as to determine the image modality weight and text modality weight; A fusion module, configured to perform weighted summation on the image embeddings and text embeddings of all teacher models respectively according to the image modality weight and text modality weight, obtain the fused image embeddings and fused text embeddings, and perform weighted aggregation on the contrast distributions of all teacher models to obtain the image-to-text fused distribution and text-to-image fused distribution; A distillation module, configured to perform feature distillation according to the image embeddings and text embeddings of the student model and the fused image embeddings and fused text embeddings; perform contrast relationship distillation according to the contrast distribution of the student model and the image-to-text fused distribution and text-to-image fused distribution, so as to complete the training of the student model and deploy the trained student model to the target device.

[0094] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0095] In more embodiments, there is also provided: An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0096] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0097] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.

[0098] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the method described in Embodiment 1.

[0099] The method in Embodiment 1 can be directly implemented by a hardware processor or by a combination of hardware and software modules in the processor. The software modules can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0100] A computer program product comprising a computer program, which, when executed by a processor, implements the method described in Embodiment 1.

[0101] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.

[0102] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0103] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0105] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A heterogeneous multi-teacher distillation deployment method for multi-modal large models on low-resource devices, characterized in that, including: In the computing cluster of the target cloud, obtain image-text pairs from a preset database based on the obtained problem domain, determine a student model according to the target device, and use a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination condition; Use the obtained image-text pairs as the input of the student model and multiple teacher models to obtain the image embeddings and text embeddings of each model. Respectively using the image embeddings and text embeddings as anchors, determine the image-to-text contrast distribution and text-to-image contrast distribution of each model; According to the image embeddings and text embeddings of the student model, calculate the similarity scores between the student model and each teacher model respectively, and thereby determine the image modality weight and text modality weight; According to the image modality weight and text modality weight, perform weighted summation on the image embeddings and text embeddings of all teacher models respectively to obtain fused image embeddings and fused text embeddings, and perform weighted aggregation on the contrast distributions of all teacher models to obtain the image-to-text fused distribution and text-to-image fused distribution; According to the image embeddings and text embeddings of the student model and the fused image embeddings and fused text embeddings, perform feature distillation; according to the contrast distribution of the student model and the image-to-text fused distribution and text-to-image fused distribution, perform contrast relationship distillation, thereby completing the training of the student model and deploying the trained student model to the target device.

2. The heterogeneous multi-teacher distillation deployment method for multi-modal large models for low-resource devices according to claim 1, characterized in that, Using the th image embedding as the anchor point, the image-to-text contrast distributions of the th teacher model and the image-to-text contrast distributions of the student model are respectively: ; ; Taking the th text embedding as the anchor point, the text-to-image comparison distributions of the th teacher model and the student model are respectively: ; ; in, represents the dot product; and All are Teacher Model and student models Temperature scaling factor in ; For the Teacher Model No. text embeddings; For the Teacher Model No. text embeddings; Model for students No. text embeddings; Model for students No. text embeddings; For the Teacher Model No. Image embedding; For the Teacher Model No. Image embedding; Model for students No. Image embedding; Model for students No. Image embedding.

3. The heterogeneous multi-teacher distillation deployment method for multi-modal large models targeting low-resource devices according to claim 1, wherein An independent projection layer is separately set in each teacher model, which is used to map the image embedding and the text embedding into the feature space of the student model with a shared dimension of and respectively through the projection functions . After projection, they are normalized to unit vectors, and finally, the similarity scores are calculated after the processing of the projection layer is completed.

4. The heterogeneous multi-teacher distillation deployment method for multimodal large models targeting low-resource devices according to claim 1, characterized in that, Image embedding according to the student model or text embedding , calculate the similarity score between the student model and the th teacher model as: : ; Determined modal weights are as follows: ; Among them, is the similarity score between the image embedding or text embedding generated by the student model and the image embedding or text embedding generated by the th teacher model for the th image; is the similarity score between the image embedding or text embedding generated by the student model and the image embedding or text embedding generated by the th teacher model for the th image, is the number of teacher models; is the sigmoid activation function; is an intermediate parameter; is the similarity weight matrix; is the scaling vector; m represents the image modality or text modality; M represents the image embedding or text embedding.

5. The heterogeneous multi-teacher distillation deployment method for multi-modal large models targeting low-resource devices according to claim 4, characterized in that, During the process of determining the modal weights, a weight regularization loss is introduced It is as follows: ; Among them, is the ideal uniform weight; is the dataset size; Weight the weight regularization loss of the image model and the weight regularization loss of the text modality to obtain the total weight regularization loss.

6. The heterogeneous multi-teacher distillation deployment method for multi-modal large models for low-resource devices according to claim 5, characterized in that During the training of the student model, the total training loss function includes the CLIP loss, contrast relationship distillation loss, feature distillation loss, and total weight regularization loss; Among them, the contrastive relation distillation loss is as follows: ; ; ; Among them, is the fusion distribution from image to text; is the fusion distribution from text to image; and are the contrast distributions from image to text and from text to image of the student model; Feature distillation loss is as follows: wherein, is the fused image embedding; is the fused text embedding; and are the image embedding and text embedding of the student model respectively.

7. A heterogeneous multi-teacher distillation deployment system for multimodal large models on low-resource devices, characterized in that including: A deployment module, configured to obtain image-text pairs from a preset database in the computing cluster of the target cloud based on the obtained problem domain, determine a student model according to the target device, and use a preset large language model as the teacher model; the target device is a device whose own computing power meets the preset low-resource determination condition; A feature extraction module, configured to use the obtained image-text pairs as the input of the student model and multiple teacher models to obtain the image embeddings and text embeddings of each model. Respectively using the image embeddings and text embeddings as anchors, determine the image-to-text contrast distribution and text-to-image contrast distribution of each model; A weight calculation module, configured to calculate the similarity scores between the student model and each teacher model respectively according to the image embeddings and text embeddings of the student model, and thereby determine the image modality weight and text modality weight; A fusion module, configured to perform weighted summation on the image embeddings and text embeddings of all teacher models respectively according to the image modality weight and text modality weight to obtain fused image embeddings and fused text embeddings, and perform weighted aggregation on the contrast distributions of all teacher models to obtain the image-to-text fused distribution and text-to-image fused distribution; A distillation module, configured to perform feature distillation based on the image embedding and text embedding of the student model, as well as the fused image embedding and fused text embedding; and perform contrast relationship distillation based on the contrast distribution of the student model, as well as the fused distribution from image to text and the fused distribution from text to image, so as to complete the training of the student model and deploy the trained student model to a target device.

8. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method according to any one of claims 1-6 is completed.

9. A computer-readable storage medium, characterized in that, It is used to store computer instructions. When the computer instructions are executed by the processor, the method according to any one of claims 1-6 is completed.

10. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Multi-modal joint representation learning method and system based on variational distillation

    CN114841335A

  • Knowledge distillation method and system based on multi-teacher multi-modal model

    CN117669693A

  • Image text generation method and device based on deep learning

    CN118397147A

  • Multi-modal style migration method, system and equipment based on knowledge distillation

    CN119741187A

  • Multi-modal sentiment analysis method and system based on knowledge distillation and dynamic fusion mechanism

    CN120046695A

Cited By

  • Internet of Things heterogeneous equipment federal learning method based on cross-modal knowledge distillation

    CN121724107A

  • An internet of things heterogeneous device federated learning method based on cross-modal knowledge distillation

    CN121724107B