Multimodal Joint Representation Learning Method and System Based on Variational Distillation
By using the variable distillation method in multimodal joint representation learning, the joint distillation training of student model and teacher model is solved, and the problems of missing and forgetfulness of the modal unified distillation method in the existing technology are achieved, and performance improvement and computational cost reduction on different modal data sets are achieved.
Patent Information
- Application Number
- CN202210062288.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-19
AI Technical Summary
The existing technology lacks a modal unified distillation method, which cannot effectively solve the forgetfulness problem caused by joint training of multiple modals, and requires a large number of additional negative samples, greatly increasing the calculation cost.
Using a multimodal joint representation learning method based on variable distillation, the student model and teacher model are deployed, and the correlation between the outputs of the student model and the teacher model is characterized by using variational mutual information, and the joint distillation training is performed using the distillation loss function, which solves the forgetfulness problem and reduces the demand for negative samples.
It surpasses the existing benchmark model on different modal data sets, reduces the information loss of the teacher model, does not require a large number of negative samples to participate in the calculation, and is simple and effective, solving the forgetfulness problem caused by multiple modal distillation.
Smart Images

Figure CN114841335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal distillation, and in particular to a multimodal joint representation learning method and system based on variational distillation. Background Art
[0002] Large-scale pre-trained models, such as BERT, GPT, and RoBERTa in the text modality, or ResNet, BiT, ViT, etc. in the image modality, have brought revolutionary progress in different modality fields. However, as the scale of pre-trained models becomes larger and larger, it becomes increasingly challenging to deploy them in resource-constrained environments. Therefore, these model compression methods that reduce the scale of pre-trained models and preserve most of their performance have attracted more and more attention.
[0003] In the field of text modality, PKD was an early exploration, which is very simple and effective. It mainly compresses the BERT model in the fine-tuning stage. Subsequently, DistillBERT, TinyBERT, and MobileBERT perform task-agnostic efficient knowledge distillation on the intermediate layer information of the BERT model using KL divergence or L2 loss function during the training stage. CoDIR distills RoBERTa based on contrastive learning during the training stage and achieves better performance. In the field of image modality, FitNet fits the outputs of the teacher model and the student model on a specific task dataset. ViD uses a heteroscedastic Gaussian distribution to replace the sample distribution to calculate the mutual information of the output feature maps of the teacher and student models. DeiT adds a distillation token different from the classification token and trains and fits from two different perspectives. CRD uses contrastive learning and uses a large number of negative samples to improve the upper bound of the mutual information between the outputs of the teacher model and the student model.
[0004] Currently, the distillation in the single-modal text field and image field is relatively mature, but there are few unified distillation frameworks for the text modality and the image modality. Considering the traditional method of fitting the probability distributions of the outputs of the teacher model and the student model through KL divergence, or using the L2 loss function to fit the representation vectors of the teacher model and the student model. Although these methods can also reduce the output differences between the teacher model and the student model, these methods have the following defects. For example, for the L2 loss function, first, dimension transformation is required, which will lose some information. Second, only the relationship between the corresponding numerical values of the representation vectors is considered, while the overall information is ignored. The currently widely used contrastive distillation also improves the upper bound of the mutual information between the outputs of the teacher model and the student model. The contrastive distillation method requires a large number of negative samples compared with other methods, thus increasing the training loss, especially doubling the training cost in multiple modalities, and is not suitable for resource-constrained situations. On the other hand, serious forgetting problems will occur during multi-modal distillation. For example, distilling text information first and then image information may cause the encoder to lose most of its text encoding ability.
[0005] Therefore, there is currently no modal unified distillation method, which cannot solve the problem of forgetting caused by joint training of multiple modalities, and requires a large number of additional negative samples, greatly increasing the computational cost. Summary of the Invention
[0006] To this end, the technical problem to be solved by the present invention is to overcome the problems existing in the prior art, and propose a multi-modal joint representation learning method and system based on variational distillation, which solves the problem of the lack of modal unified distillation method in the prior art, and surpasses the existing benchmark models on different modal datasets.
[0007] To solve the above technical problem, the present invention provides a multi-modal joint representation learning method based on variational distillation, including the following steps:
[0008] Deploy a student model and a teacher model. The teacher model includes a text teacher model and an image teacher model. The student model includes a multi-modal data unification module. Input the original multi-modal data, where the original multi-modal data includes original text modal data and original image modal data. Input the original text modal data and the original image modal data into the multi-modal data unification module to obtain text modal inputs and image modal inputs with the same input form, and perform a normalization operation on the text modal inputs and the image modal inputs.
[0009] The student model includes a modal joint representation module. Input the normalized text modal inputs and image modal inputs into the modal joint representation module respectively to obtain the text output and image output of the student model. At the same time, input the original text modal data and the original image modal data into the text teacher model and the image teacher model respectively to obtain the text output and image output of the teacher model.
[0010] Use variational mutual information to characterize the correlation between the corresponding text outputs and image outputs of the student model and the teacher model, and jointly distill and train the text outputs and image outputs using a distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model.
[0011] In an embodiment of the present invention, the multi-modal data unification module is deployed at the front end of the modal joint representation module, and the multi-modal data unification module is used to organize the original text modal data and the original image modal data into the same input form to obtain text modal inputs and image modal inputs.
[0012] In an embodiment of the present invention, organizing the original text modal data and the original image modal data into the same input form to obtain text modal inputs and image modal inputs includes:
[0013] Add the [CLS] symbol and the [SEP] symbol to the original text modal data. At the same time, add the [DIS] symbol at the end of each sentence in the original text modal data, and obtain the text modal input through the word vector matrix;
[0014] Segment the original image modal data into several image patches, stretch each image patch into a one-dimensional vector, add the [CLS] symbol and the [DIS] symbol at the beginning and end positions of the one-dimensional vector, and obtain the image modal input with the same form as the text modal input through dimensional scaling.
[0015] In one embodiment of the present invention, the modal joint representation module includes a MobileBERT model, and the MobileBERT model includes 24 layers of transformer models, and a linear layer is added to each layer of the transformer model.
[0016] In one embodiment of the present invention, the distillation loss function is the sum of the loss function of the text teacher model and the loss function of the image teacher model.
[0017] In addition, the present invention also provides a multi-modal joint representation learning system based on variational distillation, including:
[0018] A student model, the student model includes a multi-modal data unification module and a modal joint representation module, and inputs the original multi-modal data, where the original multi-modal data includes original text modal data and original image modal data. Input the original text modal data and the original image modal data into the multi-modal data unification module to obtain text modal input and image modal input with the same input form, and perform a normalization operation on the text modal input and the image modal input. Input the normalized text modal input and image modal input into the modal joint representation module respectively to obtain the text output and image output of the student model;
[0019] A teacher model, the teacher model includes a text teacher model and an image teacher model, and inputs the original text modal data and the original image modal data into the text teacher model and the image teacher model respectively to obtain the text output and image output of the teacher model;
[0020] A modal unified distillation module, which is used to characterize the correlation between the corresponding text output and image output of the student model and the teacher model by using variational mutual information, and jointly distill and train the text output and the image output by using the distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model.
[0021] In one embodiment of the present invention, the multimodal data unification module is deployed at the front end of the modal joint representation module, and the original text modal data and the original image modal data are organized into the same input form by using the multimodal data unification module to obtain the text modal input and the image modal input.
[0022] In one embodiment of the present invention, the multimodal data unification module includes:
[0023] A text modal data organization sub-module, which is used to add [CLS] symbols and [SEP] symbols to the text modal data, and at the same time add [DIS] symbols at the end of the sentences in the text modal data, and obtain the text modal input through the word vector matrix;
[0024] An image modal data organization sub-module, which is used to divide the image modal data into several picture blocks, stretch each picture block into a one-dimensional vector, add [CLS] symbols and [DIS] symbols at the beginning and end positions of the one-dimensional vector, and obtain an image modal input with the same form as the text modal input through dimension scaling.
[0025] Moreover, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned method are implemented.
[0026] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned method are implemented.
[0027] The above technical solutions of the present invention have the following advantages compared with the prior art:
[0028] 1. In view of the problem that the prior art lacks a modal unified distillation method, the present invention proposes a multimodal joint representation learning method and system based on variational distillation, which exceeds the existing benchmark models on different modal data sets;
[0029] 2. The present invention adopts variational mutual information angle distillation, which not only greatly reduces the information loss of the teacher model, but also does not require a large number of negative samples to participate in the calculation, and has simplicity and effectiveness;
[0030] 3. The present invention adopts the method of joint distillation to solve the forgetting problem caused by multiple modal distillations. Description of the Drawings
[0031] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below according to the specific embodiments of the present invention in conjunction with the drawings.
[0032] Figure 1 It is a schematic flowchart of the multi-modal joint representation learning method based on variational distillation of the present invention.
[0033] Figure 2 It is a schematic framework diagram of the modality unified distillation module in the multi-modal joint representation learning system based on variational distillation of the present invention. Detailed implementation manners
[0034] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention.
[0035] Embodiment 1
[0036] Please refer to Figure 1 and 2 As shown, this embodiment provides a multi-modal joint representation learning method based on variational distillation, including the following steps:
[0037] S1: Deploy a student model and a teacher model. The teacher model includes a text teacher model and an image teacher model. The student model includes a multi-modal data unification module. Input the original multi-modal data, where the original multi-modal data includes original text modality data and original image modality data. Input the original text modality data and the original image modality data into the multi-modal data unification module to obtain text modality inputs and image modality inputs with the same input form, and perform a normalization operation on the text modality inputs and the image modality inputs;
[0038] S2: The student model includes a modality joint representation module. Input the normalized text modality inputs and image modality inputs into the modality joint representation module respectively to obtain the text output and the image output of the student model. At the same time, input the original text modality data and the original image modality data into the text teacher model and the image teacher model respectively to obtain the text output and the image output of the teacher model;
[0039] S3: Use variational mutual information to characterize the correlation between the corresponding text outputs and image outputs of the student model and the teacher model, and perform joint distillation training on the text outputs and the image outputs using a distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model.
[0040] In a multi-modal joint representation learning method based on variational distillation disclosed by the present invention, the present invention proposes a multi-modal joint representation learning method and system based on variational distillation for the problem that the existing technology lacks a modality unified distillation method, and surpasses the existing benchmark models on different modality datasets.
[0041] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, for S1 of the above embodiment, the student model includes a multi-modal data unification module, and the multi-modal data unification module is deployed at the front end of the modal joint representation module. The original text modal data and the original image modal data are arranged into the same input form by using the multi-modal data unification module to obtain text modal input and image modal input.
[0042] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, for S1 of the above embodiment, arranging the original text modal data and the original image modal data into the same input form to obtain text modal input and image modal input includes two aspects. On the one hand, [CLS] symbol and [SEP] symbol are added to the text modal data, and [DIS] symbol is added at the end of the sentence in the text modal data, and the text modal input is obtained through the word vector matrix; on the other hand, the image modal data is segmented into several picture blocks, each picture block is stretched into a one-dimensional vector, [CLS] symbol and [DIS] symbol are added at the beginning position and the end position of the one-dimensional vector, and the image modal input with the same form as the text modal input is obtained through dimension scaling.
[0043] Specifically, on the one hand, for the original text modal data D with length L l , [CLS] symbol and [SEP] symbol are added to the original text modal data, and [DIS] symbol is added at the end of the sentence in the original text modal data for bilateral distillation, which improves the performance and accelerates the fitting while; the corresponding word segmentation serial numbers are obtained according to BPE, and the final input text word vectors are obtained through the word vector matrix with dimension d where L′ = L + 3. On the other hand, since the representation forms of text and image are different, it is difficult to directly unify text and image. Therefore, the present invention adopts the method of segmenting the image into several picture blocks for processing so as to generate the same input form as the text word vectors. For the original image input data D t , first, it is scaled to the size of 256×256×3, and then the image is segmented into 256 picture blocks according to the size of the picture block of 16×16×3 Each picture block is stretched into a one-dimensional vector to obtain Then, [CLS] symbol and [DIS] symbol are also added at the beginning position and the end position, and the dimension scaling is performed through the final linear layer to obtain the same form as the text input: Since there are some differences in the distributions of text and image data, resulting in large fluctuations in numerical values, the data is finally normalized. The present invention unifies the input forms and distributions of text modality and image modality, which is convenient for the subsequent processing of the modal joint layer.
[0044] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, for S2 of the above embodiment, the modal joint representation module includes a MobileBERT model, and the MobileBERT model includes 24 layers of transformer models. A linear layer is added to each layer of the transformer model, so the parameter scale of the transformer in MobileBERT is small. For an input of length N The output of each layer of the transformer can be obtained as To facilitate distillation, the feature representation corresponding to the [CLS] symbol is taken As the feature of the student output representation used in training, the calculation formula is
[0045] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, for S3 of the above embodiment, the distillation loss function is the sum of the loss function of the text teacher model and the loss function of the image teacher model.
[0046] Specifically, for two different modal information of text and image, the present invention unifies the distillation method from the perspective of mutual information. It will be introduced in detail in four parts: variational mutual information, knowledge distillation, summary of the distillation process, and experimental analysis.
[0047] (3.1) Variational mutual information
[0048] Mutual information can represent reducing the uncertainty of a random variable by knowing another random variable. In representation learning, it can be further used to measure the correlation between different representations. From the perspective of information theory, knowledge transfer is a process of maintaining a high mutual information between the corresponding outputs of the teacher model and the student model, and can be regarded as a process of the teacher model retaining knowledge in the student model. Given a pair of random variables (X, Y), the mutual information between X and Y can be defined as:
[0049]
[0050] Among them, \(H(X)\) is the entropy of \(X\), and \(H(X|Y)\) is the conditional entropy of the joint distribution \(P(X, Y)\). Due to the difficulty of directly calculating the joint distribution, the present invention aggregates the results of the input distributions on each layer into a joint distribution. Mutual information can adapt to distributions at a higher semantic level. However, due to the great difficulty in precisely calculating mutual information, it is difficult to maximize mutual information. One solution is to use contrastive learning to make the samples closer to the positive samples and farther from the negative samples, which can increase the lower bound of mutual information. This method is effective in the case of a large number of negative samples. However, the computational cost of this method is relatively high. Therefore, the present invention uses the variational lower bound to approximately calculate the mutual information \(I(X, Y)\). According to VID 13] which believes it is difficult to calculate the distribution \(p(X|Y)\), the present invention uses the variational distribution \(q(X|Y)\) to approximate \(p(X|Y)\). Therefore, the following formula can be further calculated:
[0051]
[0052] On the one hand, due to the non-negativity of the KL divergence, the present invention can obtain the last inequality. On the other hand, \(H(X)\) is a constant, so the present invention only needs to calculate Furthermore, since a single Gaussian distribution is too simple to approximate some complex distributions, the present invention uses a mixture of Gaussian distributions \(q(X|Y)\), \(\log q(x|y) can be further calculated as follows:
[0053]
[0054] where \(y n is the scalar component of \(y\) at the \(n\)-th subscript position, \(\mu n (x)\) is the output of the encoder network \(\mu(.)\) composed of transformers, and it is ensured to be positive through the softplus function. Where \(\sigma c is a parameter to be optimized, \(\epsilon\) is an extremely small constant and is set to be greater than 0 to ensure that the variance is positive, and constant is a constant. The final loss function for mutual information distillation is as follows:
[0055]
[0056] (3.2) Knowledge Distillation
[0057] In a preferred embodiment, BERT large is used as the teacher model for the text modality, and ResNet 152 is used as the teacher model for the image modality. Let be BERT large and ResNet152 and the output of the i-th hidden layer of MXBERT, where is BERT large and the feature vector corresponding to [CLS] in the MXBERT output feature representation, is the feature vector corresponding to [DIS] in the MXBERT output feature representation.
[0058] For the text modality, the loss function is:
[0059]
[0060] where α(i) is a coefficient function, which sets different weights for each layer and increases with the increase of the layer number, is a multi-layer non-linear transformation function used to achieve the variational effect.
[0061] For the image modality, the loss function is:
[0062]
[0063] i′ = f(i)
[0064]
[0065] where β(i) is a coefficient function, which sets different weights for each layer and increases with the increase of the layer number. Since the number of layers of ResNet and MXBERT is different, f(.) is set as a function that selects the student layer corresponding to the teacher model layer, and φ(.) is a multi-layer linear transformation, which mainly converts the three-dimensional feature map output by ResNet into a one-dimensional feature vector.
[0066] Because the output value ranges of the teacher models are the same, to solve the forgetting problem, the present invention performs joint training on the distillation of the two modalities, and the final loss function of the entire distillation is defined as follows:
[0067] L dis = L l + L T .
[0068] (3.3) Summary of the distillation process
[0069] Given two teacher models: the text teacher model BERT for the text modality large and the image teacher model ResNet for the image modality 152 , given the student model: MXBERT. The purpose of modality-unified distillation is to enable MXBERT with fewer parameters to learn the capabilities of the two teacher models simultaneously, so that MXBERT has the encoding capabilities of text and images.
[0070] The distillation process is as follows: Given a text output and an image output. The text output passes through BERT large and MXBERT respectively, and the outputs of the middle layers of the two models are fitted. The fitting method is L l (The meaning of fitting is to make two variables as equal as possible. In this aspect, it is to reduce the loss function. The lower the loss function, the more similar the two variables are, that is, to make the output of the student model more and more similar to the output of the teacher model, so that the student model can obtain the performance matching the teacher model). For the image output, it passes through ResNet 152 and MXBERT respectively, and the outputs of the middle of the two models are fitted. The fitting method is L T . Because sequential training will cause the problem of forgetting, the present invention performs joint distillation, that is, L dis =L l +L T . The present invention only needs to optimize the L dis loss function.
[0071] In actual use, the student model MXBERT can input either text or image. One only needs to use the final output of the model for downstream tasks, such as text classification, sentiment analysis, image classification, etc.
[0072] (3.4) Experimental analysis
[0073] Table 1 shows the performance comparison results between MXBERT and text modality encoders. In Table 1, the models ELMo, GPT, and BERT in the second row are pre-trained models, and the models such as MobileBERT in the third row are all benchmark models for fair comparison. It can be seen that in the text modality, the method of the present invention not only surpasses the original benchmark model MobileBERT on multiple datasets of GLUE, but further surpasses the pre-trained model BERT on most tasks base . Table 2 shows the performance comparison results between XBERT and image encoders. From Table 2, it can be seen that in the image modality, the performance of the model surpasses the benchmark model ResNet 50 , and at the same time far surpasses ResNet 18 . It can be seen from this that modality unification will not affect each other, but instead has a certain degree of complementarity. Compared with other single-modality benchmark methods, CMDIR is simple and effective, does not require additional samples to participate in the calculation, not only unifies the distillation methods of different modalities, but also matches or even surpasses the original distillation method in terms of distillation performance.
[0074] Table 1. Performance comparison between MXBERT and text modal encoders. For CoLA, the evaluation metric is Matthews correlation coefficient; for SST-2, MNLI, QNLI, and RTE, the evaluation metric is accuracy; for MRPC and QQP, the evaluation metric is the average of F1 value and accuracy; for STS-B, the evaluation metric is Pearson correlation coefficient.
[0075]
[0076] Table 2. Performance comparison between MXBERT and image encoders. For the CIFAR dataset, the evaluation metric is top-1 error rate; for ImageNet, the evaluation metric is top-5 error rate.
[0077]
[0078] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, variational mutual information angle distillation is adopted in the present invention, which not only greatly reduces the information loss of the teacher model, but also does not require a large number of negative samples to participate in the calculation, and has simplicity and effectiveness.
[0079] In a multi-modal joint representation learning method based on variational distillation disclosed in the present invention, a joint distillation method is adopted in the present invention to solve the forgetting problem caused by multi-modal distillation.
[0080] Corresponding to the above method embodiments, the embodiments of the present invention further provide a computer device, including:
[0081] A memory for storing computer programs;
[0082] A processor for implementing the steps of the above multi-modal joint representation learning method based on variational distillation when executing the computer program.
[0083] In the embodiments of the present invention, the processor may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.
[0084] The processor can call the program stored in the memory. Specifically, the processor can execute the operations in the embodiments of the multi-modal joint representation learning method based on variational distillation.
[0085] The memory is used to store one or more programs, and the program may include program codes, and the program codes include computer operation instructions.
[0086] In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device or other volatile solid-state storage devices.
[0087] Corresponding to the above method embodiments, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above multi-modal joint representation learning method based on variational distillation are implemented.
[0088] Embodiment Two
[0089] Next, a multi-modal joint representation learning system based on variational distillation disclosed in Embodiment Two of the present invention will be introduced. The multi-modal joint representation learning system described below and the multi-modal joint representation learning method described above can be correspondingly referred to each other.
[0090] Embodiment Two of the present invention discloses a multi-modal joint representation learning system based on variational distillation, including:
[0091] A student model, the student model includes a multi-modal data unification module and a modal joint representation module. Input the original multi-modal data, where the original multi-modal data includes original text modal data and original image modal data. Input the original text modal data and the original image modal data into the multi-modal data unification module to obtain text modal inputs and image modal inputs with the same input form, and perform a normalization operation on the text modal inputs and the image modal inputs. Input the normalized text modal inputs and image modal inputs into the modal joint representation module respectively to obtain the text output and the image output of the student model;
[0092] A teacher model, the teacher model includes a text teacher model and an image teacher model. Input the original text modal data and the original image modal data into the text teacher model and the image teacher model respectively to obtain the text output and the image output of the teacher model;
[0093] A modal unification distillation module, the modal unification distillation module is used to characterize the correlation between the corresponding text output and image output of the student model and the teacher model by using variational mutual information, and perform joint distillation training on the text output and the image output by using a distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model.
[0094] In a multi-modal joint representation learning system based on variational distillation disclosed in the present invention, the multi-modal data unification module is deployed at the front end of the modal joint representation module, and the original text modal data and the original image modal data are sorted into the same input form by using the multi-modal data unification module to obtain text modal inputs and image modal inputs.
[0095] In a multi-modal joint representation learning system based on variational distillation disclosed in the present invention, the multi-modal data unification module includes:
[0096] A text modality data arrangement sub-module, which is used to add [CLS] symbols and [SEP] symbols to the text modality data, and at the same time add [DIS] symbols at the end of the sentences in the text modality data, and obtain the text modality input through a word vector matrix;
[0097] An image modality data arrangement sub-module, which is used to divide the image modality data into several image patches, stretch each image patch into a one-dimensional vector, add [CLS] symbols and [DIS] symbols at the beginning and end positions of the one-dimensional vector, and obtain an image modality input with the same form as the text modality input through dimension scaling.
[0098] The multi-modal joint representation learning system based on variational distillation in this embodiment is used to implement the foregoing multi-modal joint representation learning method based on variational distillation. Therefore, the specific implementation manner of this system can be seen in the embodiment part of the multi-modal joint representation learning method based on variational distillation in the foregoing text. Therefore, its specific implementation manner can refer to the descriptions of the corresponding individual part embodiments and will not be elaborated here.
[0099] In addition, since the multi-modal joint representation learning system based on variational distillation in this embodiment is used to implement the foregoing multi-modal joint representation learning method based on variational distillation, its function corresponds to the function of the above method and will not be repeated here.
[0100] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0101] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks
[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the processes Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks
[0104] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A multi-modal joint representation learning method based on variational distillation, characterized in that Including the following steps: Deploy a student model and a teacher model, where the teacher model includes a text teacher model and an image teacher model, and the student model includes a multimodal data unification module. Input the original multimodal data, where the original multimodal data includes original text modal data and original image modal data. Input the original text modal data and the original image modal data into the multimodal data unification module to obtain text modal inputs and image modal inputs with the same input form, and perform a normalization operation on the text modal inputs and the image modal inputs; The student model includes a modal joint representation module. Input the normalized text modal inputs and image modal inputs into the modal joint representation module respectively to obtain the text output and the image output of the student model. At the same time, input the original text modal data and the original image modal data into the text teacher model and the image teacher model respectively to obtain the text output and the image output of the teacher model; Use variational mutual information to characterize the correlation between the text output and the image output corresponding to the student model and the teacher model, and jointly distill and train the text output and the image output using a distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model; Among them, the multimodal data unification module is deployed at the front end of the modal joint representation module. Use the multimodal data unification module to organize the original text modal data and the original image modal data into the same input form to obtain text modal inputs and image modal inputs; Organize the original text modal data and the original image modal data into the same input form to obtain text modal inputs and image modal inputs, including: Add [CLS] symbols and [SEP] symbols to the original text modal data, and at the same time add [DIS] symbols at the end of the sentences in the original text modal data, and obtain the text modal inputs through a word vector matrix; Segment the original image modal data into several image patches, stretch each image patch into a one-dimensional vector, add [CLS] symbols and [DIS] symbols at the beginning and end positions of the one-dimensional vector, and obtain image modal inputs with the same form as the text modal inputs through dimensional scaling.
2. The multimodal joint representation learning method based on variational distillation according to claim 1, wherein: The modal joint representation module includes a MobileBERT model, and the MobileBERT model includes a 24-layer transformer model, and a linear layer is added to each layer of the transformer model.
3. The multimodal joint representation learning method based on variational distillation according to claim 1, wherein: The distillation loss function is the sum of the loss function of the text teacher model and the loss function of the image teacher model.
4. A multi-modal joint representation learning system based on variational distillation, characterized in that, Including: Student model, the student model includes a multi-modal data unification module and a modal joint representation module, which inputs the original multi-modal data. The original multi-modal data includes original text modal data and original image modal data. The original text modal data and the original image modal data are input into the multi-modal data unification module to obtain text modal input and image modal input with the same input form, and normalization operations are performed on the text modal input and the image modal input. The normalized text modal input and image modal input are respectively input into the modal joint representation module to obtain the text output and image output of the student model. And the multi-modal data unification module is deployed at the front end of the modal joint representation module. The multi-modal data unification module is used to organize the original text modal data and the original image modal data into the same input form to obtain text modal input and image modal input; Organizing the original text modal data and the original image modal data into the same input form to obtain text modal input and image modal input includes: adding [CLS] symbol and [SEP] symbol to the original text modal data, and at the same time adding [DIS] symbol at the end of the sentence in the original text modal data, and obtaining the text modal input through the word vector matrix; dividing the original image modal data into several image patches, stretching each image patch into a one-dimensional vector, adding [CLS] symbol and [DIS] symbol at the beginning position and the end position of the one-dimensional vector, and obtaining the image modal input with the same form as the text modal input through dimensional scaling; Teacher model, the teacher model includes a text teacher model and an image teacher model. The original text modal data and the original image modal data are respectively input into the text teacher model and the image teacher model to obtain the text output and image output of the teacher model; Modal unification distillation module, the modal unification distillation module is used to characterize the correlation between the text output and image output corresponding to the student model and the teacher model by using variational mutual information, and jointly distill and train the text output and image output by using the distillation loss function, so that the student model can simultaneously obtain the ability to match the text teacher model and the image teacher model.
5. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Multilayer neural network language model training method and device based on knowledge distillation
CN111611377A
Image-text matching model compression and acceleration method based on orthogonal similarity distillation and system thereof
CN112990296A