Method and device for training image-text joint coding model
By training the joint coding model of graphic and text, using the large language model to rewrite text and combining the aggregate coding and mask cross-attention calculation of the joint coding model of the joint coding model, the problem that the cross-modal coding model cannot understand the intersection information and difference information of the image and text is solved, and the effective joint coding of the pair of graphic and text is realized, and the generalization and performance of the model is improved.
Patent Information
- Application Number
- CN202510436235.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-01
AI Technical Summary
The existing cross-modal encoding model cannot effectively understand the intersection information and difference information between images and text, and cannot jointly encode the graphics and text pairs, resulting in the inability to characterize and match the mixed graphics and text data in practical applications.
By constructing and training the joint coding model of graphic and text, using the large language model to rewritten text to generate summary text, and using the joint coding model to perform aggregate coding and mask cross attention calculation, combining reconstruction loss and comparison loss for model training, generating reconstructed text to update the model.
The joint coding of graphic and text pairs is realized, the generalization and quality of the model is improved, and a lightweight joint coding model with good speed and performance is constructed.
Smart Images

Figure CN120411986A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine learning, and in particular, to a method and apparatus for training an image-text joint encoding model. Background Art
[0002] With the development of the Internet, intelligent devices, and the multimedia field, people can more conveniently access various modal data, including text, images, videos, voices, and so on. Since different modal data are good at expressing different information, a lot of content people encounter in life is a combination of multiple modal data. For example, social media posts often contain both images and text, which can leverage the advantages of each modality to make the expression of information more accurate and easier.
[0003] Generally, there is a certain degree of semantic mapping relationship between information of different modalities. Therefore, cross-modal encoding models help to better understand the semantic mapping relationship between different modalities. However, although the cross-modal encoding models in related technologies can encode representation vectors with similar distances for different modal data with similar semantics, they still cannot jointly encode data of two modalities, such as encoding an image-text pair to generate an image-text joint representation. Therefore, a method is needed to train a model that can jointly encode image-text pairs to better represent the joint content of image-text pairs. Summary of the Invention
[0004] One or more embodiments of this specification describe a method and apparatus for training an image-text joint encoding model, which trains a model that can jointly encode image-text pairs to better represent the joint content of image-text pairs.
[0005] In a first aspect, a method for training an image-text joint encoding model is provided, including:
[0006] Input a first image and a first text included in a first image-text pair into a large language model, and instruct the large language model to rewrite the first text by introducing semantic content in the first image to obtain a first summary text;
[0007] Using the image-text joint encoding model, aggregate and encode a first image representation and a first text representation corresponding to the first image and the first text respectively to obtain a first joint representation;
[0008] Perform masked cross-attention calculation on the first joint representation and a first summary representation corresponding to the first summary text to obtain a first masked representation;
[0009] Decode the first masked representation to obtain a first reconstructed text;
[0010] Update the image-text joint encoding model according to the training loss, where the training loss at least includes a reconstruction loss determined according to the difference between the first reconstructed text and the first summary text.
[0011] In some possible implementation manners, the image-text joint encoding model includes at least one self-attention layer and a Q-Former module; obtaining the first joint representation includes:
[0012] Input the first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation;
[0013] Input the first attention representation and a trainable query representation into the Q-Former module for encoding to obtain a first joint representation.
[0014] In some possible implementation manners, the image-text joint encoding model includes at least one self-attention layer; obtaining the first joint representation includes:
[0015] Input the first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation;
[0016] Determine the representation corresponding to the first token in the first attention representation as the first joint representation.
[0017] In some possible implementation manners, the first summary representation includes representation vectors of multiple tokens respectively; obtaining the first masked representation includes:
[0018] Expand the first joint representation in the spatial dimension to determine a first expanded joint representation with the same size as the first summary representation;
[0019] Determine a query matrix according to the first expanded joint representation, determine a key matrix and a value matrix according to the first summary representation, and perform masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation.
[0020] In some possible implementation manners, performing masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation includes:
[0021] Obtain a first score matrix according to the matrix multiplication and normalization result of the query matrix and the key matrix;
[0022] Apply the first mask matrix to the first score matrix to obtain a first attention matrix;
[0023] Multiply the first attention matrix by the value matrix to obtain a first mask representation.
[0024] In some possible embodiments, the reconstruction loss is determined according to the cross-entropy loss between the first reconstructed text and the first summary text.
[0025] In some possible embodiments, the training loss further includes a contrast loss; the method further includes:
[0026] Obtain at least one positive sample and a number of negative samples of the first image-text pair;
[0027] Determine the contrast loss according to the first similarity between the first joint representation and the joint representations corresponding to the positive samples, and the second similarities between the first joint representation and the joint representations corresponding to the samples in the sample set respectively; wherein, the sample set includes the number of negative samples.
[0028] In some possible embodiments, obtaining at least one positive sample of the first image-text pair includes:
[0029] Classify the image-text pair formed by the first image and the first summary text into the positive samples.
[0030] In some possible embodiments, obtaining a number of negative samples of the first image-text pair includes:
[0031] Input the first image-text pair into a large language model, and instruct the large language model to generate a first forged text with different semantics according to the first text, and the first forged text has no semantic contradiction with the first image;
[0032] Classify the image-text pair formed by the first image and the first forged text into the negative samples.
[0033] In some possible embodiments, the first image representation includes the representation vectors of the respective image patches of the first image; obtaining at least one positive sample and a number of negative samples of the first image-text pair includes:
[0034] Perform a clustering operation on the image patches of the first image according to the first image representation to obtain a plurality of image regions;
[0035] Determine the respective image-text similarities between the first text and the respective image regions of the first image according to the first image representation and the first text representation;
[0036] Determine the positive regions for the image regions with image-text similarity greater than a preset first threshold, and determine the negative regions for the image regions with image-text similarity less than or equal to the first threshold;
[0037] The text-image pair consisting of the first image representation with the positive region masked and the first text is classified as a positive sample, and the text-image pair consisting of the first image representation with the negative region masked and the first text is classified as a negative sample.
[0038] In some possible implementation manners, clustering operations are performed on the image patches of the first image according to the first image representation to obtain a plurality of image regions, including:
[0039] Pooling operations are performed on the first feature map composed of the representation vectors of each image patch, and clustering operations are performed based on the second feature map after pooling, and each clustering cluster is determined as each image region.
[0040] In some possible implementation manners, determining the text-image similarity between the first text and each image region of the first image includes:
[0041] For any target region in each image region, according to the first image representation, the similarity between the first text representation and each image patch included in the target region is determined, and then the average similarity is determined as the text-image similarity between the first text and the target region.
[0042] In some possible implementation manners, determining the contrast loss includes:
[0043] The contrast loss is determined according to the ratio between the summation result of each first similarity and the summation result of each second similarity.
[0044] In some possible implementation manners, the training loss is determined according to the weighted summation result between the reconstruction loss and the contrast loss.
[0045] In a second aspect, a device for training a text-image joint encoding model is provided, including:
[0046] A first rewriting unit, configured to input the first image and the first text included in the first text-image pair into a large language model, and instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain a first summary text;
[0047] A first aggregation encoding unit, configured to use the text-image joint encoding model to perform aggregation encoding on the first image representation and the first text representation corresponding to the first image and the first text respectively to obtain a first joint representation;
[0048] A cross-attention unit, configured to perform masked cross-attention calculation on the first joint representation and the first summary representation corresponding to the first summary text to obtain a first masked representation;
[0049] A decoding unit, configured to decode the first masked representation to obtain a first reconstructed text;
[0050] A training unit, configured to update the text-image joint encoding model according to a training loss, where the training loss at least includes a reconstruction loss determined according to a difference between the first reconstructed text and the first summary text.
[0051] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0052] In a fourth aspect, a computing device is provided, including a memory and a processor. Wherein, an executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.
[0053] The method and device for training a text-image joint encoding model proposed in the embodiments of this specification. The method first utilizes the image understanding ability of a pre-trained large language model, inputs a text-image pair into the large language model, and makes it rewrite the text by introducing semantic content in the image to obtain a summary text containing the content of the text-image pair. Then, a text-image joint encoding model is used to perform aggregated encoding on the text-image pair to obtain a joint representation. Next, cross-attention calculation and decoding are performed on the joint representation and the summary text to obtain a reconstructed text for the summary text. Furthermore, the text-image joint encoding model is updated according to a reconstruction loss determined according to a difference between the reconstructed text and the summary text.
[0054] The method described in the embodiments of this specification uses information transfer to achieve knowledge distillation of the text-image understanding ability of the large language model, and thus constructs training data with stronger generalization ability and better quality. Furthermore, according to the training data, a text reconstruction task is used to train a relatively small-scale text-image joint encoding model to obtain a lightweight text-image joint encoding model with good speed and performance. Description of the Drawings
[0055] To more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the multiple embodiments disclosed in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0056] Figure 1 A schematic diagram showing the image data in a text-image pair according to an example;
[0057] Figure 2 A schematic flowchart showing the process of generating a summary text and a forged text according to an embodiment;
[0058] Figure 3 Schematic diagram of an implementation scenario showing a method for training a text-image joint encoding model according to an embodiment;
[0059] Figure 4 Schematic flow diagram showing a process for generating positive sample images and negative sample images according to an embodiment;
[0060] Figure 5 Flow chart showing a method for training a text-image joint encoding model according to an embodiment;
[0061] Figure 6 Flow chart showing a method for generating positive and negative samples according to an embodiment;
[0062] Figure 7 Schematic block diagram showing an apparatus for training a text-image joint encoding model according to an embodiment. Detailed implementation manners
[0063] The following describes the solution provided in this specification in conjunction with the accompanying drawings.
[0064] As described above, the information that different modalities of data are good at expressing is different, and generally speaking, there is a certain semantic mapping relationship between different modalities of information. The most common of these is the text-image pair formed by text and image. For example, there is a semantic mapping relationship between the text "cat" and an image containing a cat.
[0065] Cross-modal encoding models in the related art usually have a dual-encoder structure, including an image encoder and a text encoder, which respectively encode images and texts into representation vectors (embeddings). When training the model, the similarity between the representation vectors corresponding to texts and images with similar semantics is increased, while the similarity between the representation vectors corresponding to texts and images with different semantics is decreased. For example, the similarity between the representation vectors of the text "cat" and a cat image is increased, while the similarity between the representation vectors of the text "cat" and a dog image is decreased. Thus, for different modalities of data containing similar semantics, similar representation vectors are encoded.
[0066] Although current cross-modal encoding models have achieved a certain degree of understanding of multi-modal information to some extent, they can only receive text or image as input separately. In real life and related applications, in order to make the expressed content clearer and more vivid, text and image often appear simultaneously, such as the illustrated text in social media, product introductions on online shopping websites, and so on. In these scenarios, the information contained in the text and the image is usually related but not exactly the same, that is, there is both intersection information and difference information between the semantic information contained in the image and the text. For example, for the text "2019 high-top canvas shoes for women" andFigure 1 For the text-image pair composed of the shown image, the intersection information of the text and the image includes the semantic information of the high-top canvas shoes, the difference information of the text relative to the image includes the semantic information of "2019" and "female", and the difference information of the image relative to the text includes the semantic information of the ankles, pants, and background of the person.
[0067] After research, the inventors found that the cross-modal encoding models in the related art cannot understand the intersection information and difference information between images and texts, so they cannot directly encode the text-image pair to obtain a joint encoding representation that represents the joint semantics of the text-image pair, and thus cannot realize the representation and matching of text-image mixed data in practical applications.
[0068] To solve the above problems, the embodiments of this specification propose a method for training a text-image joint encoding model, which can encode text-image mixed data (text-image pairs) to obtain a model of joint representation vectors.
[0069] First, describe the process of constructing the sample set used for training the text-image joint encoding model. First, obtain a set of text-image pairs, and any text-image pair contains a text T and an image I. The text T and the image I may contain similar semantic information, and at the same time each contains semantic information different from the other, that is, there is both intersection information and difference information between the semantic information contained in the image T and the text I.
[0070] In the training task of the text-image joint encoding model, for any text-image pair 1 (T1, I1), its positive sample should have the same semantics as the joint text-image representation of the text-image pair 1, and at the same time the text and image contents are not exactly the same. That is, for the positive sample text-image pair (T2, I2) of the text-image pair 1, s(T1, I1) = s(T2, I2), T1 ≠ T2 or I1 ≠ I2. Where s() represents the semantic information of the data in the parentheses.
[0071] The negative sample text-image pair 3 (T3, I3) of the text-image pair 1 should have different semantics from the joint text-image representation of the text-image pair 1. This can be further classified into simple negative samples and difficult negative samples. The simple negative sample has different semantics from the joint text-image representation of the text-image pair 1, and at the same time the text and image contents are also completely different, and the text-image joint encoding model can easily distinguish the text-image pair 1 from its simple negative sample. Specifically, the simple negative sample can be any text-image pair different from the text-image pair 1 in the previously obtained set of text-image pairs.
[0072] The difficult negative sample has a different semantics from the text-image joint representation of text-image pair 1, but the text and image content are similar to those of text-image pair 1, that is, the texts of the two are similar and / or the images are similar. It is relatively more difficult for the text-image joint encoding model to distinguish text-image pair 1 from its difficult negative sample. Using difficult negative samples during model training can better train the encoding ability of the text-image joint encoding model.
[0073] The specific processes of constructing positive samples and difficult negative samples are described below.
[0074] For positive samples, the difference set information of a certain modality data in the original text-image pair can be transferred to the other modality data to generate a new text-image pair as the positive sample of the original text-image pair. For difficult negative samples, a certain modality data in the original text-image pair can be modified to generate a new text-image pair as the difficult negative sample of the original text-image pair. This can be achieved by leveraging the image and text understanding capabilities of a pre-trained large language model (LLM). Since the large language model is more proficient in text generation, in the embodiments of this specification when using the large language model to generate positive and negative samples, it is selected to make the large language model generate new text based on the original text-image pair and form new corresponding positive and negative samples with the image in the original text-image pair.
[0075] Figure 2 The flowchart showing the generation of summary text and forged text according to an embodiment is as follows. Figure 2 As shown, the text T and the image I in the original text-image pair (T, I) are input into the large language model, and a summary prompt is input, instructing the large language model to rewrite the text T by introducing the semantic content in the image I to obtain the summary text T s . The semantic information of the summary text T s generated in this way is the same as the joint semantic information of the original text-image pair (T, I), that is, s(T, I) = s(T s ). The text-image pair (T s , I) composed of the summary text T s and the image I is used as the positive sample of the original text-image pair (T, I). Exemplarily, the summary prompt can be, for example: "Please rewrite the text according to the given image, requiring that the rewritten text does not omit or change the information in the original text, and at the same time includes more information in the image."
[0076] On the other hand, the text T and the image I in the original text-image pair (T, I) are input into the large language model, and a forgery prompt is input, instructing the large language model to generate forged text T f with different semantics according to the text T, and ensuring that there is no semantic contradiction between the forged text T f and the image I. The forged text Tf The text-image pair (T f , I) composed of the image I is used as a difficult negative sample of the original text-image pair (T, I). Exemplarily, the forged prompt can be, for example: "Please modify the text in combination with the image content so that the joint semantics between the modified text and the image change, but there is no contradiction with the image."
[0077] After constructing the positive and negative samples of the original text-image pair, these samples can be used to train the text-image joint encoding model. Figure 3 FIG. shows an implementation scenario diagram of a method for training a text-image joint encoding model according to an embodiment. As Figure 3 shown, the data used to train the text-image joint encoding model includes the original text-image pair (T, I) and the summary text T s generated based on the original text-image pair (T, I). The overall architecture of model training includes Figure 3 the text-image joint encoding model to be trained on the left, and the text-conditioned reconstruction module on the right. Figure 3 The specific structures of the respective sub-modules will be described in detail in the subsequent content.
[0078] First, the pre-trained image encoder and the pre-trained text encoder are used to encode the image I and the text T respectively, to obtain the image representation f I and the text representation f T . Then, the image representation f I and the text representation f T are concatenated and input into the text-image joint encoding model for encoding, to obtain the joint representation f cls . To enable the joint representation f cls to contain the joint semantics of the text and the image, the training objective is to hope that the summary text T cls can be reconstructed based on the joint representation f s , so as to achieve supervised training of the model.
[0079] To achieve this supervision, on the other hand, the summary text T s is input into the encoding layer to obtain the summary representation f s . Then, cross-attention calculation and decoding are performed between the joint representation f cls and the summary representation f s to obtain the reconstructed text T r . Specifically, after the joint representation f cls is expanded in the spatial dimension (i.e., the joint representation f cls is repeatedly arranged in the spatial dimension several times to obtain a matrix f′ s with the same size as the summary representation f cls ), it is used as the query matrix Query in the cross-attention mechanism, and the summary representation fs As the key matrix Key and value matrix Value in the cross-attention mechanism, cross-attention calculation is performed. In this process, in order to play the role of the joint representation f cls in reconstruction, the embodiments of this specification use a mask matrix M when performing cross-attention calculation to mask the features corresponding to some tokens, and calculate the masked representation through other unmasked tokens and the joint representation f cls as shown in Formula (1):
[0080]
[0081] where Attention(Q, K, V, M) represents the masked representation, Q = f′ cls W Q , K = f s W K , V = f s W V , W Q , W K , W V are the query weight matrix, key weight matrix, and value weight matrix respectively. M is the mask matrix, The operator represents the operation of masking the values in the corresponding matrix with the mask matrix M. d k is the representation dimension.
[0082] Then the masked representation is decoded to reconstruct the summary text T s to obtain the reconstructed text T r . According to the difference between the reconstructed text T r and the summary text T s , the reconstruction loss is determined, and then the trainable parameters in the image-text joint encoding model are adjusted. When using the cross-entropy loss to measure the difference between the reconstructed text T r and the summary text T s , the reconstruction loss L r can be as shown in Formula (2):
[0083]
[0084] where log represents the natural logarithm, p θ (y i |{y j , M j ≠0}, f cls ) represents the probability that the i-th token of the text output by the image-text joint encoding model is y i , y i is the i-th token in the summary text T s , N is the number of tokens in the summary text T s , {yj ,M ij , the set of tokens that are not masked when predicting y i .
[0085] The above describes the specific process of training the image-text joint encoding model based on the text-conditioned reconstruction task. Further, the image-text joint encoding model can also be directly trained with the contrastive loss determined by the respective similarities between the original image-text pair and its corresponding positive and negative samples, together with the reconstruction loss.
[0086] For any original image-text pair P among multiple original image-text pairs i , P i 's corresponding set of positive samples can be denoted as P i 's corresponding set of negative samples can be denoted as It can include constructed hard negative samples or simple negative samples, that is, all original image-text pairs except P i , that is, {P j , j≠i}.
[0087] The contrastive loss L in one embodiment c can be as shown in formula (3):
[0088]
[0089] where M is the number of original image-text pairs, exp() represents the natural exponent, and f cls () represents the image-text joint representation obtained by encoding the image-text pair inside the parentheses through the image-text joint encoding model. <x, y> represents the similarity between representations x and y, such as inner product similarity, cosine similarity, etc. τ is a hyperparameter used to control the magnitude of the values of each similarity. L c is used to increase the similarity between the original image-text pair P i and its corresponding positive sample, while reducing the similarity between the original image-text pair Pi and its corresponding negative sample.
[0090] At this time, the total training loss L of the model can be the weighted sum form of the reconstruction loss L r and the contrastive loss L c , as shown in formula (4):
[0091] L = λ1L r + λ2L c (4)
[0092] where λ1 and λ2 are preset weight coefficients.
[0093] For the set of positive samples of the original image-text pair P i and the negative sample set The foregoing Figure 2 only describes the process of generating positive and negative samples by rewriting the text in the original text-image pair. In some possible implementation manners, more positive and negative samples can also be generated by rewriting the image in the original text-image pair. Figure 4 shows a schematic flow chart of generating a positive sample image and a negative sample image according to an embodiment. As Figure 4 shown, first, the text and the image in the original text-image pair are respectively encoded using a text encoder and an image encoder to obtain a text representation and an image representation. The image representation includes the representation vectors of each image patch in the original image. Then, the similarity between the text representation and the representation vectors of each image patch is calculated respectively to obtain a similarity matrix, which indicates the similarity degree between the text and each part (image patch) of the image.
[0094] Next, the representation vectors of each image patch are arranged according to the position of each image patch in the original image, and reshaped to obtain a feature map of the original image. The vectors at each position of the feature map represent the representation vectors of the image patches at that position in the original image. Then, pooling and clustering operations are performed on the feature map to obtain a plurality of clustering clusters, and the image is divided into a plurality of image regions according to the clustering clusters.
[0095] For any image region, the average value of the similarity values of the image patches included in it in the similarity matrix is taken as the text-image similarity between the image region and the text. The image regions with text-image similarity greater than a preset threshold are determined as positive regions, such as Figure 1 the region of the shoe part in ; and the image regions with text-image similarity less than or equal to the preset threshold are determined as negative regions. Such as Figure 1 the regions of the pants part and the background part in.
[0096] The part of the positive region of the image representation is covered with a mask as a positive sample image, which together with the text constitutes a positive sample of the original text-image pair and is added to the positive sample set; the part of the negative region of the image representation is covered with a mask as a negative sample image, which together with the text constitutes a negative sample of the original text-image pair and is added to the negative sample set.
[0097] Since the positive sample image covers the part similar to the text, and the semantic information of this part still exists in the text, the semantic information of the positive sample formed by the image representation after covering the positive region and the text is the same as that of the original text-image pair. The negative sample image covers the part different from the text, which is equivalent to lacking the unique information in the image. Therefore, the semantic information of the negative sample formed by the image representation after covering the negative region and the text is different from that of the original text-image pair.
[0098] It should be noted that Figure 3 The multiple image regions, positive regions, and negative regions in it are only examples and do not limit the protection scope of the embodiments of this specification.
[0099] The positive and negative samples of the original text-image pairs constructed according to the method as Figure 4 shown, and the positive and negative samples of the original text-image pairs constructed according to the method as Figure 2 shown are compared together to calculate the contrastive loss L c This can further enhance the representation ability of the text-image joint encoding model for text-image pairs.
[0100] The following describes the specific implementation steps of the above method for training a text-image joint encoding model in combination with specific embodiments.
[0101] Figure 5 FIG. shows a flowchart of a method for training a text-image joint encoding model according to an embodiment. The execution subject of the method can be any platform, server, or device cluster with computing and processing capabilities, etc. As Figure 5 shown, the method at least includes: Step 502, input the first image and the first text included in the first text-image pair into a large language model, and instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain a first summary text; Step 504, use the text-image joint encoding model to aggregate and encode the first image representation and the first text representation corresponding to the first image and the first text respectively to obtain a first joint representation; Step 506, perform masked cross-attention calculation on the first joint representation and the first summary representation corresponding to the first summary text to obtain a first masked representation; Step 508, decode the first masked representation to obtain a first reconstructed text; Step 510, update the text-image joint encoding model according to the training loss, where the training loss at least includes a reconstruction loss determined according to the difference between the first reconstructed text and the first summary text.
[0102] The following describes the specific execution processes of the above steps.
[0103] First, in Step 502, input the first image and the first text included in the first text-image pair into a large language model, and instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain a first summary text.
[0104] The first text-image pair can be any one of multiple text-image pairs. The first image and the first text can contain similar semantic information, and at the same time, each contains semantic information different from the other.
[0105] The large language model can use any pre-trained large language model, which is not limited here.
[0106] It can be achieved by using corresponding summarization prompting words to instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain the first summary text. In one embodiment, the summarization prompting words can be: "Please rewrite the text according to the given image, requiring that the rewritten text does not omit or change the information in the original text, and at the same time includes more information in the image." It can be understood that in other embodiments, other summarization prompting words can also be constructed as long as they can instruct the large language model to generate the first summary text.
[0107] In one embodiment, since there may be noise in the summary text generated by the large language model, it can be post-processed and filtered to remove the first summary text that does not meet the preset requirements.
[0108] In a specific embodiment, after encoding the first text and the first summary text respectively using a pre-trained text encoder, the third similarity between the representations is calculated. When the third similarity exceeds the preset second threshold, it indicates that the degree of rewriting of the first text by the large language model is too low, and then the first summary text is removed.
[0109] In another specific embodiment, the first text and the first summary text are encoded respectively using a text encoder, and the first image is encoded using an image encoder. The image encoder and the text encoder have been pre-aligned and trained, and similar representations can be encoded for semantically similar images and texts. For example, it can be the image encoder and the text encoder of the CLIP (Contrastive Language-Image Pre-training) model, or the image encoder and the text encoder of the M2Encoder model, which is not limited here.
[0110] Then, the fourth similarity between the first text and the first image is compared with the fifth similarity between the first summary text and the first image. When the fourth similarity is greater than the fifth similarity, it indicates that the first summary text does not contain more image information compared to the first text, so the first summary text is removed.
[0111] Since the large language model itself has a certain ability to comprehensively understand image and text information, but since it is a generative model and cannot directly output the joint representation of the image-text pair, so in step 502, a summary text is constructed by using the large language model to achieve knowledge distillation of the image and text understanding ability contained in the large language model.
[0112] Then, in step 504, using the image-text joint encoding model, the first image representation and the first text representation corresponding to the first image and the first text are aggregated and encoded to obtain the first joint representation.
[0113] The first image representation can be obtained by encoding the first image into a pre-trained image encoder; the first text representation can be obtained by encoding the first text into a pre-trained text encoder. The image encoder and the text encoder can be pre-aligned and trained, and similar representations can be encoded for semantically similar images and texts. For example, it can be the image encoder and text encoder of the CLIP model, or the image encoder and text encoder of the M2Encoder model, which is not limited here.
[0114] The image-text joint encoding model can be implemented based on various model structures. In one embodiment, the image-text joint encoding model includes at least one self-attention layer and a Q-Former module. In this embodiment, obtaining the first joint representation in step 504 includes:
[0115] Inputting the first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation; inputting the first attention representation and a trainable query representation into the Q-Former module for encoding to obtain a first joint representation.
[0116] The self-attention layer can be the self-attention layer in the Transformer model, which performs self-attention calculation after receiving the input representation and outputs the self-attention representation. The Q-Former module can be a module in the BLIP2 (Bootstrapping Language-Image Pre-training 2) model, which receives the input representation and a set of trainable query representations (queries) and outputs the encoded representation.
[0117] In another embodiment, the image-text joint encoding model includes at least one self-attention layer. Obtaining the first joint representation in step 504 includes:
[0118] Inputting the first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation; determining the representation corresponding to the first token in the first attention representation as the first joint representation.
[0119] In this embodiment, the first token in the first attention representation is a pre-set special token [CLS], which is used to summarize the global semantics of the image-text pair.
[0120] In step 504, the first image representation and the first text representation can be in matrix form, where each row of the matrix represents the representation vector corresponding to the corresponding token (or image patch) in the image / text. The first joint representation is in vector form and represents the final representation obtained after encoding the first image-text pair through the image-text joint encoding model.
[0121] Next, in step 506, a masked cross-attention calculation is performed between the first joint representation and the first summary representation corresponding to the first summary text to obtain a first masked representation.
[0122] The first summary representation can be obtained by inputting the first summary text into a pre-trained text encoder or encoding layer for encoding. The encoding layer can be a single-layer neural network or a multi-layer perceptron.
[0123] In one embodiment, the first summary representation includes the representation vectors of multiple tokens, i.e., in matrix form. Obtaining the first masked representation in step 506 includes:
[0124] Expanding the first joint representation in the spatial dimension to determine a first expanded joint representation with the same size as the first summary representation; determining a query matrix based on the first expanded joint representation, determining a key matrix and a value matrix based on the first summary representation, and performing a masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation.
[0125] The first joint representation can be a 1×d-dimensional vector, and the first summary representation can be an N×d-dimensional matrix, where n is the number of tokens included in the first summary text. Expanding the first joint representation in the spatial dimension means repeating the first joint representation in the spatial dimension n times to obtain a first expanded joint representation that is also n×d-dimensional.
[0126] The first mask matrix can be an N×N-dimensional matrix, where multiple positions contain randomly generated masks, for example, represented by 0. The meaning of the mask at the i-th row and j-th column of the first mask matrix can be to mask the context information of the j-th token for the i-th token.
[0127] Determining the query matrix based on the first expanded joint representation can be multiplying the first expanded joint representation by a query weight matrix to determine the query matrix. Determining the key matrix and the value matrix based on the first summary representation can be multiplying the first summary representation by a key weight matrix and a value weight matrix respectively to determine the key matrix and the value matrix.
[0128] In a more specific embodiment, performing a masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation includes:
[0129] Obtain a first score matrix based on the result of matrix multiplication and normalization of the query matrix Q and the key matrix K Apply the first mask matrix M to the first score matrix to obtain a first attention matrix Multiply the first attention matrix by the value matrix to obtain a first masked representation Attention(Q, K, V, M), as shown in formula (1).
[0130] In other embodiments, a first masked representation can also be obtained by using the method of masked cross-attention calculation. For example, using the first summary representation as the query matrix and the first joint representation as the key matrix and value matrix, etc., which are not limited here.
[0131] Then, in step 508, decode the first masked representation to obtain a first reconstructed text.
[0132] Specifically, input the first masked representation into a pre-trained decoding layer for decoding to obtain a first reconstructed text. The decoding layer can be a single-layer neural network or a multi-layer perceptron.
[0133] Finally, in step 510, update the text-image joint encoding model according to the training loss, where the training loss at least includes a reconstruction loss determined based on the difference between the first reconstructed text and the first summary text.
[0134] In one embodiment, the reconstruction loss in step 510 is determined according to the cross-entropy loss between the first reconstructed text and the first summary text, as shown in formula (2).
[0135] Updating the text-image joint encoding model according to the training loss can be based on the gradient descent method to adjust each trainable parameter in the text-image joint encoding model, which will not be elaborated here.
[0136] It can be understood that steps 502 to 510 only show the process of updating the text-image joint encoding model according to a first text-image pair among multiple text-image pairs. In other embodiments, multiple text-image pairs can also be used to update the text-image joint encoding model one or more rounds respectively based on the processes described in steps 502 to 510, which will not be elaborated here.
[0137] The above describes the process of knowledge distillation for large language models through information transfer to construct training data with higher generalization and quality, and then training the text-image joint encoding model based on the method of text-conditioned reconstruction.
[0138] In some possible implementation manners, the training loss further includes a contrastive loss. The method for training the text-image joint encoding model further includes step (a) and step (b).
[0139] In step (a), at least one positive sample and a number of negative samples of the first image-text pair are obtained.
[0140] In one embodiment, obtaining at least one positive sample of the first image-text pair in step (a) includes:
[0141] Classifying the image-text pair formed by the first image and the first summary text into the positive samples.
[0142] The image-text pair formed by the first image and the first summary text corresponds to the positive sample constructed by rewriting the text as shown in Figure 2 shown.
[0143] In another embodiment, obtaining a number of negative samples of the first image-text pair includes:
[0144] Inputting the first image-text pair into a large language model, instructing the large language model to generate a first forged text with different semantics according to the first text, where the first forged text has no semantic contradiction with the first image; classifying the image-text pair formed by the first image and the first forged text into the negative samples.
[0145] In this embodiment, the process of instructing the large language to generate the first forged text can be achieved by using forged prompt words. In a specific embodiment, the forged prompt words can be: "Please modify the text in combination with the image content so that the combined semantics of the modified text and the image change, but there is no contradiction with the image." In other embodiments, other forms of forged prompt words can also be constructed as long as they can instruct the large language model to generate a first forged text with different semantics and no semantic contradiction with the first image.
[0146] The image-text pair formed by the first image and the first forged text corresponds to the difficult negative sample constructed by rewriting the text as shown in Figure 2 shown.
[0147] In yet another embodiment, obtaining a number of negative samples of the first image-text pair includes:
[0148] Obtaining multiple image-text pairs and classifying the multiple image-text pairs into the negative samples.
[0149] The multiple image-text pairs are the simple negative samples corresponding to the first image-text pair.
[0150] In yet another embodiment, the method shown in Figure 4 can also be used to construct positive and negative samples by rewriting the image.
[0151] In this embodiment, the first image representation includes the representation vectors of the respective image patches of the first image. Step (a) of obtaining at least one positive sample and several negative samples of the first text-image pair includes steps 602 to 608 as shown in Figure 6 Figure 602 to Figure 608 shown below. Figure 6 Figure 4 shows a flowchart of a method for generating positive and negative samples according to an embodiment.
[0152] In step 602, clustering operations are performed on the image patches of the first image according to the first image representation to obtain a plurality of image regions.
[0153] Specifically, a pooling operation is performed on the first feature map composed of the representation vectors of each image patch, and a clustering operation is performed based on the second feature map after pooling, and each clustering cluster is determined as each image region.
[0154] The pooling operation can be average pooling, max pooling, etc., which is not limited here.
[0155] The clustering operation can be based on any density-based clustering algorithm, such as the K-means clustering algorithm, hierarchical clustering algorithm, etc., which is not limited here.
[0156] Then, in step 604, according to the first image representation and the first text representation, the respective text-image similarities between the first text and the respective image regions of the first image are determined.
[0157] Specifically, for any target region in each image region, according to the first image representation, the similarities between the first text representation and the respective image patches included in the target region are determined, and then the average similarity is determined as the text-image similarity between the first text and the target region.
[0158] For the k image patches included in the target region, the respective representation vectors of the k image patches are determined from the first image representation. Then, the similarities between the first text representation and the respective k image patches are calculated respectively, and the average value is taken as the text-image similarity between the first text and the target region.
[0159] Next, in step 606, the image regions with text-image similarity greater than a preset first threshold are determined as positive regions, and the image regions with text-image similarity less than or equal to the first threshold are determined as negative regions.
[0160] The positive region is the region related to the first text, that is, the part of the intersection information of the image and the text; the negative region is the region unrelated to the first text, that is, the part of the difference set information of the image and the text.
[0161] Finally, at step 608, the image-text pair formed by the first image representation of the positive region covered with the mask and the first text is classified as a positive sample, and the image-text pair formed by the first image representation of the negative region covered with the mask and the first text is classified as a negative sample.
[0162] The positive and negative samples constructed in steps 602 to 608 can be used for the contrast training of the image-text joint encoding model in subsequent step (b). The negative sample among them is the hard negative sample corresponding to the first image-text pair.
[0163] Then, at step (b), according to the first similarity between the first joint representation and the joint representations corresponding to each positive sample, and the second similarity between the first joint representation and the joint representations corresponding to each sample in the sample set, the contrast loss is determined; wherein, the sample set includes the several negative samples.
[0164] The joint representations corresponding to each positive sample can be obtained by encoding each positive sample respectively through a process similar to steps 502 to 504. The joint representations corresponding to each sample in the sample set can be obtained by encoding each sample respectively through a process similar to steps 502 to 504. Details are not elaborated here.
[0165] The contrast training needs to increase each first similarity while decreasing each second similarity.
[0166] In one embodiment, determining the contrast loss includes:
[0167] Determining the contrast loss according to the ratio between the summation result of each first similarity and the summation result of each second similarity.
[0168] Preferably, each first similarity and each second similarity can be exponentiated respectively, then the exponentiation results are summed respectively, and then the ratio of the two summation results is taken, and then the logarithmic result of the ratio is determined as the contrast loss.
[0169] In other embodiments, the contrast loss can also be determined according to other methods, as long as each first similarity can be increased while each second similarity can be decreased. For example, the contrast loss can also be determined according to the difference between the summation result of each first similarity and the summation result of each second similarity.
[0170] After determining the contrast loss, the contrast loss and the reconstruction loss can be used to jointly train the image-text joint encoding model.
[0171] In one embodiment, the training loss is determined according to the weighted summation result between the reconstruction loss and the contrast loss. As shown in formula (4).
[0172] Updating the text-image joint encoding model according to the training loss can adjust each trainable parameter in the text-image joint encoding model based on the gradient descent method, which will not be elaborated here.
[0173] It can be understood that steps (a) and (b) only show the process of determining the contrast loss according to a first text-image pair among multiple text-image pairs. In other embodiments, multiple text-image pairs can also be used to respectively determine the intermediate process values of the contrast loss based on the processes described in steps (a) and (b), and the average value of the intermediate process values corresponding to each text-image pair can be used as the final contrast loss.
[0174] The similarity between the representations described in each of the above embodiments can be determined according to, for example, cosine similarity or inner product similarity, which is not limited here.
[0175] The method for training a text-image joint encoding model proposed in the embodiments of this specification constructs positive samples and difficult negative samples of the original text-image pairs by using a large language model and image masking enhancement, and by rewriting the text and the image. Then, based on the conditional text reconstruction task and / or contrast learning task, the text-image joint encoding model is trained to enable it to obtain the ability to jointly encode text-image pairs, and finally a lightweight text-image joint representation model with excellent speed and performance is obtained.
[0176] According to an embodiment of another aspect, there is also provided an apparatus for training a text-image joint encoding model. Figure 7 The schematic block diagram of an apparatus for training a text-image joint encoding model according to an embodiment is shown. This apparatus can be deployed in any device, platform, or device cluster with computing and processing capabilities. As Figure 7 shown, the apparatus 700 includes:
[0177] A first rewriting unit 702 configured to input a first image and a first text included in a first text-image pair into a large language model, and instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain a first summary text;
[0178] A first aggregation encoding unit 704 configured to use the text-image joint encoding model to perform aggregation encoding on a first image representation and a first text representation corresponding to the first image and the first text respectively to obtain a first joint representation;
[0179] A cross-attention unit 706 configured to perform masked cross-attention calculation on the first joint representation and a first summary representation corresponding to the first summary text to obtain a first masked representation;
[0180] A decoding unit 708 configured to decode the first masked representation to obtain a first reconstructed text;
[0181] The training unit 710 is configured to update the text - image joint encoding model according to the training loss, where the training loss at least includes a reconstruction loss determined according to the difference between the first reconstructed text and the first summary text.
[0182] In some possible implementation manners, the apparatus 700 further includes:
[0183] An acquisition unit, configured to acquire at least one positive sample and a plurality of negative samples of the first text - image pair;
[0184] A loss determination unit, configured to determine the contrast loss according to the first similarity between the first joint representation and the joint representations corresponding to the positive samples, and the second similarities between the first joint representation and the joint representations corresponding to the samples in the sample set respectively, where the sample set includes the plurality of negative samples.
[0185] According to an embodiment of another aspect, there is also provided a computer - readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in any of the above embodiments.
[0186] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method described in any of the above embodiments is implemented.
[0187] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiment, since it is basically similar to the method embodiment, it is described relatively simply, and the relevant parts can be referred to the partial description of the method embodiment.
[0188] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In certain implementation manners, multitasking and parallel processing are also possible or may be advantageous.
[0189] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0190] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0191] The specific embodiments described above further elaborate on the objectives, technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for training a text-image joint encoding model, comprising: Inputting a first image and a first text included in a first text-image pair into a large language model, and instructing the large language model to rewrite the first text by introducing semantic content in the first image to obtain a first summary text; Using the text-image joint encoding model, aggregating and encoding a first image representation and a first text representation corresponding to the first image and the first text respectively to obtain a first joint representation; Performing masked cross-attention calculation on the first joint representation and a first summary representation corresponding to the first summary text to obtain a first masked representation; Decoding the first masked representation to obtain a first reconstructed text; Updating the text-image joint encoding model according to a training loss, wherein the training loss at least includes a reconstruction loss determined according to a difference between the first reconstructed text and the first summary text.
2. The method according to claim 1, wherein, The text-image joint encoding model includes at least one self-attention layer and a Q-Former module; obtaining the first joint representation includes: Inputting a first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation; Inputting the first attention representation and a trainable query representation into the Q-Former module for encoding to obtain a first joint representation.
3. The method according to claim 1, wherein, The text-image joint encoding model includes at least one self-attention layer; obtaining the first joint representation includes: Inputting a first concatenated representation obtained by concatenating the first image representation and the first text representation into the at least one self-attention layer for self-attention processing to obtain a first attention representation; Determining a representation corresponding to the first token in the first attention representation as the first joint representation.
4. The method according to claim 1, wherein The first summary representation includes representation vectors of multiple tokens; obtaining the first masked representation includes: Expanding the first joint representation in the spatial dimension to determine a first expanded joint representation having the same size as the first summary representation; Determining a query matrix according to the first expanded joint representation, determining a key matrix and a value matrix according to the first summary representation, and performing masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation.
5. The method according to claim 4, wherein Performing masked cross-attention calculation based on a preset first mask matrix to obtain a first masked representation, including: Obtaining a first score matrix according to a matrix multiplication and normalization result of the query matrix and the key matrix; Applying the first mask matrix to the first score matrix to obtain a first attention matrix; Multiplying the first attention matrix by the value matrix to obtain a first masked representation.
6. The method according to claim 1, wherein The reconstruction loss is determined according to a cross-entropy loss between the first reconstructed text and the first summary text.
7. The method according to claim 1, wherein The training loss further includes a contrastive loss; the method further includes: Obtaining at least one positive sample and a plurality of negative samples of the first text-image pair; Determine the contrastive loss according to the first similarity between the first joint representation and the joint representations corresponding to each positive sample, and the respective second similarities between the first joint representation and the joint representations corresponding to each sample in the sample set; wherein, the sample set includes the plurality of negative samples.
8. The method according to claim 7, wherein Obtain at least one positive sample of the first text-image pair, including: Classify the text-image pair composed of the first image and the first summary text into the positive samples.
9. The method according to claim 7, wherein Obtain a plurality of negative samples of the first text-image pair, including: Input the first text-image pair into a large language model, and instruct the large language model to generate a first forged text with different semantics according to the first text, where the first forged text has no semantic contradiction with the first image; Classify the text-image pair composed of the first image and the first forged text into the negative samples.
10. The method according to claim 7, wherein The first image representation includes the representation vectors of each of the multiple image patches of the first image; obtaining at least one positive sample and a plurality of negative samples of the first text-image pair includes: Perform a clustering operation on the image patches of the first image according to the first image representation to obtain multiple image regions; Determine the respective text-image similarities between the first text and each image region of the first image according to the first image representation and the first text representation; Determine the positive regions as the image regions with the text-image similarity greater than a preset first threshold, and determine the negative regions as the image regions with the text-image similarity less than or equal to the first threshold; Classify the text-image pair composed of the first image representation with the positive regions covered by the mask and the first text into the positive samples, and classify the text-image pair composed of the first image representation with the negative regions covered by the mask and the first text into the negative samples.
11. The method according to claim 10, wherein, Performing a clustering operation on the image patches of the first image according to the first image representation to obtain multiple image regions, including: Perform a pooling operation on the first feature map composed of the representation vectors of each image patch, and perform a clustering operation based on the second feature map after pooling, and determine each clustering cluster as each image region.
12. The method according to claim 10, wherein, Determine the respective text-image similarities between the first text and each image region of the first image, including: For any target region in each image region, determine the similarity between the first text representation and each image patch included in the target region according to the first image representation, and further determine the average similarity as the text-image similarity between the first text and the target region.
13. The method according to claim 7, wherein, Determine the contrastive loss, including: Determine the contrastive loss according to the ratio between the summation result of each first similarity and the summation result of each second similarity.
14. The method according to claim 7, wherein, The training loss is determined according to the weighted summation result between the reconstruction loss and the contrastive loss.
15. An apparatus for training a text-image joint encoding model, including: A first rewriting unit, configured to input the first image and the first text included in the first text-image pair into a large language model, and instruct the large language model to rewrite the first text by introducing the semantic content in the first image to obtain a first summary text; The first aggregation encoding unit is configured to perform aggregation encoding on the first image representation and the first text representation corresponding to the first image and the first text respectively by using a text-image joint encoding model to obtain a first joint representation; The cross-attention unit is configured to perform masked cross-attention calculation on the first joint representation and the first summary representation corresponding to the first summary text to obtain a first masked representation; The decoding unit is configured to decode the first masked representation to obtain a first reconstructed text; The training unit is configured to update the text-image joint encoding model according to the training loss, where the training loss at least includes a reconstruction loss determined according to the difference between the first reconstructed text and the first summary text.
16. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-14.
17. A computing device, comprising a memory and a processor, wherein, Executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-14 is implemented.