Pre-trained Image-Text Model Processing Method and Image-Text Retrieval System

Through image chunking embedding and multi-level encoder decoder structure combined with self-built residual subnet, the problem of insufficient fine-grained image information learning in multimodal pre-training is solved, and pixel-level reconstruction and accurate graphic and text retrieval are realized, which is suitable for the e-commerce field.

CN114821223BActive Publication Date: 2025-07-08HANGZHOU ALIBABA INT INTERNET IND CO LTD

Patent Information

Application Number
CN202210327383.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-07-08
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

Existing multimodal pre-training solutions lack accurate learning of fine-grained image information and text-related optimizations, and cannot meet the growing diverse needs.

Method used

The model structure of image chunking embedding combined with multi-stage downsampling encoder and corresponding upsampling decoder step by step, combined with self-built residual subnetwork, realizes pixel-level reconstruction of masked image blocks in the pre-trained image language network, and is achieved through end-to-end multimodal pre-training.

Benefits of technology

It realizes more accurate graphic and text retrieval, image classification and pixel-level reconstruction, which is especially suitable for e-commerce application scenarios, and improves the ability to understand fine-grained image information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821223B_ABST
    Figure CN114821223B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for processing a pre-trained image-text model and an image-text retrieval system. The method includes: obtaining a masked training sample pair obtained by masking words and image patches in an image-text sample pair; inputting the masked training sample pair into a pre-trained image-text model to obtain loss values output for the masked words, the masked image patches, and an image-text task, wherein the pre-trained image-text model includes a multi-level downsampling encoder and a multi-level upsampling decoder; and adjusting parameters in the pre-trained image-text model according to the loss values. The present invention realizes pixel-level reconstruction of masked image patches in a pre-trained image-language network through a model structure that combines block embedding of images with a multi-level downsampling encoder and a gradually corresponding upsampling decoder. Further, by combining a self-built residual sub-network that realizes embedding of input images and texts with the above encoder-decoder structure, end-to-end multi-modal pre-training is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning, and particularly to a method for processing a pre-trained image-text model and an image-text retrieval system. Background Art

[0002] With the advent of the information age, multimedia data (including text, images, voice, video, etc.) in the Internet has deeply penetrated into all aspects of people's daily lives. How to efficiently parse effective content that follows human semantic understanding from massive multimedia data and give accurate relevant content feedback according to the behavior habits of specific users has become a research hotspot in the academic and industrial fields in recent years.

[0003] For example, traditional single-modal technologies for pure image or pure text retrieval, due to their single data form, can no longer meet the growing diverse needs. In contrast, receiving diverse perceptual content enables an artificial intelligence entity to understand things more comprehensively and efficiently, which also more conforms to the multi-sensory cognitive mode of humans.

[0004] Therefore, a solution for more deeply and accurately mining image-text related information is needed. Summary of the Invention

[0005] One technical problem to be solved by the present disclosure is to provide a method for processing a pre-trained image-text model. This method realizes pixel-level reconstruction of masked image patches in a pre-trained image-language network through a model structure that combines block embedding of images, a multi-level downsampling encoder, and a corresponding upsampling decoder level by level. Further, by combining a self-built residual sub-network that realizes embedding of input images and text with the above encoder-decoder structure, end-to-end multi-modal pre-training is achieved.

[0006] According to a first aspect of the present disclosure, there is provided a method for processing a pre-trained image-text model, including: obtaining a masked training sample pair for masking words and image patches in an image-text sample pair; inputting the masked training sample pair into the pre-trained image-text model to obtain loss values output by the pre-trained image-text model for masked words, masked image patches, and image-text tasks, wherein the pre-trained image-text model includes a multi-level downsampling encoder and a multi-level upsampling decoder; and adjusting parameters in the pre-trained image-text model according to the loss values.

[0007] Optionally, inputting the masked training sample pair into the pre-trained image-text model includes: using an embedding transformation sub-network to generate a text embedding vector and an image embedding vector from the masked training sample pair, wherein the image in the image-text sample pair is divided into multiple blocks, and the image embedding vector is generated based on the divided blocks.

[0008] Optionally, inputting the masked training sample pair into the pre-trained image-text model includes performing the following operations in each level of the downsampling encoder: concatenating the text embedding vectors with the flattened image embedding vectors to obtain a composite vector, sending the composite vector into an encoder unit composed of a multi-head attention sub-network and a feed-forward network after spatial reduction; and splitting the composite vector processed by the encoder unit, obtaining the downsampled text embedding vectors, and reconstructing the remaining part to obtain the downsampled image embedding vectors.

[0009] Optionally, in each level of the downsampling encoder, the image embedding vectors are processed by a convolutional module before being flattened.

[0010] Optionally, inputting the masked training sample pair into the pre-trained image-text model includes: sending the image embedding vectors that have undergone multi-level downsampling into the multi-level upsampling decoder to obtain processed image embedding vectors with the same dimension as the input image.

[0011] Optionally, use the processed image embedding vectors to obtain the loss value of the pre-trained image-text model for the masked image patch, which is used for pixel-level reconstruction of the masked image patch.

[0012] Optionally, set the number of partition blocks occupied by the masked image patch based on the granularity of the image features to be extracted.

[0013] Optionally, obtaining the loss values output by the pre-trained image-text model for the masked word, the masked image patch, and the image-text task includes: based on the output of the multi-level downsampling encoder, calculating the loss values for the masked word and the image-text task output; and based on the output of the multi-level upsampling decoder, calculating the loss value for the masked image patch.

[0014] According to a second aspect of the present disclosure, there is provided an image-text retrieval system, including: a query information acquisition module for acquiring text and / or image information input by a user; and a pre-trained image-text model obtained according to the method of the first aspect, for outputting matching image-text information based on the text and / or image information input by the user.

[0015] According to a third aspect of the present disclosure, there is provided a computing device, including: a processor; and a memory storing executable code thereon, which when executed by the processor, causes the processor to execute the method as described in the first aspect above.

[0016] According to a fourth aspect of the present disclosure, there is provided a non-transitory machine-readable storage medium storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described in the first aspect above.

[0017] Thus, through the block division of the end-to-end image, the variable-granularity masking strategy, and the multi-stage pyramid encoding combined with the subsequent inverted pyramid decoding, this solution realizes the learning of variable-granularity graphic and text information, and is particularly applicable to application scenarios involving more fine-grained feature learning such as fashion pictures. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, wherein in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0019] Figure 1 FIG. shows a schematic flowchart of a pre-trained image-text model processing method according to an embodiment of the present invention.

[0020] Figure 2 FIG. shows an example of a masked training sample pair.

[0021] Figure 3A -B FIG. shows two implementation examples of the pre-trained network structure of the present invention.

[0022] Figure 4 FIG. shows an example of obtaining text tokens and image patches based on the original image-text sample pair.

[0023] Figure 5 FIG. shows a masked vision-language transformer for fashion cross-modal representation.

[0024] Figure 6 FIG. shows a comparative example of the prior art and the PVT-based masking strategy of the present invention.

[0025] Figure 7 FIG. shows an example of the image-text retrieval system of the present invention.

[0026] Figure 8 FIG. shows a schematic structural diagram of a computing device that can be used to implement the above pre-trained image-text model processing method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure will be more thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0028] As described above, how to efficiently parse effective content that follows human semantic understanding from a large amount of multimedia data and give accurate relevant content feedback according to the behavior habits of specific users (including application scenarios such as search, recommendation, and advertising) has become a research hotspot in academia and industry in recent years.

[0029] Traditional unimodal technologies (such as pure image or pure text retrieval) can no longer meet the growing diverse needs due to their single data form. In contrast, receiving diverse perceptual content enables an artificial intelligence agent to understand things more comprehensively and efficiently, which is also more in line with the multi-sensory cognitive mode of humans. Multimodal technologies have also been favored by the industry for their excellent performance in many rich semantic understanding tasks.

[0030] In the context of deep learning, how to make a machine give feedback in a way similar to human intelligence is a hot issue in the field. Most existing multimodal pre-training solutions focus on the improvement of the Transformer module and mostly focus on general field image-text tasks in applications. There is a lack of an optimized solution for accurately learning fine-grained image information and relating it to text.

[0031] Therefore, the present invention provides a pre-training image-text model processing method, which realizes pixel-level reconstruction of masked image patches in a pre-trained image-language network through a model structure that combines block embedding of images, a multi-level downsampling encoder, and a corresponding upsampling decoder. Further, by combining a self-built residual sub-network that realizes the embedding of input images and text with the above encoder-decoder structure, end-to-end multimodal pre-training is achieved. A system using this model can achieve more accurate image-text retrieval, image classification, and even pixel-level reconstruction of masked parts, and is particularly suitable for various e-commerce application scenarios.

[0032] Figure 1 FIG. shows a schematic flowchart of a pre-training image-text model processing method according to an embodiment of the present invention.

[0033] In step S110, a masked training sample pair for masking words and image patches in an image-text sample pair is obtained. Figure 2 FIG. shows an example of a masked training sample pair. As shown in the figure, the image-text sample pair can be Figure 2 the image shown on the left and the corresponding text description ("Women’s Sleeveless Long Dress"), and a masked training sample pair (such as T and V shown in the figure) can be obtained by masking specific image patches in the image (such as the dashed box shown in the figure) and specific words in the text description (such as [MASK] shown in the figure).

[0034] In step S120, the masked training sample pair is input into the pre-trained image-text model. The pre-trained image-text model may include a multi-level downsampling encoder and a multi-level upsampling decoder, and thereby obtain the loss values output by the pre-trained image-text model for the masked word, the masked image patch, and the image-text task. Subsequently, in step S130, according to the loss values, the parameters in the pre-trained image-text model are adjusted.

[0035] It can be seen that the pre-trained image-text model in the present invention has a pyramid-shaped multi-level encoder (i.e., the input feature map becomes smaller and smaller) and an inverted pyramid-shaped multi-level decoder (i.e., the reduced feature map gradually increases), and can have three different training tasks, the prediction task for the masked word (i.e., the language task, for example, the MLM task described below), the image-language task (also called the vision-language task, i.e., the VL task, such as the ITM task described below), and the task for the masked image patch (i.e., the vision task). In the present application, the vision task can be particularly implemented as a masked image reconstruction task (i.e., the MIR task described in detail below).

[0036] In one embodiment, the outputs of the multi-level encoder can be directly used for the language task, the VL task, and the vision task. The image-language pre-trained network trained thereby can also perform well in downstream tasks.

[0037] In a preferred embodiment, the image embedding vector output by the multi-level encoder can be further input into the multi-level upsampling decoder to obtain a processed image embedding vector with the same dimension as the input image. At this time, obtaining the loss values output by the pre-trained image-text model for the masked word, the masked image patch, and the image-text task may include: obtaining the loss values output for the masked word and the image-text task based on the output of the multi-level downsampling encoder; and obtaining the loss value for the masked image patch based on the output of the multi-level upsampling decoder. Thus, the loss value for the masked image patch of the pre-trained image-text model is obtained using the processed image embedding vector for pixel-level reconstruction of the masked image patch.

[0038] Figure 3A -B shows two implementation examples of the pre-trained network structure of the present invention. As Figure 3A and 3BAs shown, the input text and the image as the text description object are processed to obtain multiple text tokens and multiple image patches. These text tokens and image patches are then converted into text embeddings ("embeddings" or "embedding vectors") and image embeddings. Subsequently, the image embedding vectors and text embedding vectors are sent to the multimodal PVT (Pyramid Image Transformer). Here, multimodality refers to the processing of two types of inputs, text and images, and PVT refers to the pyramid model structure composed of multi-level downsampling Transformer encoders. In Figure 3A In the example, the loss value can be directly calculated based on the output of multimodal PVT, thereby realizing the language task, VL task and image task of the pre-trained model. Figure 3B In the preferred embodiment shown, the image embedding vector output by the multimodal PVT can be subjected to a multi-level upsampling Transformer decoder of an inverted pyramid type, thereby obtaining a processed image embedding vector of the same dimension as the input image. Subsequently, the MIR (image reconstruction) task of pixel-level reconstruction can be achieved.

[0039] Figure 4 An example of obtaining text tokens and image patches based on original image-text sample pairs is shown. As shown in the figure, the original sample includes an image ( Figure 4 The image of a man wearing a hooded zip-up shirt and jeans on the lower right side) and the description of the image ( Figure 4 The description of the zip-up hooded sweater worn by the man on the lower left side is “Long sleeve hoodle in black. Drawstring closure at hood. Zip closure and patch packets atfront. Rib knit sleeve cuffs and hem…”. The original sample above can be obtained from an e-commerce website selling clothing, for example. The image is a product picture, and the description is the product title or description.

[0040] Subsequently, in order to obtain a text token, the input text may be first tokenized word by word to obtain a token sequence, and then the entire word is masked to obtain a masked text token.

[0041] Similarly, the patch processing of the image is different from the common RoI (Region of Interest) method in the art. Instead, the image is sliced into small blocks (usually squares) of the same pixel size, and each patch can be regarded as an "image token". For each patch, the output of PatchNet (Patch Network) can be regarded as patch features, and then these features can be masked to obtain the masked image patches. These patches are naturally sorted, and the spatial position of the patches can be used for position embedding vectors in subsequent processing (see Figure 5 ).

[0042] In one embodiment of the present invention, text tokens and image patches can be fed into a standard ResNet to obtain text embedding vectors and image embedding vectors. In another more preferred embodiment, an embedding transformation sub-network included in the pre-trained image-text model of the present invention can be used to achieve the embedding vectorization of the input sample pair. Thus, inputting the masked training sample pair into the pre-trained image-text model may include: using the embedding transformation sub-network (e.g., a ResNet with a simplified structure written by the inventor himself) to generate text embedding vectors and image embedding vectors from the masked training sample pair, wherein the image in the image-text sample pair is divided into multiple patches, and the image embedding vectors are generated based on the divided patches.

[0043] This enables the input to the multi-level downsampling encoder to be Figure 3A the embedding shown in -B, that is, the text embedding vector and the image embedding vector. By using a self-written embedding transformation sub-network (e.g., a residual network with a simplified structure compared to the standard ResNet), the offline processing of the input image-language pair can be avoided, thus making it possible to perform downstream tasks online. Further, although the embedding transformation sub-network of the present invention has a simplified structure, since this sub-network can be regarded as included in the pre-trained image-text model of the present invention, that is, the parameters of this sub-network can be trained together with the parameters of the downstream PVT and the optional multi-level decoder, it can still perform excellently in downstream tasks.

[0044] In the present invention, each stage of the downsampling encoder can be regarded as a Transformer encoder capable of performing downsampling. To this end, inputting the masked training samples into the pre-trained image-text model may include the following operations in each stage of the downsampling encoder: concatenating the text embedding vectors with the flattened image embedding vectors to obtain a composite vector; feeding the composite vector into an encoder unit composed of a multi-head attention sub-network and a feed-forward network after spatial reduction; and splitting the composite vector processed by the encoder unit to obtain the downsampled text embedding vector (which can also be referred to as the "text embedding vector"), and reconstructing the remaining part to obtain the downsampled image embedding vector. Thus, the concatenation of the text and image embedding vectors is fed into the encoder for processing, and after processing, splitting and shape reconstruction (only for the image embedding) are performed to complete the processing of each stage. Preferably, in each stage of the downsampling encoder, the image embedding vector is processed by a convolutional module (such as a Conv2D block) before flattening. Thus, the downsampling-friendly characteristics of the convolutional operation are utilized to process the image embedding. Further, the image embedding vector that has undergone multiple stages of downsampling can be combined as described above Figure 3B and fed into the multi-stage upsampling decoder to obtain a processed image embedding vector with the same dimension as the input image. In this way, the loss value of the pre-trained image-text model for the masked image patch can be obtained using the processed image embedding vector for pixel-level reconstruction of the masked image patch.

[0045] In one embodiment, the size of the masked image patch may be different from the size of the patch. In other words, the number of partition blocks occupied by the masked image patch can be set based on the granularity of the image features to be extracted. For example, it can be determined whether the masked image patch is to mask 1 (1x1), 4 (2x2), 9 (3x3), or even more patches based on the setting of the α value described below.

[0046] To better clarify the principle of the present invention, a preferred end-to-end model embodiment of the present invention will be described below in combination with the Figure 5 encoder structure described in detail. Figure 5 An example of an end-to-end VL pre-training model according to the present invention is shown.

[0047] Figure 5A masked vision - language transformer (M - ViLT) for cross - modal representation in fashion (especially in the e - commerce field) is shown. In the model of the present invention, the Pyramid Vision Transformer (PVT) is used instead of BERT, and a masked image reconstruction (MIR) strategy is introduced in PVT, making M - ViLT the first end - to - end framework in the fashion field. M - ViLT is an extensible and convenient architecture that can accept raw multi - modal inputs without an additional pre - processing model (such as ResNet), implicitly models visual - language alignment, and can be easily generalized to various downstream matching and generation tasks.

[0048] To solve the problems of insufficient granularity and poor transferability in the prior art, the present invention constructs an end - to - end VL framework specifically for the fashion field. The overall process of M - ViLT is as Figure 5 shown. Figure 5 The example uses a four - level encoder and generates features of different sizes. Two keys of the proposed architecture are the multi - modal encoder and the pre - training objective.

[0049] Multimodal Encoder

[0050] As Figure 5 shown, M - ViLT allows visual and language inputs. In terms of language, first, the title of a fashion product (e.g., clothing accessories and other goods) is tokenized, and a specific token [MASK] is used to randomly mask the tokens of the title at a masking ratio r l After the masking process, a series of word tokens are obtained. Then, a specific [CLS] token is inserted at the head of this sequence. In addition, if the length is less than 128, the sequence is padded to a unified length L using [PAD] tokens. This process generates language input ids

[0051] In terms of vision, can be regarded as the visual input, where H and W represent the height and width of the given input. This input is divided into multiple grid - like patches (small pieces) where is the total number of patches, and P represents the patch size. Similarly, the masking ratio r v is used to mask the segmented patches. More details about the masking strategies for the above - mentioned language and vision parts are provided in the following description of "pre - training objective". The above - mentioned multi - modal inputs are embedded and fed into the subsequent four VL interaction stages (i.e., k ∈ {1, 2, 3, 4}). In the first stage, text embedding vectors T 1 and visual embedding vectors V 1。Regarding the subsequent stage, only the k-th stage is considered to simplify the description. As Figure 5 shown at the bottom, first, the text embedding vector is embedded into the language hidden feature (i.e., the linear Embed shown in the figure), and the formula is:

[0052]

[0053] where and are the learnable linear embedding and position embedding matrices respectively. D k is the size of the hidden feature embedding. The visual embedding vector is where R k represents the spatial reduction factor of the visual embedding. To obtain the pyramid visual features, V k is then embedded and flattened (flattened) into the visual hidden feature through a two-dimensional projection (i.e., the Conv2D block) (i.e., the spatial Embed shown in the figure). Specifically, this projection uses the convolutional kernel (where the kernel size is K k , and the stride is S k ) to force the network to reduce the equivalent spatial dimension from to This can be expressed as:

[0054]

[0055] where represents the position embedding matrix. Subsequently, the two VL hidden features z k = <m k ; n k > are concatenated (Concat) and fed into multiple (M k ) VL transformer encoders. Each encoder contains a multi-head self-attention layer with spatial reduction (i.e., the reduction box shown in the figure), a multi-layer perceptron, and layer normalization. Finally, the encoded multi-modal feature z k+1 = <m k+1 ; n k+1 > is obtained and divided into the language part T k+1 = m k+1 and the visual part V k+1 = n k+1 , where the Reshape(·) operation includes restoring the spatial dimension of the given feature.

[0056] After four VL interaction levels, four text embedding vectors and four pyramid visual embedding vectors The following table shows the detailed hyperparameter settings.

[0057] Hyperparameter k=1 k=2 k=3 k=4 <![CDATA[Number of layers M k > 2 2 2 2 <![CDATA[Hidden dimension D k > 64 128 320 512 <![CDATA[Reduced size R k > 4 8 16 32 <![CDATA[Core size K k > 4 2 2 2 <![CDATA[Step S k > 4 2 2 2

[0058] Pretraining Objective

[0059] To obtain a multimodal representation with sufficient discriminability, three pre-training tasks are adopted to establish the relationships within and between the most primitive VL modalities, including a visual task (implemented as masked image reconstruction, MIR), a language task (implemented as masked language modeling, MLM), and a vision-language task (implemented as image-text matching, ITM).

[0060] Task 1: Masked Image Reconstruction (MIR).

[0061] The present invention attempts to construct pixel-to-pixel relationships from the perspective of a generation task, thereby promoting the scalability of visual representations, and designs masked image reconstruction (MIR) to achieve pixel-level reconstruction. To help the model better learn through MIR, the pyramid features of the PVT architecture are utilized to design a flexible masking strategy. The PVT-based architecture masks the input image according to a masked unit matrix containing small-grained patches. Given a patch sequence The masked sequence V \Φ can be defined as:

[0062]

[0063] where represents a function (or process) of the masking strategy, q is the randomly selected area of the masking unit, and [ZERO] indicates using the pixel value zero (0) to fill the selected area. The masking unit is derived from the following metric function:

[0064]

[0065] where each value in a set of integers Φ is randomly selected from the range [1, Q] at a masking ratio r v is the total number of masking units. is the total number of masking units. Figure 6 Shows a comparative example of the prior art and the PVT-based masking strategy of the present invention. Compared with the ViT-based masking strategy on the left side of the figure that can only mask one Patch (size PxP), the PVT-based method in the present invention is more flexible. As shown on the right side of the figure, the PVT-based method can be based on finer-grained Patches and flexibly select the masking coefficient. For example, the masking coefficient α can be defined as 1 to 8, thereby enabling the basic masking unit (MaskingUnit) to mask (α × P) 2Range of images. In the implementation of the present invention, α = 4 can be default set to capture finer-grained semantics.

[0066] Since the smooth-l1 loss is less sensitive to outliers, it can be chosen as the pre-training objective to reconstruct the entire image through the masked sequence V \Φ This pre-training task is defined as:

[0067]

[0068] where I′ (x,y) and I (x,y) represent the pixels at the coordinates (x, y) of the reconstructed image I′ and the input image I, respectively. Parameterized by the learnable weight W MIR The function represents a standard four-level U-Net decoder that accepts four pyramid visual embedding vectors as inputs. In other words, although not shown in Figure 5 , the visual embedding vectors fed into the MIR are not the vectors of the fourth-level output shown in Figure 5 , but the vectors output after being followed by a four-level inverted pyramid decoder as shown in Figure 3B . Therefore, in the process of obtaining the MIR task loss, the finally obtained processed visual embedding vectors have passed through a four-level pyramid encoder and a four-level inverted pyramid decoder, and thus are obtained through the U-Net network.

[0069] Task 2: Image-Text Matching (ITM).

[0070] The additional classification embedding vector in the last-level text embedding vector T 4 can be used to couple the representations from the VL multimodality. The function can be used to represent the fully connected (FC) and softmax layers, parameterized by the weight W ITM . Output a binary classification probability vector indicating whether the input image and text description match (i.e., positive pair) or not (i.e., negative pair). The positive pairs are selected from the same product category, while the negative pairs are randomly selected from different entries. This task is finally constrained by the binary cross-entropy loss function:

[0071]

[0072] where y ITM represents the true label, i.e., 1 indicates a matching pair and 0 indicates a non-matching pair.

[0073] It should be understood that in other implementations, other tasks such as text-image matching (TIM) can also be used to train the model for learning information between multiple modalities.

[0074] Task 3: Masked Language Modeling (MLM)

[0075] Specific tokens [MASK] can be randomly used to replace the original text tokens. The goal of MLM is to use the unmasked tokens and patches to predict the text content of the masked tokens. Given a tokenized sequence T = {t1, …, t L}, the masked sequence is denoted as T \i = {t1, …, [MASK] i , …, t L}. Cross-entropy loss can be used to model this goal:

[0076]

[0077] where represents the prediction probability for each masked token [MASK] \i using T i . The function represents the parameters W MLM of the classifier. The final pre-training goal of the proposed M-ViLT is a combination of three goals:

[0078]

[0079] Overall, the M-ViLT based on the above preferred embodiments of the present invention can achieve end-to-end training, that is, it provides an end-to-end framework including a feature extractor and a multi-modal matching network. This framework can especially incorporate a large amount of prior information in the e-commerce field (such as information about merchants, products, categories, etc.), thereby further improving the multi-modal matching effect in downstream tasks. Further, the pixel-level image reconstruction achieved through the MIR task can enable the model to have a finer-grained understanding of the image side, thus significantly improving the model performance.

[0080] In the multi-modal encoder part of the present invention, the picture and text information are encoded. By introducing the PVT pre-trained model, the multi-modal information of pictures and texts is fully interacted in the PVT. In terms of pre-training tasks, three pre-training tasks are introduced. The MLM task is mainly responsible for language model learning, the TIM is mainly responsible for image and text matching learning, and the MIR is mainly responsible for image reconstruction task learning.

[0081] The pre-trained image-text model processed according to the present invention can be used for different downstream tasks, such as picture and text mutual retrieval, image classification, and image generation.

[0082] For example, a downstream task can be text-image retrieval (TIR). The TIR task requires the model to find the text with the highest similarity value to different query images. Specifically, the product title and its corresponding image can be used as a positive image-text pair, while the negative image pairs are randomly selected from a pool of unmatched images. To increase accuracy, a set of image-text candidates (i.e., one positive pair and 100 negative pairs) can be restricted to the same subcategory to make them as similar as possible.

[0083] Another downstream task is image-text retrieval (ITR). As the inverse process of the TIR task, the ITR task aims to retrieve the matching images for a given sequence of product description text entries. The above two-way retrieval tasks (i.e., TIR and ITR) can become the main application scenarios of cross-modal.

[0084] Other downstream tasks can also include category recognition tasks and masked image generation (MIG) tasks. Category recognition can include: main category recognition (M-CR) and subcategory recognition (S-CR). These tasks provide the specific category of the query product, and the model should be able to recognize the differences at different granularity levels through different downstream tasks: for example, 48 main categories and 122 subcategories. After the class embedding vector in the last text embedding vector T 4 two independent FC layers can be added to generate the final probabilities for two different recognition tasks, and this process requires additional fine-tuning using the recognition labels. The MIG task can be regarded as a pixel-level reconstruction task. Then, using the uncovered areas as visual cues, the model can be required to recreate the entire image.

[0085] After further training for these downstream tasks, the pre-trained image-text model of the present invention can be deployed in a text-image retrieval system. For this purpose, further, the present invention can also be implemented as a text-image retrieval system. Figure 7 An example of the text-image retrieval system of the present invention is shown. The text-image retrieval system can include a query information acquisition module for acquiring the text and / or image information input by the user. Further, the text-image retrieval system can include a pre-trained image-text model (PLM) obtained according to the method described above for outputting the matching text-image information based on the text and / or image information input by the user. The pre-trained image-text model of the present invention can be deployed in the retrieval system of an e-commerce website so that more accurate information can be returned when the user performs text-image retrieval. The pre-trained image-text model processed according to the present invention can deeply learn the mutual relationship between images and text, thereby improving the accuracy of text-image retrieval.

[0086] Figure 8FIG. shows a schematic structural diagram of a computing device that can be used to implement the above-mentioned pre-trained image-text model processing method according to an embodiment of the present invention.

[0087] Refer to Figure 8 , the computing device 800 includes a memory 810 and a processor 820.

[0088] The processor 820 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 820 can include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, the processor 820 can be implemented using custom circuits, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0089] The memory 810 can include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 820 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 810 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 810 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), a read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, a super density disc, a flash memory card (such as SD card, min SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0090] Executable code is stored on the memory 810, and when the executable code is processed by the processor 820, it can cause the processor 820 to execute the pre-trained image-text model processing method described above.

[0091] The pre-trained image-text model processing method according to the present invention and the resulting image-text retrieval system have been described in detail above with reference to the accompanying drawings.

[0092] The present invention innovatively eliminates the need for complex preprocessing of images in the multi-modal encoder part (i.e., uses a self-written simplified residual network instead of the standard ResNet). M-ViLT realizes consistent offline and online logic through end-to-end online patch splitting of pictures and combines a masking strategy, which is beneficial to the actual deployment of the model, such as in the e-commerce field. Further, the MIR task defined by the present invention can force a more fine-grained understanding of the image through pixel-level feature reconstruction, and thus better meets the needs of the e-commerce field that relies on a more fine-grained understanding of pictures.

[0093] In addition, the method according to the present invention can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing the above steps defined in the above method of the present invention.

[0094] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium), on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), it causes the processor to execute each step of the above method according to the present invention.

[0095] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both.

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0097] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements to the technology in the market, or to enable other ordinary skilled persons in the art in the field to understand the embodiments disclosed herein.

Claims

1. A method for processing a pre-trained image-text model, comprising: Obtaining a masked training sample pair for masking words and image patches in an image-text sample pair, wherein the masked training sample pair is obtained by masking specific image patches in the image and specific words in the text description; Inputting the masked training sample pair into the pre-trained image-text model, and obtaining a loss value output by the pre-trained image-text model for the masked words, masked image patches, and image-text tasks, wherein the pre-trained image-text model includes a multi-level downsampling encoder and a multi-level upsampling decoder, and wherein, based on the output of the multi-level downsampling encoder, a loss value for the masked words and the image-text tasks is obtained, and based on the output of the multi-level upsampling decoder, a loss value for the masked image patches is obtained; Adjusting the parameters in the pre-trained image-text model according to the loss value.

2. The method according to claim 1, wherein Inputting the masked training sample pair into the pre-trained image-text model includes: Using an embedding transformation sub-network to generate a text embedding vector and an image embedding vector from the masked training sample pair, wherein the image in the image-text sample pair is divided into multiple patches, and the image embedding vector is generated based on the divided patches.

3. The method according to claim 2, wherein, Inputting the masked training sample pair into the pre-trained image-text model includes performing the following operations in each level of the downsampling encoder: Concatenating the text embedding vector with the flattened image embedding vector to obtain a composite vector; Feeding the composite vector through spatial reduction into an encoder unit composed of a multi-head attention sub-network and a feed-forward network; And Splitting the composite vector processed by the encoder unit to obtain the downsampled text embedding vector, and reconstructing the remaining part to obtain the downsampled image embedding vector.

4. The method according to claim 3, wherein, In each level of the downsampling encoder, the image embedding vector is processed by a convolutional module before being flattened.

5. The method according to claim 3, wherein Inputting the masked training sample pair into the pre-trained image-text model includes: Feeding the image embedding vector that has undergone multi-level downsampling into the multi-level upsampling decoder to obtain a processed image embedding vector with the same dimension as the input image.

6. The method according to claim 5, wherein Using the processed image embedding vector to obtain the loss value of the pre-trained image-text model for the masked image patches for pixel-level reconstruction of the masked image patches.

7. The method according to claim 5, wherein Setting the number of divided patches occupied by the masked image patches based on the image feature granularity to be extracted.

8. A graphic-text retrieval system, comprising: A query information acquisition module for acquiring text and / or image information input by a user; And A pre-trained image-text model obtained according to the method as described in any one of claims 1-7, for outputting matching graphic-text information based on the text and / or image information input by the user.

9. A computing device, comprising: A processor; And A memory having executable code stored thereon, which when executed by the processor causes the processor to execute the method as described in any one of claims 1-7.

10. A non-transitory machine-readable storage medium having executable code stored thereon, which when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Semantic representation model pre-training method and device and electronic equipment

    CN114186564A

Cited By

  • A cross-modal image-text retrieval system based on contrastive learning

    CN122615101A