Methods, apparatus, devices, and media for image coding learning and application
Patent Information
- Application Number
- CN202310032610.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-01-10
AI Technical Summary
[0009]应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其他特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN116030318B_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and in particular to methods, apparatuses, devices, and computer-readable storage media for image coding learning and applications. Background Technology
[0002] With the development of machine learning technology, machine learning models can now be used to perform tasks in a variety of application environments. Model-based vision tasks are used to process visual data, such as images and videos. Examples of vision tasks include, but are not limited to, image classification, object detection, and semantic segmentation. In vision task models, the challenge lies in how to extract features that accurately represent image data. Models used to extract feature representations of images are typically called image encoders. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for learning image encoding is provided. The method includes: extracting image feature representations of sample images using an image encoder to be trained; extracting text feature representations of sample text sequences associated with sample images using a text encoder; generating predicted text sequences based on the text feature representations and image feature representations using a text decoder; and training the image encoder based at least on the text error between the predicted text sequences and the sample text sequences.
[0004] In a second aspect of this disclosure, a method for image coding applications is provided. The method includes: acquiring an image encoder trained according to the method of the first aspect; extracting an image feature representation of a target image using the acquired image encoder; and performing a predetermined visual task for the target image based on the image feature representation.
[0005] In a third aspect of this disclosure, an apparatus for image encoding learning is provided. The apparatus includes: an image feature extraction module configured to extract image feature representations of a sample image using an image encoder to be trained; a text feature extraction module configured to extract text feature representations of a sample text sequence associated with the sample image using a text encoder; a text generation module configured to generate a predicted text sequence based on the text feature representations and the image feature representations using a text decoder; and a training module configured to train the image encoder based at least on a text error between the predicted text sequence and the sample text sequence.
[0006] In a fourth aspect of this disclosure, an apparatus for image coding applications is provided. The apparatus includes: an encoder acquisition module configured to acquire an image encoder trained according to the method of the first aspect; a feature extraction module configured to extract an image feature representation of a target image using the acquired image encoder; and a task execution module configured to perform a predetermined visual task on the target image based on the image feature representation.
[0007] In a fifth aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the methods of the first aspect and / or the second aspect.
[0008] In a sixth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the methods of the first and / or second aspects.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0012] Figures 2A to 2C A schematic diagram of the example model training architecture is shown;
[0013] Figure 3 A schematic diagram of a training architecture for image coding learning according to some embodiments of the present disclosure is shown;
[0014] Figure 4 A schematic diagram of a simplified training architecture for image encoding learning according to some embodiments of the present disclosure is shown;
[0015] Figure 5 A schematic diagram of an image encoding / decoding process during training is shown according to some embodiments of the present disclosure;
[0016] Figure 6A schematic diagram of attention matrix processing for text encoding during training is shown according to some embodiments of the present disclosure;
[0017] Figure 7 A flowchart of a process for learning image coding according to some embodiments of the present disclosure is shown;
[0018] Figure 8 A flowchart illustrating a process for an image encoding application according to some embodiments of the present disclosure is shown;
[0019] Figure 9 A block diagram of an apparatus for image coding learning according to some embodiments of the present disclosure is shown;
[0020] Figure 10 A block diagram of an apparatus for an image encoding application according to some embodiments of the present disclosure is shown; and
[0021] Figure 11 A block diagram of an electronic device that may implement one or more embodiments of the present disclosure is shown. Detailed Implementation
[0022] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0023] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0027] As an optional but non-limiting embodiment, in response to receiving a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understood that the above notification and user authorization acquisition process is merely illustrative and does not constitute a limitation on the embodiments of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the embodiments of this disclosure.
[0029] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0030] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.
[0031] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values to determine the corresponding output.
[0032] Figure 1 A schematic diagram is shown of a model training and application environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 The environment 100 illustrates three distinct phases of the model: a pre-training phase 102, a fine-tuning phase 104, and an application phase 106. A testing phase, not shown in the figure, may also occur after the pre-training or fine-tuning phases.
[0033] In the pre-training phase 102, the model pre-training system 110 is configured to perform pre-training of the image encoder 105 using the training dataset 112. At the start of pre-training, the image encoder 105 may have initial parameter values. The pre-training process involves updating the parameter values of the image encoder 105 to desired values based on the training data. During pre-training, one or more pre-training tasks 107-1, 107-2, etc., can be designed. These pre-training tasks are used to assist in updating the parameters of the image encoder 105. Some pre-training tasks may require connecting the image decoder 105 to the image decoder associated with the pre-training task.
[0034] In the pre-training phase 102, the image encoder 105 can learn strong generalization capabilities using large-scale training data. After pre-training, the parameter values of the image encoder 105 have been updated to include the pre-trained parameter values. The pre-trained image encoder 105 can extract feature representations of images with relatively high accuracy.
[0035] The pre-trained image encoder 105 can be provided to the fine-tuning stage 104, where it is fine-tuned by the model fine-tuning system 120 for different downstream tasks. Downstream tasks can involve various visual tasks, such as image classification, object detection, and semantic segmentation. In some embodiments, depending on the specific downstream task, the pre-trained image encoder 105 can be connected to the image decoder 127 required by the downstream task to construct a downstream task model 125. This is because the required output may differ for different downstream tasks.
[0036] In the fine-tuning phase 104, the parameter values of the image encoder 105 are further adjusted using the training dataset 122. If necessary, the parameters of the image decoder 127 may also be adjusted. The image encoder 105 can extract feature representations from the input image and text data and provide them to the image decoder 127 to provide the output for the corresponding task.
[0037] During fine-tuning, the corresponding training algorithm is also used to update and adjust the parameters of the overall model. Since the image encoder 105 has learned a lot from the training data during the pre-training phase, a downstream task model that meets the expectations can be obtained using a small amount of training data during the fine-tuning phase 104. In some embodiments, during the pre-training phase 102, a specific image decoder may have been constructed according to the goal of the pre-training task. In this case, if the image decoder required in the downstream task is the same as the image decoder constructed during pre-training, the pre-trained image encoder 105 and the image decoder can be directly used to compose the corresponding downstream task model. In this case, the downstream task model may not require fine-tuning, or may only require fine-tuning with a small amount of training data.
[0038] In application phase 106, the obtained downstream task model 125, with trained parameter values, can be provided to the model application system 130 for use. In application phase 106, the downstream task model 125 can be used to process corresponding inputs in the real-world scene and provide corresponding outputs. For example, the image encoder 105 in the downstream task model 125 receives the input target image 132 to extract corresponding feature representations. The extracted feature representations are provided to the image decoder 127 to determine the corresponding visual task output.
[0039] exist Figure 1 In this system, the model pre-training system 110, the model fine-tuning system 120, and the model application system 130 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0040] It should be understood that Figure 1The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model pre-training system 110, the model fine-tuning system 120, and the model application system 130 may be integrated in the same system or device. Implementations of this disclosure are not limited in this respect.
[0041] In some embodiments, the training phase of the image encoder 105 may not be divided into... Figure 1 Instead of the pre-training and fine-tuning stages shown, it can directly build downstream task models based on the task and use a large amount of training data to train the feature extraction model.
[0042] Visual self-supervised learning (VSM) demonstrates superior performance in visual representation learning. Instead of requiring labeled information, VSM utilizes the input samples (e.g., images or image pairs) themselves as supervision to pre-train the model (e.g., an image encoder). Compared to random initialization, using a model trained through self-supervised learning as initialization can yield significant advantages on many visual tasks.
[0043] However, the amount of data and the supervision method in self-supervised learning affect the amount of information learned. In self-supervised learning, it is expected that the pre-trained model should learn a large amount of information, which is more conducive to extracting useful information from downstream tasks. Therefore, the amount of data and the supervision method in self-supervised learning affect the effect of pre-training transfer learning.
[0044] Figures 2A to 2C A schematic diagram of the example model training architecture is shown.
[0045] Figure 2A The training architecture is based on a unimodal discriminative pre-training / training method. This method involves argumentation of sample images, such as rotation, cropping, and flipping, and uses the enhanced images as positive samples, while other sample images in the training dataset are used as negative samples. After constructing the positive and negative samples, discriminative learning is performed based on a contrastive loss function. The image encoder 210 is used to extract feature representations of a pair of images (the sample image and the enhanced image). The goal of discriminative learning is to determine whether the extracted feature representations match. It is expected that the extracted feature representations for positive sample pairs are similar and matching, while the extracted feature representations for negative sample pairs are dissimilar and mismatched. However, this pre-training method only uses image data and cannot utilize massive amounts of image and text data. Furthermore, this pre-training architecture is based on discriminative contrastive learning, resulting in a limited amount of information learned by the image encoder, which is not conducive to downstream fine-tuning tasks.
[0046] Figure 2B The training architecture is based on a unimodal generative pre-training / training method. This method involves an image encoder 220 extracting feature representations from the input, and an image decoder 222 generating images based on these extracted feature representations. The only training data required in this process is the sample images. For example, the sample images can be partially masked and input into the image encoder 220 for feature extraction, requiring the image decoder 222 to decode the original sample images from the extracted feature representations. While this training method is beneficial for learning more information, the amount of information learned is limited because it only utilizes image data during training.
[0047] Figure 2C The training architecture is based on a multimodal discriminative pre-training / training method. This training architecture is consistent with... Figure 2A The training architecture is similar, but the difference lies in that it does not use single-modality data to construct positive samples. Instead, it uses pre-obtained matched sample image and text pairs as positive samples, and other sample images or text pairs within the training dataset as negative samples. After the positive and negative samples are constructed, discriminative learning is performed based on a contrastive loss function. For a pair of sample images and text, the image encoder 230 extracts the image feature representation of the sample image, while the text encoder 240 extracts the text feature representation of the sample text. The goal of discriminative learning is to distinguish whether the extracted image feature representation and text feature representation match. It is expected that the extracted image and text feature representations for positive sample pairs are similar and matched, while the extracted image and text feature representations for negative sample pairs are dissimilar and mismatched. Although Figure 2C The training architecture uses image and text data, but this method is based on contrastive learning, and the amount of information the model learns from image and text data is relatively limited, which is not friendly to downstream fine-tuning tasks.
[0048] As mentioned earlier, the amount of information a model can learn during training is related to the amount of training data and also to the supervision method used. Some discriminative training architectures cannot learn sufficient visual information from the training data, and... Figure 2B For example, a single-modal generative training architecture can only learn from single-modal training data.
[0049] According to an example embodiment of this disclosure, an improved scheme for image encoder learning is provided. This scheme utilizes multimodal training data (image modality and text modality data) and performs training for the image encoder based on generative self-supervised learning. The training data includes associated (or matched) sample images and sample text images. In the training of the image encoder, the image feature representations extracted by the image encoder from the sample images are used to guide the text encoder / decoder in performing a text generation task. The loss from the text generation task is used to train the image encoder. This training method can utilize multimodal data, enabling the image encoder to learn useful feature information from image and text data. The trained image encoder has stronger feature extraction capabilities and can be adapted to various downstream vision tasks.
[0050] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0051] Figure 3 A schematic diagram of a training architecture 300 for image encoding learning according to some embodiments of the present disclosure is shown. In some embodiments, the training architecture 300 can be implemented in Figure 1 The model pre-training system 110 is used for pre-training the image encoder. In some embodiments, in addition to pre-training and fine-tuning, the image encoder and image decoder can also be directly trained and used as a whole model. In such embodiments, the training architecture 300 can be implemented in other model training systems. Hereinafter, a pre-training architecture will be used as an example.
[0052] like Figure 3 As shown, the training architecture 300 involves an image encoder 310, an image decoder 320, a text encoder 330, and a text decoder 340. The training objective of this architecture 300 is to enable the image encoder 310 to learn as much information as possible from the training data, thereby extracting accurate features from the input image and making it applicable to different visual tasks. In this paper, the encoder, also referred to as a feature extractor, is configured to extract feature representations of the corresponding modal input. The decoder is configured to generate the corresponding output from the feature representations provided by the encoder.
[0053] Training architecture 300 performs training based on a self-supervised learning approach. Training data includes sample images and sequences of sample images. Each pair of sample images and sequences can be correlated or matched. Here, correlation or matching means that the sample text sequence accurately describes the visual information presented by the sample image. In some embodiments, sample images can include static images (e.g., a single image) or dynamic images (e.g., video clips). Individual video frames of a video clip can be considered as single images.
[0054] The training architecture 300 utilizes a generative training method. During training, for each pair of sample images 301 and sample text sequences 303, at least a text generation task is performed. In some embodiments, for each pair of sample images 301 and sample text sequences 303, an image generation task is also performed.
[0055] Specifically, an image feature representation 312 of the sample image 301 is extracted using an image encoder 310, and a text feature representation 332 of the sample text sequence 303 is extracted using a text encoder. The image feature representation 312 and the text feature representation 332 are provided to a text decoder 340. The feature representation can typically be in the form of a multi-dimensional vector. In this paper, "feature representation" or simply "feature" is also referred to as encoded representation, vector representation, etc. Using the text decoder 340, a predicted text sequence 342 is generated based on the text feature representation 332 and the image feature representation 312. Thus, a text error can be constructed between the generated predicted text sequence 342 and the sample text sequence 303, and the image encoder 310 can be trained based on this text error. For example, a loss function based on the text error can be constructed, and the image encoder 310 can be trained based on the loss value of this loss function.
[0056] Training the image encoder 310 may include updating the parameter values of the image encoder 310 in a direction that continuously reduces the text error (e.g., the corresponding loss function) to a desired or minimum value. Since the image feature representation extracted by the image encoder 310 is used to guide the text generation of the text encoder 340, the loss in the text generation task can, in turn, guide the training of the image encoder 310. During training, multiple pairs of sample images and sample text sequences can be iteratively input into the image encoder 310 and the text encoder 330 to iteratively update the image encoder 310 until the text error between the predicted text sequence and the sample text sequence is reduced to a desired or minimum value. The training of the image encoder 310 can be considered complete at least when the text error is reduced to a desired or minimum value.
[0057] In some embodiments, in addition to text generation tasks, training of the image encoder 310 can also be performed based on image generation tasks. Specifically, using the image decoder 320, a predicted image 322 is generated based on the image feature representation 312 extracted by the image encoder 310. The image encoder 310 is trained based on the image error between the predicted image 322 and the sample image 301. For example, a loss function based on the image error can be constructed, and the image encoder 310 can be trained based on the loss value of this loss function. In some embodiments, a total loss function can be constructed based on the image error and the text error for training the image encoder 310. The training of the image encoder 310 can be considered complete when the sum of the text error and the image error is reduced to the expected value or the minimum value.
[0058] In some embodiments, while training the image encoder 310, the text encoder 330 can also be jointly trained based on text errors (and image errors), and the image decoder 320 and text decoder 340 may also need to be trained. In this way, the parameter values of the text encoder 330, image decoder 320 and text decoder 340 are also updated together to reduce the text errors (and image errors) to the desired or minimum values.
[0059] Figure 4 A schematic diagram of a simplified training architecture for image encoding learning according to some embodiments of the present disclosure is shown. As shown, the image feature representation extracted by the image encoder 310 is used by the text decoder 340 to perform a text generation task, and by the image decoder 320 to perform an image generation task. The training of the image encoder 310 aims to learn useful feature information so as to reduce or minimize the text error between the predicted text sequence generated in the text generation task and the sample text sequence, and to reduce or minimize the image error between the predicted image generated in the image generation task and the sample image.
[0060] In some embodiments, the image encoder 310 and image decoder 320 can be configured as machine learning models or neural networks suitable for processing visual data. The text encoder 330 and text decoder 340 can be configured as machine learning models or neural networks suitable for processing text data. In some embodiments, the image encoder 310, image decoder 320, text encoder 330, and / or text decoder 340 can each be implemented based on one or more Transformer blocks or various variations of Transformer blocks. Some embodiments described below will be illustrated using Transformer block models as examples. In addition to Transformer blocks, one or more of the image encoder 310, image decoder 320, text encoder 330, and / or text decoder 340 can be based on other types of models or neural networks, such as the BERT architecture, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc. The specific type of model structure can be selected according to the actual application requirements.
[0061] In some embodiments, various image generation tasks can be built upon the sample image 301. As an example, the sample image 301 can be modified and then input into the image encoder 310. In this way, the image encoder 310 and image decoder 320 learn to reconstruct the original, undisturbed sample image 301 through feature extraction and decoding processes. That is, the predicted image 322 is required to be as close as possible to the sample image 301. Such an image encoder can be called a denoising autoencoder (DAE). In some embodiments, the modification methods for the sample image 301 may include masking, removing color channels, etc.
[0062] Figure 5 A schematic diagram of the image encoding / decoding process during training according to some embodiments of the present disclosure is shown. In this example, it is assumed that image self-supervised learning is supported based on a masking approach. For ease of discussion, in Figure 5 Example sample image 502 is used as an example for illustration, but this image does not imply any limitation on the embodiments of this disclosure.
[0063] The sample image 502 can be divided into multiple image patches, and at least one image patch must be masked. For example, the sample image 502 can be masked according to a certain masking probability (denoted as mask_ratio∈[0,1]). Assume that the sample image 502 is divided into N image patches, and the number of remaining image patches after masking is M = (1-mask_ratio)×N. Assume that each image patch is mapped to a k-dimensional embedding vector representation. The M unmasked image patches 512 (or k-dimensional embedding vector representations) are input to the image encoder 310 for feature extraction, resulting in feature representations 522 for the M image patches. Assume that the dimension of the feature representation corresponding to each image patch is d.
[0064] Based on the M d-dimensional feature representations, mask feature representations are filled according to the positions of the unmasked image patches in the sample image 502 to obtain feature representations 532 for each of the N image patches. In some embodiments, the filled mask feature representations are also learnable d-dimensional vectors and are shared between different masked image patches (i.e., the same mask feature representation is filled for each masked image patch).
[0065] The feature representations 532 of the padded N image patches are provided to the image decoder 320, which decodes N k-dimensional embedding vector representations 542. These N k-dimensional embedding vector representations 542 can be mapped to a predicted image 552. The image error can be calculated based on the pixel-wise difference (e.g., minimum mean square error) between the predicted image 552 and the unmasked sample image 502. The parameter values of the image encoder 310 (and the image decoder 320, text encoder 330, and text decoder 340) can be updated on the loss function corresponding to this image error using various model training methods, such as gradient descent.
[0066] The aforementioned image decoder is also known as a masked autoencoder (MAE). It should be understood that, except as otherwise provided in the reference... Figure 5 In addition to the embodiments discussed, other image codec structures and other image generation tasks can be used to perform at least the training task of the image encoder, based on the sample images.
[0067] The following section will continue to provide specific examples of the use of image feature representations introduced in text generation tasks. As mentioned earlier, sample text sequence 303 is associated with sample image 301 and is a description of the visual information presented by sample image 301. For example, for Figure 5 The example sample image can be associated with the English text sequence "a parrotcombing feathers".
[0068] In specific text encoding and decoding operations, multiple text units (e.g., single characters or words) in the sample text sequence 303 can be tokenized to convert them into embedded vector representations. For example, a vocabulary can be defined, which includes V text units. Thus, each word in the vocabulary can be converted into a V-dimensional one-hot vector (i.e., only one position in the V dimensions is 1, and the rest are 0). However, a language model (e.g., a fully connected (FC) network) can be learned to map the V-dimensional one-hot vector to a smaller D-dimensional vector (V >> D). In this way, each text unit can be uniquely mapped to a D-dimensional vector. By performing this mapping on each text unit in the sample text sequence, the sample text sequence 303 can be mapped into a feature sequence.
[0069] The feature sequence corresponding to sample text sequence 303 is input into text encoder 330 (denoted as f). text Feature extraction is performed to obtain the text feature representation 332[t1,t2,...,t L ], where L is the number of text units in the sample text sequence 303.
[0070] In some embodiments, the text generation task can be defined as predicting the nth text unit using the feature information of the first n-1 text units in the input text sequence; that is, predicting the subsequent possible text units based only on the previous text units. To achieve this prediction objective, for any given text unit among the L text units, the text encoder 330 can focus on the feature information of the given text unit and the text units preceding it during feature extraction, without focusing on the information of the subsequent text units.
[0071] Such attention constraints can be achieved by adding self-attention constraints to the text encoder 330. For example, for a Transformer block-based text encoder 330, which relies on an attention mechanism to perform feature extraction, the feature extraction of each text unit can be constrained to focus only on the text unit itself and the preceding text units by processing the self-attention weight matrix into a lower triangular self-attention weight matrix during the calculation of the self-attention matrix of each Transformer block.
[0072] Figure 6 A schematic diagram of attention matrix processing for text encoding during training is shown according to some embodiments of this disclosure. Figure 6 In this context, assuming X is the input to the self-attention module in the text encoder 330, and W' represents the self-attention weight matrix, which can be expressed as W' = X TX is used for calculation. A mask is applied to the self-attention weight matrix W' to set the elements in the upper triangular region of matrix W' to negative infinity (i.e., for the element W'ij in the i-th row and j-th column, if j>i, its value becomes negative infinity). The masked self-attention weight matrix W' is then processed by the softmax function to obtain the lower triangular self-attention matrix W, where the self-attention weights in the lower triangular region are retained, while the self-attention weights in the upper triangular region become 0. The lower triangular self-attention matrix W is applied to the input X to obtain the output Y of the current module, where Y... T =WX T .
[0073] Figure 6 The processing of a single self-attention module in the text encoder 330 is illustrated. The text encoder 330 may include multiple self-attention modules, and may also include other types of modules, which will not be described in detail here.
[0074] By constraining the self-attention weight matrix, the feature representation 332[t1,t2,...,t] extracted by the text encoder 330 is made more efficient. L In the diagram, the feature representation t corresponding to each text unit is... i It may only represent the feature information of this text unit and the preceding text units.
[0075] The text decoder 340 is configured to predict the next text unit at each position in the sample text sequence based on the feature representation 332 provided by the text decoder 330 and the image feature representation 312 of the image encoder 310. In some embodiments, self-attention weights for the sample image 301 can be determined based on the image feature representation 312 and the text feature representation 332, and the predicted text sequence 342 can be generated based on the image feature representation and the self-attention weights. For example, the self-attention weights may include weights for the feature representations corresponding to each image patch in the image feature representation. The self-attention weights can be applied to the feature representations of the corresponding image patches. Such image feature representations constrain the text generation task to rely more on the output of the image encoder 310, indirectly increasing the information content of the image feature representations.
[0076] In some embodiments, if the text decoder 340 includes one or more Transformer blocks, then the text feature representation is defined as the query features input to each Transformer block, and the image feature representation is defined as the key and value features input to each Transformer block. The processing of the Transformer blocks can be represented as follows:
[0077]
[0078] Where Q represents the query feature, K represents the key feature, V represents the value feature, and d k The column number of Q and K represents the feature dimension. The above processing can be understood as calculating a self-attention weight matrix using the query feature Q and the key feature K, and then using this self-attention weight matrix to perform a weighted summation on the value feature V. In the general Transformer block processing, Q, K, and V are different projections of the same feature. In some embodiments of this disclosure, after introducing image feature representation, Q can be defined as the output of the text encoder 330, i.e., text feature representation 332, and K and V can be defined as the output of the image encoder 310, i.e., image feature representation 312.
[0079] Based on the above processing, the specific decoding process in text decoder 340 will be discussed. Text decoder 340 (denoted as g) text The input to ) is the text feature representation 332[t1,t2,...,t] output by the text encoder 330. L ] and the image feature representation 312[i1,i2,...,i generated by the image encoder 310. M ](For example, Figure 5 The example is a representation of the image features extracted for M masked image patches (522).
[0080] The output of text decoder 340 is predicted text column 342, which can be represented as a feature sequence [t'1,t'2,...,t'] of the same length as the sample text sequence 303. L ], After post-processing, such as connecting a language model (e.g., a fully connected (FC) network) after the text decoder 340, the D-dimensional feature representation can be remapped back to a V-dimensional one-hot vector [t”1,t”2,...,t”]. L ], Furthermore, the mapped V-dimensional one-hot vector is normalized to a sum of 1 using the softmax function (probabilistic normalization). For example, after processing the V-dimensional one-hot vector t”1 of the first text unit with the softmax function, the predicted probability distribution [p] can be obtained. 11 ,p 12 ,...,p 1L ], where p 11 +p 12 +...+p 1L =1. The predicted probability distribution for the i-th text unit [p i1 ,p i2 ,...,p iLThe probability at each position indicates the probability that the text unit at the corresponding position in the sample text sequence 301 will be predicted. The text unit corresponding to the position with the highest probability is the predicted text unit.
[0081] As mentioned earlier, the text generation task involves using the feature information of the first n-1 text units in an input text sequence to predict the nth text unit. Based on this task objective, label information can be constructed, for example, for predicting text units w1, w2, ..., w in the text sequence. L-1 w L Each position requires prediction of the text unit at the next position. Thus, the label information for the entire sample text sequence is w2, w3, ..., w L w EOS , where w EOS Indicates the end of the sequence.
[0082] Thus, in the text decoding process, at any given text unit position, the text decoder 340 attempts to determine a predicted text unit using the text feature representation for that text unit and the image feature representation 312 provided by the image encoder. This predicted text unit is a prediction of the text units following that text unit in the sample text sequence. For example, the text decoder 340 generates a predicted text unit for text unit w2 based on the text feature representation and image feature representation 312 extracted for the first text unit w1. Similarly, for the last text unit in the sample text sequence, the text decoder 340 generates a prediction for text unit w2 based on the text feature representation and image feature representation 312 extracted for the last text unit, i.e., predicting whether that text unit marks the end of the sample text sequence.
[0083] When calculating the text error, the predicted probability distribution of the i-th text unit output by the text decoder 340 is [p i1 ,p i2 ,...,p iL The corresponding tag word at that word position (e.g., the i-th text unit w) i The corresponding tag word is w i+1 The cross-entropy of the V-dimensional one-hot vector (with a value of 1 at the (i+1)th position and 0 for the rest) is calculated and used as the loss for the text generation task. Based on the loss function corresponding to this text error, various model training methods, such as gradient descent, can be used to update the parameter values of the image encoder 310 (and image decoder 320, text encoder 330, and text decoder 340).
[0084] exist Figure 3In the training architecture, the image feature representation extracted by the image encoder 310 needs to be able to complete not only the image generation task, but also the text generation task. This is equivalent to imposing a greater constraint on the image encoder 310's image feature learning, so that the image encoder 310 needs to learn information that can better represent image features in order to complete both text generation and image generation tasks at the same time.
[0085] In contrast, because image feature representations provide additional feature information to help text decoder 320 complete the text generation task, text encoder 330 may not learn the same amount of information as image encoder 310. For downstream task performance considerations, a trained image encoder 310 can be provided for downstream tasks (e.g., various downstream vision tasks). Text encoder 330 can be discarded. In some embodiments, text decoder 340 can be discarded. Depending on the actual task requirements, image decoder 320 can be retained or discarded.
[0086] According to embodiments of this disclosure, the training of the image encoder is performed by a generative task, rather than a discriminative pre-training / training task. In some embodiments, text generation and image generation tasks are performed simultaneously. Compared to unimodal generation methods that rely solely on image data, the training method proposed in the embodiments of this disclosure utilizes multimodal data (image data and text data) simultaneously. This training approach is more conducive to learning sufficient information from multimodal data. The trained image encoder can be better transferred to various downstream tasks.
[0087] In some embodiments, if Figure 3 The training architecture is a pre-trained architecture, and the trained image encoder 310 can be provided to the model fine-tuning system for fine-tuning according to the needs of downstream tasks. The image encoder 310 can be connected to the image decoder required by the downstream task. In some embodiments, during the fine-tuning stage, for example if the downstream task is also an image generation task, a pre-trained image encoder 310 can also be used. Figure 3 A similar training architecture is used to further fine-tune the image encoder 310. The fine-tuned image encoder 310, along with the image decoder, is provided together to complete the actual task.
[0088] In some embodiments, the image encoder 310 can also be directly applied to actual downstream tasks. In the application of downstream tasks (e.g., by the model application system 130), the image encoder 310 will be used to extract image feature representations of the target image.
[0089] The extracted image feature representations are used to perform a predetermined visual task on the target image. For example, if the visual task is image classification, a corresponding image decoder can be constructed to classify the target image into one of several predetermined categories based on the extracted image feature representations. As another example, if the visual task is semantic segmentation, a corresponding image decoder can be constructed to determine a semantic segmentation map for the target image based on the extracted image feature representations, indicating which semantic category each pixel in the target image is assigned to. In some examples, the predetermined visual task may also include an image generation task, such as extracting image feature representations from a corrupted input image by the image encoder 310, for use by the corresponding image decoder to reconstruct the original image based on the extracted image feature representations. Because the image encoder 310 learns sufficient information during training to accurately extract feature information from various images, it can help improve the quality of task completion in downstream tasks.
[0090] Figure 7 A flowchart of a process for learning image encoding according to some embodiments of the present disclosure is shown. Process 700 may be implemented at model pre-training system 110 and / or model fine-tuning system 120. For ease of discussion, reference will be made to... Figure 1 The environment 100 is used to describe the process 700.
[0091] In box 710, the model pre-training system 110 and / or the model fine-tuning system 120 extract image feature representations of sample images using an image encoder to be trained. In box 720, the model pre-training system 110 and / or the model fine-tuning system 120 extract text feature representations of sample text sequences associated with sample images using a text encoder.
[0092] In box 730, the model pre-training system 110 and / or the model fine-tuning system 120 utilize a text decoder to generate a predicted text sequence based on text feature representations and image feature representations. In box 740, the model pre-training system 110 and / or the model fine-tuning system 120 train the image encoder based at least on the text error between the predicted text sequence and the sample text sequence.
[0093] In some embodiments, training an image encoder includes: generating a predicted image based on an image feature representation using an image decoder; and further training the image encoder based on the image error between the predicted image and a sample image.
[0094] In some embodiments, training the image encoder includes jointly training the image encoder and the text encoder, at least based on text errors. In some embodiments, process 700 further includes providing the trained image encoder for a downstream task, wherein the text encoder is discarded.
[0095] In some embodiments, extracting image feature representations includes: masking at least one image patch in a sample image; and extracting image feature representations from at least one unmasked image patch in an image using an image encoder.
[0096] In some embodiments, the sample text sequence includes a plurality of text units, and wherein extracting text feature representations includes: for a given text unit among the plurality of text units, extracting text feature representations for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.
[0097] In some embodiments, generating a predicted text sequence includes: for a given text unit among a plurality of text units, determining a predicted text unit from text feature representations and image feature representations for the given text unit, wherein the predicted text unit is a prediction of a text unit in the sample text sequence that follows the given text unit.
[0098] In some embodiments, if a given text unit is the last text unit in a sample text sequence, the predicted text unit is a prediction of the end of the sample text sequence.
[0099] In some embodiments, the sample text sequence comprises a plurality of text units, and wherein generating the predicted text sequence comprises: determining self-attention weights for the sample image based on image feature representations and text feature representations; and generating the predicted text sequence based on image feature representations and self-attention weights.
[0100] In some embodiments, the text decoder includes a converter block, and text feature representations are defined as query features input to the converter block, while image feature representations are defined as key and value features input to the converter block.
[0101] Figure 8 A flowchart of a process for an image encoding application according to some embodiments of the present disclosure is shown. Process 800 can be implemented at model application system 130. For ease of discussion, reference will be made to... Figure 1 The environment 100 is used to describe the process 800.
[0102] In block 810, model application system 130 acquires an image encoder trained according to any embodiment of process 700. In block 820, model application system 130 uses the acquired image encoder to extract an image feature representation of the target image. In block 830, model application system 130 performs a predetermined visual task for the target image based on the image feature representation.
[0103] Figure 9A block diagram of an apparatus 900 for image encoding learning according to some embodiments of the present disclosure is shown. The apparatus 900 may be implemented, for example, in or included in a model pre-training system 110 and / or a model fine-tuning system 120. Various modules / components in the apparatus 900 may be implemented by hardware, software, firmware, or any combination thereof.
[0104] As shown in the figure, the device 900 includes an image feature extraction module 910, configured to extract image feature representations of sample images using an image encoder to be trained. The device 900 also includes a text feature extraction module 920, configured to extract text feature representations of sample text sequences associated with sample images using a text encoder. The device 900 further includes a text generation module 930, configured to generate predicted text sequences based on the text feature representations and image feature representations using a text decoder. The device 900 also includes a training module 940, configured to train the image encoder based at least on the text error between the predicted text sequence and the sample text sequence.
[0105] In some embodiments, the training module 940 includes: an image generation module configured to generate a predicted image based on image feature representation using an image decoder; and an image error-based training module configured to further train an image encoder based on the image error between the predicted image and the sample image.
[0106] In some embodiments, the training module 940 includes a joint training module configured to jointly train an image encoder and a text encoder based at least on text errors. In some embodiments, the apparatus 900 further includes an encoder providing module configured to provide a trained image encoder for a downstream task, wherein the text encoder is discarded.
[0107] In some embodiments, the image feature extraction module 910 includes: an image masking module configured to mask at least one image block in a sample image; and a post-masking extraction module configured to extract image feature representations from at least one unmasked image block in the image using an image encoder.
[0108] In some embodiments, the sample text sequence includes a plurality of text units, and the text feature extraction module 920 is configured to: for a given text unit among the plurality of text units, extract text feature representations for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.
[0109] In some embodiments, the text generation module 930 is configured to: for a given text unit among a plurality of text units, determine a predicted text unit from the text feature representation and image feature representation for the given text unit, wherein the predicted text unit is a prediction of the text unit following the given text unit in the sample text sequence.
[0110] In some embodiments, if a given text unit is the last text unit in a sample text sequence, the predicted text unit is a prediction of the end of the sample text sequence.
[0111] In some embodiments, the sample text sequence includes a plurality of text units, and the text generation module 930 includes: a weight determination module configured to determine self-attention weights for a sample image based on image feature representations and text feature representations; and a weight-based text generation module configured to generate a predicted text sequence based on image feature representations and self-attention weights.
[0112] In some embodiments, the text decoder includes a converter block, and text feature representations are defined as query features input to the converter block, while image feature representations are defined as key and value features input to the converter block.
[0113] Figure 10 A block diagram of an apparatus 1000 for an image encoding application according to some embodiments of the present disclosure is shown. The apparatus 1000 may be implemented in or included in, for example, a model application system 130. Various modules / components in the apparatus 1000 may be implemented by hardware, software, firmware, or any combination thereof.
[0114] As shown in the figure, the apparatus 1000 includes an acquisition module 1010 configured to acquire an image encoder trained according to any embodiment of the apparatus 900. The apparatus 1000 also includes a feature extraction module 1020 configured to extract image feature representations of a target image using the acquired image encoder. The apparatus 1000 further includes a task execution module 1030 configured to perform a predetermined visual task on the target image based on the image feature representations.
[0115] Figure 11 A block diagram of an electronic device 1100 that may implement one or more embodiments of the present disclosure is shown. It should be understood that... Figure 11 The electronic device 1100 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 1100 may, for example, be used to implement a model pre-training system 110, a model fine-tuning system 120, and / or a model application system 130.
[0116] like Figure 11As shown, electronic device 1100 is in the form of a general-purpose computing device. Components of electronic device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1100.
[0117] Electronic device 1100 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1130 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 1100.
[0118] Electronic device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 11 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1120 may include computer program product 1125 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0119] Communication unit 1140 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 1100 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 1100 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0120] Input device 1150 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1160 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1100 can also communicate with one or more external devices (not shown) via communication unit 1140 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1100, or with any device that enables electronic device 1100 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0121] According to exemplary embodiments of the present disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary embodiments of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0122] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0123] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0124] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0126] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various embodiments disclosed herein.
Claims
1. A method for image coding learning, comprising: The image feature representation of the sample image is extracted using the image encoder to be trained; A text encoder is used to extract text feature representations of a sample text sequence, which is associated with the sample image and is used to describe the visual information presented by the sample image. Using a text decoder, a predicted text sequence is generated based on the text feature representation and the image feature representation; and The image encoder is trained based at least on the text error between the predicted text sequence and the sample text sequence, wherein training of the image encoder is performed by a generative task, and training the image encoder includes: A predicted image is generated by performing the generative task based on the image feature representation using an image decoder. as well as The image encoder is also trained based on the image error between the predicted image and the sample image.
2. The method of claim 1, wherein training the image encoder comprises: The image encoder and the text encoder are jointly trained based at least on the text error; The method further includes: The trained image encoder is provided for downstream tasks, wherein the text encoder is discarded.
3. The method according to claim 1, wherein extracting the image feature representation comprises: Mask at least one image patch in the sample image; as well as The image feature representation is extracted from at least one unmasked image block in the image using the image encoder.
4. The method of claim 1, wherein the sample text sequence comprises a plurality of text units, and wherein extracting the text feature representation comprises: For a given text unit among the plurality of text units Extract text feature representations for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.
5. The method of claim 4, wherein generating the predicted text sequence comprises: For the given text unit among the plurality of text units A predicted text unit is determined from the text feature representation and the image feature representation for the given text unit, wherein the predicted text unit is a prediction of the text unit following the given text unit in the sample text sequence.
6. The method of claim 5, wherein if the given text unit is the last text unit in the sample text sequence, the predicted text unit is a prediction of the end of the sample text sequence.
7. The method of claim 1, wherein the sample text sequence comprises a plurality of text units, and wherein generating the predicted text sequence comprises: The self-attention weights for the sample image are determined based on the image feature representation and the text feature representation; as well as The predicted text sequence is generated based on the image feature representation and the self-attention weights.
8. The method of claim 1, wherein the text decoder includes a converter block, and the text feature representation is defined as query features input to the converter block, and the image feature representation is defined as key features and value features input to the converter block.
9. A method for image coding applications, comprising: Obtain an image encoder trained according to any one of claims 1 to 8; The image feature representation of the target image is extracted using the acquired image encoder; as well as A predetermined visual task for the target image is performed based on the image feature representation.
10. An apparatus for image coding learning, comprising: The image feature extraction module is configured to extract image feature representations of sample images using the image encoder to be trained. The text feature extraction module is configured to extract text feature representations of sample text sequences using a text encoder, the sample text sequences being associated with the sample image, and the sample text sequences being used to describe the visual information presented by the sample image; The text generation module is configured to generate a predicted text sequence based on the text feature representation and the image feature representation using a text decoder; as well as A training module is configured to train the image encoder based at least on the text error between the predicted text sequence and the sample text sequence, wherein training of the image encoder is performed by a generative task, and the training module includes: An image generation module is configured to generate a predicted image by performing the generative task based on the image feature representation using an image decoder. as well as The image error-based training module is configured to also train the image encoder based on the image error between the predicted image and the sample image.
11. The apparatus of claim 10, wherein the training module comprises: A joint training module is configured to jointly train the image encoder and the text encoder based at least on the text error; The device further includes: An encoder providing module is configured to provide a trained image encoder for downstream tasks, wherein the text encoder is discarded.
12. The apparatus according to claim 10, wherein the image feature extraction module comprises: An image masking module is configured to mask at least one image patch in the sample image; as well as The post-masking extraction module is configured to extract the image feature representation from at least one unmasked image block in the image using the image encoder.
13. The apparatus of claim 10, wherein the sample text sequence comprises a plurality of text units, and wherein the text feature extraction module is configured to: for a given text unit among the plurality of text units, Extract text feature representations for the given text unit from the given text unit and at least one text unit preceding the given text unit in the sample text sequence.
14. The apparatus of claim 10, wherein the sample text sequence comprises a plurality of text units, and wherein the text generation module comprises: The weight determination module is configured to determine the self-attention weights for the sample image based on the image feature representation and the text feature representation; as well as The weighted text generation module is configured to generate the predicted text sequence based on the image feature representation and the self-attention weights.
15. The apparatus of claim 10, wherein the text decoder includes a converter block, and the text feature representation is defined as query features input to the converter block, and the image feature representation is defined as key features and value features input to the converter block.
16. An apparatus for image coding applications, comprising: The acquisition module is configured to acquire an image encoder trained by the method according to any one of claims 1 to 8; The feature extraction module is configured to extract image feature representations of the target image using the acquired image encoder; as well as The task execution module is configured to perform a predetermined visual task on the target image based on the image feature representation.
17. An electronic device comprising: At least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 8 and / or the method according to claim 9 when executed by the at least one processing unit.
18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 8 and / or the method according to claim 9.
Citation Information
Patent Citations
Visual language task processing system, training method and device, equipment and medium
CN113792112A
Training method and device of image-text matching model, and method and device for realizing image-text retrieval
CN113836333A
Image interpolation method and device, processing equipment and storage medium
CN114494011A
End-to-end text image translation model training method
CN114626392A