Method and apparatus for visual language understanding

By introducing cross-entropy text to text loss in the visual language model, learning soft prompts brings them close to hand-designed text prompts, solving the problem of base category overfitting and significantly improving the recognition accuracy of novel categories.

CN119948496APending Publication Date: 2025-05-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067554.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-12
Filing Date
2023-09-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has the problem of base category overfitting when training visual language models, resulting in a decrease in recognition accuracy on novel categories.

Method used

By introducing cross-entropy text to text loss, learning soft cues brings them close to hand-designed text cues in the embedding space, thus alleviating base category overfitting.

Benefits of technology

Significantly improves recognition accuracy on novel categories, outperforms existing soft-cue learning methods, and enables training of models without visual samples available.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948496A_ABST
    Figure CN119948496A_ABST
Patent Text Reader

Abstract

The present disclosure provides a computer-implemented method for training a visual language machine learning (ML) model to classify an image depicting a novel or known category. The method includes obtaining a first training dataset including a plurality of category names and a training visual language ML model. The training method comprises the following steps: generating at least one enhanced text prompt; inputting at least one enhanced text hint into a frozen text encoder; outputting the first text embedding of each enhanced text prompt; generating a plurality of first inputs by cascading each of the plurality of learnable soft cues; inputting the category name and the plurality of first inputs into a frozen text encoder; outputting the second text embedding of each first input; and minimizing cross-entropy text-to-text loss between the first text embedding and the second text embedding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to methods and apparatus for visual language understanding. Specifically, the present application provides a method for training a visual language model to classify images using novel categories, and a method for using the trained visual language model to classify images for a series of downstream tasks. Background Art

[0002] Large-scale pre-training of machine learning models has recently led to the construction of a range of foundational models for language and visual-language understanding. Unlike previous generations of neural networks, such models can better capture the distribution of the world from which new properties and behavioral features emerge. One such property is their ability to perform zero-shot understanding guided by textual prompts. Initially, recent methods, first for language and more recently for visual-language, have gradually moved from hand-crafting prompts to automatically learning prompts using a few labeled samples.

[0003] Accordingly, Applicants have recognized a need for an improved method of training visual-language models. Summary of the invention

[0004] Technical Solution

[0005] In a first approach of the present technology, a computer-implemented method for training a visual-language machine learning (ML) model to classify images depicting novel or known categories is provided, the method comprising: obtaining a first training dataset comprising a plurality of category names; and training the visual-language ML model by: generating at least one enhanced text prompt for each category name in the first training dataset to condition the visual-language ML model to output the category name of an object detected in the image; inputting at least one enhanced soft prompt into a frozen text encoder of the visual-language ML model; and outputting a first enhanced text prompt from the frozen text encoder for each enhanced text prompt. A text embedding, a first text embedding representing a category name in an enhanced text prompt; generating multiple first inputs by concatenating each of a plurality of learnable soft prompts to each category name in a first training dataset; inputting the category names in the first training dataset and the generated multiple first inputs into a frozen text encoder of a visual language ML model; outputting a second text embedding for each first input from the frozen text encoder of the visual language ML model, the second text embedding representing the category name in each first input; and minimizing a cross-entropy text-to-text loss between the first text embedding and the second text embedding, thereby training the learnable soft prompt to be similar to the enhanced text prompt. The method can be performed by a server.

[0006] In a second method of the present technology, a device is provided for classifying images depicting novel or known categories using a trained visual language machine learning (ML) model, the device comprising: an interface for receiving at least one input text data item containing at least one category name; a storage storing a list of category names and corresponding text embeddings; and at least one processor coupled to the storage, arranged to use the trained visual language ML model to: generate a text embedding representing each category name in the input text data item using a text encoder of the visual language ML model, compare the generated text embedding with the text embeddings of the category names in the stored list to determine whether the generated text embedding is novel, and when the generated text embedding of the category name is novel, add the category name and the corresponding text embedding to the list of category names in the storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments of the present technology will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0008] FIG1A is a schematic diagram of a visual language model of the present technology;

[0009] FIG. 1B is a block diagram illustrating a text encoder of the ML model of FIG. 1A ;

[0010] FIG1C is a schematic diagram showing how to use category names to generate enhanced text prompts;

[0011] FIG1D shows a schematic diagram of how to use enhanced text prompts to train an ML model;

[0012] FIG. 1E is a block diagram illustrating an image encoder of an ML model.

[0013] FIG. 1F is a schematic diagram showing how to use the outputs of an image encoder and a text encoder to obtain a cross entropy loss;

[0014] FIG2A is a flow chart of example steps for training the current vision-language ML model;

[0015] FIG2B illustrates a process of generating at least one enhanced text prompt for each category name in the first training data set;

[0016] FIG2C shows a process of training an ML model using a first training dataset;

[0017] FIG2D shows the process of training the ML model using the second training dataset;

[0018] Figure 3 is a schematic block diagram of how a trained ML model can be used to classify images using novel categories;

[0019] 4A to 4D show experimental data of the performance of the present visual-language ML model compared to existing models;

[0020] Figure 5 Experimental results comparing the present technology with the current state of the art are shown;

[0021] Figure 6 Experimental results showing the impact of different components of LASP;

[0022] FIG7A shows experimental results on the selection of enhanced templates in the DTD dataset;

[0023] FIG7B shows experimental results on the selection of enhanced templates in the EuroSAT dataset;

[0024] FIG7C shows the experimental results on the selection of enhanced templates in the UCF101 dataset;

[0025] Figure 8 Experimental results on varying loss selection are shown;

[0026] FIG9A shows experimental results on the influence of out-of-domain distractors in the EuroSAT dataset;

[0027] FIG9B shows the experimental results on the influence of out-of-domain distractors in the Food101 dataset;

[0028] FIG9C shows experimental results on the influence of out-of-domain distractors in the Flowers102 dataset;

[0029] FIG9D shows experimental results on the influence of out-of-domain distractors in the OxfordPets dataset;

[0030] FIG10A shows experimental results on the influence of intra-domain distractors in the Food101 dataset;

[0031] FIG10B shows experimental results on the influence of intra-domain distractors in the Flowers102 dataset;

[0032] Fig.11 is a flowchart of example steps for training a visual-language ML model using this technique;

[0033] Fig.12 is a flowchart of example steps for using a trained vision-language ML model to perform a vision-language task; and

[0034] Fig.13 is a block diagram of a system 300 for training and using ML models to perform vision-language tasks. DETAILED DESCRIPTION

[0035] Broadly speaking, embodiments of the present technology provide a method and apparatus for visual language understanding. In particular, the present application provides a method for training a visual language model to classify images using novel categories, and a method for using the trained visual language model to classify images for a range of downstream tasks.

[0036] Large-scale pre-training of neural networks has recently led to the construction of a large number of basic models for language and vision and language (V&L) understanding. Unlike previous generation neural networks, such models can better capture the distribution of the world from which new advantageous properties and features emerge. Of particular interest in this work are V&L models trained using contrastive learning (i.e., CLIP-like models), which have achieved seamless few-shot and even zero-shot adaptation to new downstream tasks and datasets. The applicant proposes a simple but efficient way to significantly improve soft-hint learning for few-shot adaptation of V&L models to given downstream tasks.

[0037] Similar to their NLP counterparts, prompt engineering and learning have become one of the most powerful techniques for adapting V&L to new tasks. Initially, in early models, a manually-defined set of hand-engineered templates (or prompts), such as a photo of {cls_name} or a black-and-white photo of {cls_name}, was passed through the text encoder of the V&L model to create category-specific weights for the category cls_name that can be used for zero-shot recognition. Following research in NLP, subsequent work has proposed replacing the manually picked templates with a sequence of learnable vectors, also creating soft prompts, which are fed as input to the text encoder along with the category name cls_name. The soft prompts are learned from a few training examples, where the entire V&L model remains frozen. The entire process can be viewed as effectively fine-tuning the parameters of the model on a small training dataset.

[0038] However, a clearly identifiable problem with hint learning is overfitting to the base class: while the accuracy for the classes used for training (base classes) increases significantly, the accuracy for classes not seen during training (novel) decreases significantly. This is somewhat expected, since the soft hints are learned from the few examples belonging to the base classes. Notably, direct zero-shot recognition using hand-crafted hints outperforms all existing soft hint learning methods on novel classes.

[0039] To mitigate overfitting of base categories, the applicants propose a solution motivated by the following observation: since hint learning improves accuracy on base categories, but hint design is significantly better on novel categories, it is proposed to learn soft hints by adding a cross-entropy text-to-text loss, which forces the learned hints to be close to the text hints in the embedding space, thereby leveraging the inherent information captured by the text encoder. The proposed text-to-text loss makes language-only optimization of V&L model adaptation possible for the first time. This is in contrast to existing soft hint learning methods that only capture V&L interactions.

[0040] This technology provides at least the following contributions:

[0041] (1) A novel framework for soft prompting learning, called Language-Aware Soft Prompting (LASP) in this paper;

[0042] (2) Language-only optimization for V&L model adaptation. Specifically, a novel text-to-text cross-entropy loss that maximizes the probability that the learned prompt is correctly classified relative to the hand-crafted prompt. The effectiveness of this new loss in alleviating overfitting of the base classes is demonstrated below;

[0043] (3) a grouped speech-aware cue representation, where each group of cues is dedicated to a different subset of the predefined manual templates;

[0044] (4) a recalibration mechanism based on (a) layer normalization fine-tuning and (b) learning class-agnostic bias to account for it; and

[0045] (5) A novel training method for LASP, which involves training with category names for which no visual samples are available.

[0046] Our techniques provide new state-of-the-art for few-shot and zero-shot image classification on 11 datasets, significantly outperforming all soft-hint prior work. Importantly, we present a hint-learning approach that outperforms very strong baselines based on hand-crafted hints and CLIP (i.e., the zero-shot setting) for identifying new categories for most of the test datasets (8 out of 11).

[0047] Before describing the present technology, some related prior works are described.

[0048] Contrastive Visual-Language Models: Recently, large-scale visual-language pre-training with contrastive learning has shown that it is possible to learn models with robust representations that are transferable to new tasks in zero-shot and few-shot settings. Such networks typically consist of an image encoder (e.g., a visual transformer) and a transformer-based text encoder. In an embodiment, the image encoder can be a visual encoder. A highly parameterized instantiation of this architecture is then trained on a large corpus of over 400 million visual-language pairs using contrastive learning-based methods. In contrast, the present technique takes an orthogonal direction to this, attempting to improve the downstream performance of this approach on few-shot and zero-shot scenarios by learning a set of cues that supplement the input of a frozen contrastive language-image pre-training (CLIP) model.

[0049] Hint learning involves adapting a pre-trained base model on a downstream task, typically in a zero-shot or few-shot setting. Initially studied in the context of large language models (LMs), hints in their simplest form initially involved prepending handcrafted instructions / examples to the task input, such that the LM generates appropriate outputs conditioned on the input or augments the text embedding representation of the CLIP to make it close to the image representation on the chosen downstream dataset. In the context of LMs, techniques aim to reformulate the downstream task as a cloze task using handcrafted patterns (or templates), thereby avoiding the need to train task-specific classifiers. Since finding the optimal pattern is laborious, recent work has attempted to address this problem by learning a collection of soft (continuous) hints.

[0050] In vision-language based models like CLIP, category names are used to create handcrafted hints that are fed as input to a text encoder, enabling zero-shot visual recognition. CoOp extends the work on soft hint optimization to the vision-language domain by learning a set of M hints that are used as input to a text encoder along with category names. Hints are learned by minimizing the classification error with respect to a training set consisting of a given base category. A major limitation of CoOp is weak generalization: the learned hints overfit the base categories and do not work well when tested on novel categories. The CoOP technique minimizes the distance between visual and language embeddings using only hints (p) and category names (CLS). To mitigate this, CoCoOp proposes a dynamic version of CoOp, where a small network is trained to produce visual features from the input image, which are added to the learned hints, thus making them input-specific (i.e., dynamic).

[0051] In contrast to the prior art, and as described in more detail below, the loss of the present technique is a pure text-to-text loss, further allowing for the incorporation of virtual categories. The present technique outperforms CLIP on novel categories.

[0052] This technique proposes to exploit the joint visual-language space learned by CLIP by optimizing cues in both the visual-language domain and the language-language domain. In addition to minimizing the distance between visual and language embeddings, this technique (referred to herein as LASP) exploits the intrinsic model of the world learned by the text encoder by minimizing the loss between soft learnable cues (p) and the category name (CLS) and the sentence (G) into which the CLS token is inserted. This formulation can incorporate novel category names for which no visual information is available (i.e., categories for which there are no corresponding images).

[0053] The advantages are manifold: (a) Similar to image augmentation, which allows the model to generalize to changes in scale, pose, rotation, etc., text can be augmented by adding background or distractors. The model is then forced to correctly classify the text cue given a learned text cue. (b) It acts as a regularizer. (c) It allows for mixed few-shot and zero-shot training, where the model can be trained to be aware of categories for which no visual data exists. Due to these properties, the present technique can outperform methods that condition on both the cue and the image features.

[0054] The present technique mitigates base class overfitting and significantly improves upon previously reported state-of-the-art results, without resorting to dynamic approaches as in CoCoOp. In its basic version, LASP deploys a text-to-text loss that enforces the learned cues to be “close” to a manually defined set of text cues in the text encoder space. Importantly, basic LASP can be extended in three important ways: (1) by allowing the incorporation of virtual categories, i.e., novel category name information for which no (visual) training data is available (LASP-V). This is shown to significantly improve the robustness of the learned cues at no additional cost during inference; (2) by allowing the use of grouped cue representations within the proposed language-aware training, which is shown to increase the representation power of the learned cues; and (3) by performing a further optimization of the visual encoder such that the visual and text embeddings are realigned, leading to significant accuracy gains. Notably, in contrast to CoCoOp, which requires recomputing all category-related text embeddings every time a new image is to be classified, the present approach is very efficient (as efficient as CoOp).

[0055] Methods: Background

[0056] The hint design enables zero-shot visual recognition using a visual language model trained using contrastive learning (CLIP in this work), as follows: Given a set V of C class names, class_name c , c∈{1,...,C}, hints, i.e., manually designed templates concatenated with category names, such as h c = {class_namec} photos, passed through V&L's text encoder g T (.) to calculate the category-specific text features (weights) t h c =g T (h c ). In addition, the image x to be classified is passed through the V&L image encoder g I (.) to calculate the image-specific feature f = g I (x). The probability distribution over the class labels is given by:

[0057]

Mathematical formula 1

[0058]

[0059] where τ is the temperature factor and cos is the cosine similarity. Finally, the category of x is given by Note that in order to calculate t h c , does not require training with category-specific image data, thus enabling zero-shot recognition for any given category name.

[0060] Soft hint learning involves learning M learnable vectors by using a small number of labeled samples Specifically, the manually selected prompt h c is replaced by m The sequence and word embedding w of class_namec c The new learnable hint r formed by cascading c , that is: r c ={p 1 ,p 2 ,...,p M ,w c}, and, finally, obtain the category-specific text features t r c =g T (r c ). The probability distribution over the class labels is:

[0061]

Mathematical formula 2

[0062]

[0063] The hint can be learned by minimizing the cross entropy loss:

[0064]

Mathematical formula 3

[0065]

[0066] Note that the visual language model is kept completely frozen during training. Furthermore, since the soft cues are generally shared across all categories, they can be directly used for zero-shot evaluation on additional novel categories.

[0067] Method: Language Aware Soft Prompt (LASP)

[0068] Despite their strong performance on base (i.e., seen) categories, current vanilla soft-cue based optimization methods perform poorly on novel categories (i.e., zero-shot setting). While CoCoOP partially mitigates this by learning a small neural network conditioned on image features, the zero-shot accuracy still lags behind that of CLIP with handcrafted cues. Furthermore, it needs to pass cues for all categories through the text encoder whenever a new image is to be classified.

[0069] FIG1A is a schematic diagram of the visual language model of the present technology. The present technology uses the proposed text-to-text loss L TT (Equation (5)) jointly models image-text and text-text interactions. That is, departing from the current paradigm of only considering visual-linguistic interactions between cues and images, the present technique exploits the intrinsic model of the world learned by the text encoder. This (1) exploits the relational representations encoded in the language model, (2) enables augmentation of text data, and (3) creates a hybrid few-shot + zero-shot training process, which enables the incorporation of categories that do not have labeled samples.

[0070] There exists a hint p learned by a group G of text encoders j i , to form G text embeddings t that summarize the input j Then L TT Losses are applied to different sets of text embeddings and text hints. Furthermore, to mitigate data distribution shift and visual-language misalignment, the LN layer of the visual encoder is fine-tuned and the embeddings are “corrected” at the output space by a learnable vector b shared by all categories. The text encoder is kept completely frozen. It is worth noting that LASP can exploit virtual category training by including category names for which no visual samples are available during training.

[0071] Our technique contributes novel language-only optimizations for vision-language downstream adaptation. This is in contrast to existing soft-cue learning methods that only capture vision-language interactions. Specifically, since hand-crafted textual cues outperform learnable soft cues for zero-shot settings, to avoid overfitting of base categories and enforce generalizability, our learnable cues are trained such that they can be correctly classified in language space, where the category weights are given by the textual cues. In other words, the model is forced to correctly classify the learnable cues as the corresponding hand-crafted cues.

[0072] To this end, a second cross-entropy loss is used to minimize the distance between the encoded learned soft hints and the encoded textual soft hints. Specifically, recall that t h c =g T (h c) is the class weight of class c obtained by encoding the corresponding textual prompt. Assuming that there are L manually defined textual prompts available, we have t h,l c ,l=1,...,L. In addition, t r is a learnable cue for the encoding to be classified into one of the C categories. Finally, the cue t r The probability of being classified as category y is:

[0073]

Mathematical formula 4

[0074]

[0075] The language-aware training loss is calculated similarly to the V&L loss:

[0076]

Mathematical formula 5

[0077]

[0078] The overall training objective is defined as:

[0079]

Mathematical formula 6

[0080]

[0081] where α VL and α TT Is to control L VL and L TT A user-defined scaling factor for the magnitude of the loss. As already mentioned, the proposed learning formula is called Language Aware Soft Prompts (LASP).

[0082] FIG. 1A is now described in more detail. The visual-language ML model includes a text encoder 10 and an image encoder 12. The text encoder 10 is used to process text inputs 16, 18 and output text embeddings representing the text inputs, while the image encoder 12 is used to process image inputs 14 and output image embeddings representing the image inputs. The two embeddings represent the text input and the image input in a joint embedding / representation space of text and image. The text encoder and the image encoder are pre-trained to align similar text and image concepts in the joint embedding / representation space. Then, during training of the current visual-language ML model, the text encoder 10 is frozen, while training of the ML model occurs to train learnable soft hints. The image encoder 12 is further trained during training of the current visual-language ML model to generate image embeddings for each image input. The various components or blocks of the visual-language ML model are now described in more detail with respect to FIGS. 1B to 1F.

[0083] 1B is a block diagram showing a text encoder 10 of an ML model. The input of the text encoder 10 is text. As described above, the text encoder 10 outputs a d-dimensional feature vector (f), which is a text embedding corresponding to a category name.

[0084] The text encoder 10 of the ML model is itself pre-trained to learn how to output text embeddings. In the present technique, the text encoder 10 is not further trained. Instead, the training of the visual language ML model includes learning a set of vectors, referred to herein as "learnable soft hints", which are used to provide appropriate inputs to the text encoder 10 so that the text encoder 10 can generate text embeddings. Therefore, the text encoder 10 is frozen (remains fixed) during the training of the ML model, i.e., is not retrained or further trained.

[0085] The learnable soft hints start with some initial values ​​that may have been set during pre-training of the text encoder 10. The learnable soft hints are trained / updated in two ways. In the first way, the learnable soft hints are updated based on training the ML model using a first training data set 16 including a plurality of category names 16A. That is, the training of the ML model involves using a training data set 16 that contains only text category names 16A but not corresponding images. In this case, training the visual language ML model includes generating at least one enhanced text hint 20 for each category name 16A in the first training data set to adjust the visual language ML model to output the category name of the object detected in the image. At least one enhanced text hint 20 is input into the frozen text encoder 10 of the ML model, and the frozen text encoder 10 outputs a first text embedding for each enhanced text hint 20, the first text embedding representing the category name in the enhanced text hint 20. The method includes generating a plurality of first inputs by cascading each of the plurality of learnable soft hints 18 to each category name 16A in the first training data set 16. The category names 16A in the first training dataset 16 and the generated plurality of first inputs are input into the frozen text encoder 10 of the ML model, and the frozen text encoder 10 outputs a second text embedding for each category name 16A. Then, the method involves minimizing the cross entropy text-to-text loss (L ) between the first text embedding and the second text embedding. TT ), thereby training the learnable soft prompt to be similar to the enhanced text prompt. Minimizing the cross entropy text-to-text loss between the first text embedding and the second text embedding may include adjusting the learnable soft prompt. Thus, the output of this training using the first training data set results in an updated soft prompt. The learnable soft prompt 18 may be updated using a gradient descent that uses a text-to-text loss (LTT) between the first text embedding and the second text embedding.

[0086] While fine-tuning models pre-trained on large-scale datasets using self-supervised contrastive learning results in strongly supervised classifiers, the ability to classify unseen categories is severely compromised during the fine-tuning process. Therefore, to mitigate this, the text encoder 10 is kept frozen and only a small set of learnable parameters ("soft hints") 18 are trained. As described above, the learnable parameters 18 are passed as input into the text encoder 10 along with the category names 16A.

[0087] 1C is a schematic diagram showing how the category names 16A are used to generate enhanced text prompts. The method includes generating at least one enhanced text prompt 20 for each category name 16A in a first training data set. That is, a list / data set 16 of category names 16A for which corresponding visual examples may or may not exist is used to generate the enhanced text prompt 20. Generating the at least one enhanced text prompt may include: selecting at least one manually defined enhancement template from a plurality of enhancement templates, each enhancement template being a text phrase into which the category name may be inserted, and inserting the category name 16A in the first training data set 16 into the selected at least one enhancement template, thereby generating the at least one enhanced text prompt 20. This is useful in the same way that it is useful to have a set of image transformation techniques - the enhancement template is a defined way of enhancing the training data to generate new training inputs.

[0088] Selecting at least one enhancement template may include selecting at least one group of enhancement templates. That is, a plurality of enhancement templates may be divided into groups, and each group may be associated with a certain sequence or set of prompt types that produce category-specific text features. For example, enhancement templates may be divided into groups depending on the aspects they cover (e.g., pose, photo quality, viewing angle, etc.). In this case, the output is a set of enhancement category names, one set for each group of templates.

[0089] Current soft-hint learning methods restrict augmentation to the visual domain, where random transformations such as rotation, color jittering, or scaling increase the robustness of the system, especially for cases with a limited number of training samples. However, in the prior art, augmentation is not performed in the language domain. Ideally, it is desired that prompt-conditioned text embeddings are also robust, capturing the full space of each category. The present technique achieves this through targeted prompts, where certain features can be specified and / or text-based transformations can be applied to category names, e.g.: "sketch of a dog" or "rotated photo of a dog".

[0090] FIG1D shows a schematic diagram of how to train an ML model using an enhanced text prompt 20. At least one enhanced text prompt 20 is input into the frozen text encoder 10 of the ML model, and the frozen text encoder 10 outputs a first text embedding for each enhanced text prompt 20. Therefore, the input to the text encoder 10 is the enhanced category name / enhanced text prompt 20 and the category name descriptor adjusted using a shared learnable bias "b". The output of the text encoder 10 is a list of d-dimensional feature vectors encoded by the text encoder, one vector for each category name. As described above, the method then involves minimizing the cross-entropy text-to-text loss LTT, which enforces the enhanced text prompt 20 to be correctly classified given the category name text descriptor t. Note that unlike previous work, this is the first time that text-to-text optimization has been used for image visual recognition. As described above, although vanilla soft prompt learning has strong performance on the base category, vanilla soft prompt learning performs poorly on novel categories (i.e., zero-shot settings). To alleviate this situation, the present technology proposes language-only optimization for visual language downstream adaptation for the first time. This is in contrast to existing soft-cue learning methods that only capture visual-language interactions. Specifically, since hand-crafted textual cues outperform learnable soft cues for the zero-shot setting, to avoid overfitting of the base categories and to enforce generalizability, we propose that learnable soft cues18 should be trained such that they can be correctly classified in the language space, where the category weights are given by the textual cues. In other words, the model is forced to correctly classify the learnable cues as the corresponding hand-crafted cues.

[0091] As shown in FIG. 1A , the ML model includes an image encoder 12. The training method includes training the image encoder 12 while also training / updating the learnable soft hint 18. This is a second way of training / updating the learnable soft hint 18. To this end, the training of the ML model also includes obtaining a second training data set including a plurality of data pairs, wherein each data pair includes an image depicting an object and a class name of the object. Therefore, the second training data set is different from the first training data set because the second training data set includes text and images. The class name in each data pair corresponds to an object in the image of the data pair. Therefore, the images in the second training data set are substantially labeled images.

[0092] 1E is a block diagram of an image encoder 12 of an ML model. The input to the image encoder 12 is an image 14. During training, the image 14 is obtained from a labeled dataset (i.e., a second training dataset containing pairs of data items, each pair including an image and text). During inference, the image 14 can be obtained from any source and input into the ML model for any visual language task. The image 14 can be an RGB color image. The image encoder 12 processes the input image 14, identifies each object in the image, and for each object, outputs a d-dimensional feature vector (f). The feature vector is then used to classify the identified object and to generate an image embedding corresponding to the category assigned to the identified object.

[0093] In order to implement the training of both the image encoder and the text encoder, the ML model is trained using data pairs, wherein each data pair includes an image 14 of an object and a corresponding text category name. However, the input of the text encoder is not just the text category name in each data pair. Instead, the input of the text encoder is the text category name and multiple text inputs based on the text category name. The text category name and the learnable soft prompt 18 are used to generate the text input. Therefore, in the case where the second training data set is available, training the visual language ML model can also include generating multiple second inputs by cascading each of the multiple learnable soft prompts 18 to each category name in the data pair in the second training data set. The category name in the second training data set and the generated multiple second inputs are then input into the frozen text encoder 10 of the ML model. The frozen text encoder 10 outputs a third text embedding for each second input, and the third text embedding represents the category name in each second input. The image encoder 12 of the training ML model includes inputting the image 14 in each data pair of the second training data set into the image encoder 12, and outputting the image embedding of the object in each input image in the data pair from the image encoder 12. The training method then involves minimizing the cross entropy image-to-text loss (L VL ).

[0094] By minimizing the cross entropy image-to-text loss (L VL ) to train the visual language ML model may include adjusting the learnable soft hints 18 used to generate the second input into the frozen text encoder so that for each data pair, the third text embedding is similar to the image embedding. Thus, the learnable soft hints are adjusted via two training processes.

[0095] Training the visual language ML model by minimizing the cross entropy image-to-text loss between the third text embedding and the image embedding may include adjusting the learnable parameters of the layer norm layer of the image encoder 12 of the ML model. That is, since updating all parameters of the model results in overfitting and loss of zero-shot capability (i.e., the ability to recognize novel categories not seen during pre-training), the present technique adapts the image encoder 12 to the target subdomain by adjusting the statistics learned by the layer norm present in the encoder. Specifically, the layer normalization of the image encoder is fine-tuned, thereby training the image encoder to output an image embedding similar to the third text embedding for each data pair. The encoder itself is typically represented by a visual transformer architecture and is pre-trained in an unsupervised manner via contrastive learning.

[0096] Training the visual language ML model may also include learning how to reduce the impact of data distribution shift. Data distribution shift may occur when the image data used to pre-train the image encoder is different from the images in the data pair of the second data set that can be used for a specific downstream task. Therefore, layer normalization fine-tuning is performed to combat data distribution shift, as described above. However, after performing fine-tuning, the visual encoder and the text encoder may not be aligned. Therefore, the method also includes learning an offset or deviation at the output of the text encoder, which can be used to realign the two encoders. Training the ML model may also include reducing the impact of data distribution shift by: learning an offset at the output of the text encoder to realign the visual encoder and the text encoder; and adding the offset to the weight of the text encoder. In other words, the ML model is trained to correct the deviation. The fine-tuning and deviation correction steps can preferably be performed simultaneously.

[0097] FIG1F is a schematic diagram showing how the output (f) of the image encoder 12 and the text encoder 12 is used to obtain the cross entropy loss. Here, f is the image descriptor generated by the image encoder 12 given the input image 14. In FIG1D , t is the text descriptor generated by the text encoder 12 by taking as input the (soft prompt, category name) pair. There are a total of c (the same as the number of categories) text prompts, one for each category. The output of training using the second training data set is the cross entropy loss LVL, which forces the descriptor f to be correctly classified. The cross entropy loss forces the image feature f to be correctly classified as category i, in effect maximizing the similarity between the image descriptor and the corresponding text descriptor corresponding to the i-th category name.

[0098] Figure 2A is a flow chart of example steps for training the present visual language ML model. The process is explained in more detail with reference to Figures 2B to 2D.

[0099] FIG. 2B illustrates a process of generating at least one enhanced text prompt 20 for each category name 16A in the first training data set 16. As described above, text enhancement is similar to image enhancement, where during training, an image is transformed using a set of random transformations. Here, the category name is enhanced. The enhanced text prompt is a sentence or phrase containing the category name of interest. For example, if the first training data set 16 includes the category name "bamboo forest", the enhanced text prompt generated using the category name may include "rotated image of bamboo forest", "enlarged image of bamboo forest", "sketch of bamboo forest", etc. In this way, the category name is enhanced by forming a sentence or short descriptor using the category name. Since the text encoder 10 is part of a visual language model, the enhanced text prompt can be image-based because the model attempts to classify images. These enhanced text prompts are then used together with the original category names of the first training data set to train the text encoder.

[0100] Figure 2C shows the process of training an ML model using the first training dataset 16. The learnable parameters are first optimized only for the image recognition problem in the text domain (the "learnable hint" in the figure). This allows the incorporation of novel categories for which no visual samples exist, thereby significantly increasing zero-shot performance.

[0101] As shown, training the visual language ML model includes generating at least one enhanced text prompt 20 for each category name 16A in the first training data set. At least one enhanced text prompt 20 is input into the frozen text encoder 10 of the ML model, and the frozen text encoder 10 outputs a first text embedding for each enhanced text prompt 20. The method includes generating multiple first inputs by cascading a learnable soft prompt in a plurality of learnable soft prompts 18 to each category name 16A in the first training data set 16. The category names 16A in the first training data set 16 and the generated multiple first inputs are input into the frozen text encoder 10 of the ML model, and the frozen text encoder 10 outputs a second text embedding for each category name 16A. Then, the method involves minimizing the cross entropy text-to-text loss between the first text embedding and the second text embedding. Minimizing the cross entropy text-to-text loss between the first text embedding and the second text embedding may include adjusting the learnable soft prompt. Therefore, the output of this training using the first training data set results in an updated soft prompt. The learnable soft prompt 18 may be updated using a gradient descent that uses the text-to-text loss between the first text embedding and the second text embedding.

[0102] 2D illustrates the process of training an ML model using a second training dataset 22. This involves realigning the image encoder 12 to the domain of interest by updating only the layer norm weights of the image encoder 12. This prevents overfitting, adapting the image encoder 12 to the new data distribution while maintaining the zero-shot capability of the model. The input to the image encoder 12 is an image 14. During training, the image 14 is obtained from a labeled dataset (i.e., a second training dataset containing pairs of data items, each pair including an image and text). The image encoder 12 processes the input image 14, identifies each object in the image, and outputs a d-dimensional feature vector (f) for each object. The feature vector is then used to classify the identified object.

[0103] In order to implement the training of both the image encoder and the text encoder, the ML model is trained using data pairs, wherein each data pair includes an image 14 of an object and a corresponding text category name. However, the input of the text encoder is not just the text category name in each data pair. Instead, the input of the text encoder is the text category name and multiple text inputs based on the text category name. The text category name and the learnable soft prompt 18 are used to generate the text input. Therefore, in the case where the second training data set is available, training the visual language ML model can also include generating multiple second inputs by cascading each of the multiple learnable soft prompts 18 to each category name in the data pair in the second training data set. The category name in the second training data set and the generated multiple second inputs are then input into the frozen text encoder 10 of the ML model. The frozen text encoder 10 outputs a third text embedding for each category name. The image encoder 12 of the training ML model includes inputting the image 14 in each data pair of the second training data set into the image encoder 12, and outputting the image embedding of the object in each input image in the data pair. Then, the training method involves minimizing the cross entropy text-to-text loss between the third text embedding and the image embedding.

[0104] As shown in FIG2D , the input of the image encoder 12 may be an enhanced version of the image 14 in the second training data set 22. One or more possible enhancement techniques are used to enhance the image 14 in the data pair. The enhancement may result in a synthetically generated version of the original image in the data pair. The image encoder then processes the enhanced image to output an image embedding for each object in the enhanced image. The text category name in the corresponding data pair provides a ground truth. The image in the data pair may be enhanced by applying at least one image transformation (such as rotation, tilt, flipping around an axis, magnification, etc.). The image in the data pair may be enhanced by applying one or more of the following to change the appearance of the image: changes in color distribution, noise, Gaussian noise, blur, motion blur, zoom blur, simulated weather effects, simulated lighting changes, etc. The image in the data pair may be enhanced by inserting an object into the image, which may not overlap or partially overlap with the object in the original image.

[0105] Figure 3 is a schematic block diagram of how a trained ML model may be used to classify images using images depicting novel or known categories. One input to the trained ML model is at least one input text data item containing at least one category name. The trained visual language ML model generates a text embedding representing each category name in the input text data item using a text encoder 10 of the ML model. The trained ML model has access to an existing list of known category names. The trained ML model compares the generated text embedding with text embeddings of known category names to determine whether the generated text embedding / each generated text embedding is novel. When the text embedding generated for a category name is novel, the trained ML model adds the category name and the corresponding generated text embedding to the list. For example, in Figure 3 , the model has access to a list of category names, and the user input text data items include novel categories "French butter" and "oat milk".

[0106] The model can receive at least one image, which can be a still image or a frame of a video. Figure 3, an image of the inside of a refrigerator. The ML model identifies at least one object in each input image using the image encoder 12 of the ML model, and generates an image embedding for each identified object using the image encoder. The ML model compares the image embedding of each identified object with the text embeddings corresponding to each category in the list of category names to determine similarity, and using the comparison, outputs a category name for each identified object in the input image based on the most similar embedding. In other words, the ML model outputs a category name for each identified object in the input image, where the category names may be category names that the model was trained to recognize, or novel category names that the model has received after training. In this example, the ML model outputs, for example, "oat milk" and "red peppers," which have been recognized and classified within the input image.

[0107] Visual language models can be used for a variety of purposes. For example, one use case involves a robot or other electronic device with a camera (e.g., a smart refrigerator with a camera, a robot vacuum cleaner, etc.). An ML model is provided on the device to enable the device to automatically recognize new objects of interest without requiring visual examples of those novel objects provided by the user. Instead, a list containing the names of the categories of interest is simply provided to the device. This allows the device to recognize new objects without having to be updated first. Advantageously, this enables novel objects to be recognized based solely on a list of category names. The names can be user-defined or automatically obtained (e.g., by downloading product names found on a retailer's web page). That is, the user can enter a text data item, or the text data item can be downloaded from an external source. For example, an external source can be an online food recipe that the user is following or from a supermarket or other store. Alternatively, the model also has the ability to adapt on the device based on a few examples (images + their category names).

[0108] like Figure 3 As shown, the visual language model can be used in conjunction with another model. In this example, the visual language model is used in conjunction with an AI assistant, where the category names output from the visual language ML model are input into the AI ​​assistant for processing. Thus, in some cases, the output of the visual language model can be output directly to the user, while in other cases the output can be used to control another entity or device, and in other cases the output can be input into another model (such as an AI assistant model). In this example, the AI ​​assistant can notify the user which products they need to purchase so that the user does not purchase food products they already have, or so that the user can be reminded to purchase what they have run out of. The AI ​​assistant can notify the user whether they have all the necessary ingredients to cook the recipes they plan to cook.

[0109] In some cases, the AI ​​assistant may be an AI cooking assistant. In this case, the ML model may receive multiple input text data items including ingredients in a target recipe, and at least one image of a user following the target recipe. The trained visual language ML model may identify objects corresponding to the ingredients in at least one image. The AI ​​assistant may then use the identified objects to provide instructions to the user. For example, at least one image may be a frame of a live video of a user performing the steps of a recipe to make an apple pie. The user may be performing the steps of a recipe that involve cutting apples for an apple pie. The visual language ML model may receive a list of ingredients as an input text data item. The model may also receive method steps as input text data items (which involve the ingredients and utensils / equipment required). The visual language ML model may use only the input text data items to identify and classify objects in the video frame corresponding to the categories in the input text data items. For example, the ML model may identify apples and knives. The classified images may be provided to the AI ​​cooking assistant so that the AI ​​cooking assistant model may determine which step of the recipe the user is in, and provide the user with a prompt to perform the step. The AI ​​cooking assistant may also be able to provide the user with instructions on the next step to be performed.

[0110] The electronic device may be a smart phone having a storage including a plurality of images. For example, the storage may include a photo gallery. In this case, the ML model may receive an input text data item indicating an object or category to be recognized in a plurality of images, and the trained visual language ML model may analyze the images in the storage and output at least one image of the plurality of images corresponding to the input text data item. Thus, the ML model enables a fast photo gallery search to be performed based solely on the input text data item (i.e., a novel category in text form).

[0111] The visual language model LASP can be interpreted in a number of ways as now described.

[0112] LASP as a language-based augmentation. Current soft-cue based methods restrict augmentation to the visual domain, where random transformations such as rotation, color jittering, or scaling increase the robustness of the system, especially for cases with a limited number of training samples. However, no augmentation is performed in the language domain. Ideally, it is desired that the prompt-conditioned text embeddings are also robust, capturing the full space of each category. In practice, this can be achieved via target prompts, where certain features can be specified, or by applying text-based transformations to the category names, e.g.: “sketch of a dog” or “rotated photo of a dog”. At training time, as reflected in Equation 6, the model needs to correctly distinguish between the prompts t iand C cues formed by combining the C category names with one of the sentences from the set ζ. Additionally, the category label distribution is computed for each l-th template and then averaged over all templates. Note that the sentences / templates used for augmentation are not mixed, as the model needs to focus on the category information, rather than relying on augmentation (i.e.: the model can more easily distinguish between “sketch of a dog” and “photo of a wolf” than “sketch of a dog” and “photo of a wolf” in the first case, styles can be used as an additional queue). This has been experimentally verified and was found to be the case on the test dataset. Mixing templates was found to impact the performance on novel categories by 0.5%.

[0113] LASP as a regularizer. Even when using a small number of parameters, especially on the few-shot setting, the resulting model (cue) can still exhibit symptoms consistent with overfitting to the base class. Since the proposed language-aware loss encourages new learned cues to be close to the textual cues in the embedding space, the present technique can naturally be viewed in part as a regularizer that prevents the cues-conditioned features from deviating too much from the predefined handcrafted linguistic query.

[0114] LASP acts as a feature discriminator learner. By optimizing both image and text features, the technique produces class centroids that are more discriminative and have higher separation margins. This effect can be visualized in Figures 4A to 4D, where cosine distances are measured between embeddings for each class in the test set. In general, the technique learns embeddings / class centroids with higher cosine distances than the baseline CoOp.

[0115] Figures 4A to 4D show experimental data on the performance of the present visual language ML model ("ours", i.e., LASP) compared to an existing model (CoOP). Specifically, Figures 4A to 4D show the cosine distances between all text category embeddings generated by the CLIP text encoder on the Eurosat, DTD, Flowers102, and Caltech101 datasets. Since the underlying image features between the two methods are the same, clusters that are farther apart are more discriminative. Each figure also indicates the average cosine distance between categories. Brighter colors / shading indicate larger cosine distances.

[0116] LASP as data-free distillation. Typically, knowledge distillation requires a training set of images, where a teacher network provides the training signal for the student. The text-to-text loss of the present technique can also be interpreted as data-free distillation (without using any image data), where a learnable cue defines a “sample”. As CLIP learns a joint visual-language space, similar concepts are close between the two domains. Therefore, optimizing for concepts or objects in the language domain using the proposed loss should also help make steps in the visual domain, thereby improving classification of images.

[0117] Grouped LASP. Grouped convolution and multi-head attention have been shown to learn strong representations. The number of groups or heads, respectively, can also be interpreted as a set of experts that are then combined to produce strong features. Drawing inspiration from them, a grouped prompt representation is proposed, where each group is optimized with respect to a separate subset of textual prompts. Effectively, the prompts from each group will learn a transformation that is specific to its corresponding subset (similar to the aforementioned techniques that are also specific to a portion of the signal). In particular, the set of L templates is partitioned into G subsets of equal size. Furthermore, each subset is associated with a sequence r of M prompts. c g ={p 1 g ,p 2 g ,...,p M g ,w c},g=1,2,...,G are associated, each producing a category-specific text feature t r,g c =g T (r c g ). Finally, the text-to-text loss in Equation (5) becomes:

[0118]

Mathematical formula 7

[0119]

[0120] Where P g rh Similar to equation (4) is calculated for each group. The segmentation is generated randomly. Since text templates are usually semantically independent, no preferred grouping will emerge. At test time, the final result is calculated by taking the average of the cosine similarity scores between each group and the visual feature f.

[0121] Realigning LASP: Adversarial Data Distribution Shift. For some downstream tasks, there may be a data distribution shift between the downstream image dataset and the dataset used by CLIP during training. Therefore, it is desirable that this aspect be captured by the downstream adaptation method. To this end, some optimizations of the visual encoder can be performed; however, if after training, the visual-language embeddings are pushed away from the joint space learned by CLIP, this can very easily lead to overfitting of the base classes. For example, preliminary results with visual adapters suggest that they hurt zero-shot accuracy. Instead, it was found that layer normalization (LN) fine-tuning is a much more robust way to adapt the visual encoder. In summary, fine-tuning the LN of the CLIP encoder is proposed as a way to combat distribution shift.

[0122] Combating visual-linguistic misalignment. Since vision and language are not guaranteed to remain aligned after LN fine-tuning, our technique learns a “correction” at the output of the text encoder in the form of a learnable offset (bias) that aims to realign the two modalities. Let W be the set of weights of the linear classifier obtained by passing the learning hints from the text encoder. Our technique learns a vector b∈R that is simply added to W. d , i.e. W = W + b. Importantly, the learned offset is shared across all classes, and in this way it can also be easily applied to the case of novel classes.

[0123] Language-Aware Soft Hints with Virtual Categories (LASP-V). A straightforward observation that can be drawn from Equation 4 is that, in practice, it is not necessary to restrict the set V to only category names for which there are annotated images, since its values ​​are independent of the input image. To this end, the present technique learns hints using both annotated image-text pairs and category names that are outside the base set, for which no images are available. This creates a hybrid few-shot and zero-shot setting that combines the best of two words: guidance in newly annotated samples and the prevalence of zero-shots. It is shown in the Results section that this small addition can significantly improve the robustness of the learned hints. To distinguish LASP from being trained only on few-shot data, a variant of this method that is trained on a mixture of zero-shot and few-shot data is called LASP trained with virtual categories (LASP-V). Note that training with virtual categories does not violate the zero-shot setting. Furthermore, from a practical point of view, if novel category names are unknown during initial training, the model can simply be retrained in a zero-shot manner when they become available.

[0124] Experiments. Following Radford et al. and Zhou et al. (CoCoOP), the accuracy of generalization of the present technique to novel categories (i.e., zero-shot recognition) is evaluated on 11 datasets. Each dataset is split into two equal partitions with disjoint categories, called base and novel categories. The present model (LASP) is trained using text-image pairs in the base category and tested on both the base and novel categories.

[0125] Model. For all experiments, unless otherwise stated, the pre-trained CLIP model was used with the ViT-B / 16 image encoder, M = 4 learnable cues and 16 samples per class. For all experiments, the average of three runs is reported. The number of groups G (when used) was set to 3. In all experiments, the average of 3 runs is reported.

[0126] After CoOP, we conduct experiments on 11 datasets also used in CLIP, namely: ImageNet, Caltech101, Oxford-Pets, Stanford Cars, Flowers102, Food101, FGVC Aircraft, SUN397, DTD, EuroSAT, and UCF-101.

[0127] We trained extensively, using the training procedure described in CoOp and CoCoOp (i.e., the same image augmentation, SGD with an initial learning rate of 0.002, and a cosine annealing scheduler with 1 warm-up epoch). In Equation 8, α VL is set to 1, α T is set to 20. The number of soft hints is 4 and the number of text templates of size L is 34, where the templates are taken from CoOP and CLIP (see the supplementary material for a complete list). All training and testing are done on a single nVIDIA V100 GPU, and the code is implemented in PyTorch.

[0128] result. Figure 5 The present technique is compared to the current state of the art, where results for LASP and LASP trained with virtual classes (LASP-V) are shown. All learning-based methods use 4 soft cues and are trained with 16 samples from each base class. The present technique, LASP, significantly outperforms the direct baseline CoOP and stronger image conditioning methods such as CoCoOP, despite not using any of them. Furthermore, the results of LASP match and exceed the performance of CLIP on novel classes.

[0129] As shown in the results, our technique significantly outperforms its direct baseline, CoOp on novel categories, on all 11 datasets, while largely matching the accuracy on base categories. Furthermore, the proposed LASP+, which combines few-shot (on base categories) and zero-shot learning (for novel categories), surpasses CoOp and a stronger successor, CoCoOp, conditioned on image features. Note that stronger gains are observed on datasets with informative category names (such as EuroSAT or UCF101), and lower gains are observed on datasets containing less informative or more challenging category names (such as FGVCAircraft), where isolated model information is less available. Finally, our technique is the first to consistently outperform the CLIP baseline with hand-crafted hints on novel categories.

[0130] As the results further show, the present technique outperforms all methods by a large margin in terms of the harmonic mean. On average, it outperforms the next best (ProDA) by >2%. The improvements on specific datasets are even greater (e.g., >3% on Flowers102, >11% on EuroSAT, >3% on UCF101).

[0131] On novel categories, the present technique outperforms all methods by a large margin. It is the first reported method to outperform CLIP by 0.68% (but note that CLIP performs very poorly on base categories). It also outperforms ProDA (third best) by >2.5%. Again, the improvements on specific datasets are even larger than ProDA (e.g. >5% on Flowers102, >3% on Food101, >11% on EuroSAT, >6% on UCF101). On new categories, LASP with virtual categories (LASP-V) has a significant impact on specific datasets. These include datasets with informative category names such as EuroSAT and DTD, where the improvements over LASP are ~5.5% and ~4.0%, respectively.

[0132] Generalized zero-shot setting: The current evaluation protocol used in CoCoOp considers base and novel categories independently to compute accuracy. A more realistic evaluation protocol should jointly consider categories across both subsets (i.e., base and novel). Detailed results for this setting are provided in the supplementary material, but generally the same conclusions as above hold.

[0133] Ablation study: The impact of different LASP components. LASP proposes multiple contributions that are evaluated incrementally. The starting point is the text-to-text loss proposed by Eq. (7). To this, we incrementally apply the grouped hint representation (Eq. (9)) and then the realignment module. This yields LASP. Finally, virtual categories are added, yielding LASP-V, and the baseline is CoOp. Figure 6 As shown in the results in , our technique outperforms the baseline by a large margin. Specifically, our technique improves CoOp by an average of 4.5%, demonstrating its effectiveness. Figure 6 The results in further show that all components are needed to obtain high accuracy.

[0134] Effect of Size and Content on Textual Prompts: Here, the effect of the size L and content of the set of textual prompts used by the present technique in Eq. (6) is investigated. For simplicity, results are reported only using the text-to-text loss (Eq. (7)). By including the rest of the prompts defined in CLIP, the number of handcrafted templates increases to 100, while by using photo templates their number is reduced to 1. Random templates are generated by sampling grammatically plausible random sentences containing incoherent words of length between 5 and 20 words. The category names are inserted at the end of these random templates (e.g., see the supplementary material). All variants use the same training scheduler and hyperparameters, except in the case of random templates, where α TT =5.

[0135] As shown in the results in Figures 7A, 7B, and 7C, the exact choice of template may be less important for the few-shot setting. The results further show that for the case of novel classes, both the number and content of templates are important for obtaining high accuracy. Note that the accuracy on the base class remains similar across all settings (not shown in Figures 7A, 7B, and 7C).

[0136] Influence of loss type: Figure 8 In this technique, the choice of loss is varied. The cross entropy (CE) is L 2 and L 1 Again, for simplicity, only the text-to-text loss is used to report the results. Figure 8 The results in further demonstrate that the proposed formulation based on CE loss outperforms other losses in the context of this technique.

[0137] Impact of Out-of-Domain Distractors: Inspired by other work showing that the performance of CLIP degrades as the number of categories used for testing increases, a new evaluation setting is introduced: First, 4 test datasets with clearly disjoint domains are selected: EuroSAT (10 satellite terrain types), Food101 (101 food names), Flowers102 (102 flower names), and OxfordPets (37 dog and cat breed names). At test time, the classifier is defined on the union of the categories across all 4 datasets (250 categories in total). Note that LASP-V is the only method that benefits from the knowledge of this extended vocabulary during training. From Figures 9A, 9B, 9C, and 9D, it can be concluded that the present technique is somewhat robust to out-of-domain distractors. Specifically, the decreased inaccuracy is moderate (typically 1-2%). The exception is EuroSAT, where the number of categories increases 25X. Importantly, LASP-V manages to recover the lost accuracy to a large extent.

[0138] Impact of In-Domain Distractors: Extending the ideas from the previous section, here, the performance of current soft hint methods is tested with in-domain distractors. Unlike the case of out-of-domain distractors, in-domain distractors are chosen such that they are closely related to the current dataset / categories being part of the same supercategory. Experiments are performed on two datasets: Food101 and Flowers102. For Flowers102, we added 65 new category names, while for Food101, we added 53 new categories. Note again that except for LASP-V, categories are only used as distractors at test time to expand the C-way classifier by 65 and 53, respectively. The list of added categories can be found in the supplementary material. From the results shown in Figures 10A and 10B, it can be concluded that in-domain distractors significantly increase the difficulty of the problem. Specifically, the decreased inaccuracy is large (4-7%). LASP-V manages to recover part of the lost accuracy.

[0139] Fig.11is a flowchart of example steps for using the present technique to train a visual-language ML model to classify images depicting novel or known categories. The training method includes: obtaining a first training data set including a plurality of category names (S100); and training a visual language ML model by: generating at least one enhanced text prompt for each category name in the first training data set (S102) to adjust the visual language ML model to output the category name of the object detected in the image; inputting the at least one enhanced text prompt into a frozen text encoder of the ML model (S104); outputting a first text embedding of each enhanced text prompt from the frozen text encoder (S106), the first text embedding representing the category name in the enhanced text prompt; generating a plurality of first inputs by cascading each of a plurality of learnable soft prompts to each category name in the first training data set (S108); inputting the category names in the first training data set and the generated plurality of first inputs into the frozen text encoder of the ML model (S110); outputting a second text embedding of each first input from the frozen text encoder of the ML model (S112), the second text embedding representing the category name in each first input; and minimizing the cross entropy text-to-text loss between the first text embedding and the second text embedding (S114), thereby training the learnable soft prompt to be similar to the enhanced text prompt. The method may be executed by a server.

[0140] Fig.12 is a flowchart of example steps for performing a visual language task using a trained visual language ML model. The method may include: receiving at least one input text data item containing at least one category name (S200). The method may include using the trained visual language ML model to: generate a text embedding representing each category name in the input text data item using a text encoder of the ML model (S202); compare the generated text embedding with the text embedding of the category name in a stored list of category names and corresponding text embeddings to determine whether the generated text embedding is novel (S204); and when the generated text embedding of the category name is novel, add the category name and the corresponding text embedding to the list of category names in storage (S206).

[0141] Fig.13 is a block diagram of a system 300 for training and using ML models to perform vision-language tasks.

[0142] The system 300 includes a server 302 for training a visual language machine learning ML model to classify an image. The server 302 includes at least one processor 304 coupled to a memory 306. The at least one processor 304 may be arranged to perform the method described above with respect to FIGS. 1A to 2D.

[0143] The system 300 includes an apparatus 312 for performing a visual-language task using a trained visual-language machine learning ML model 318, the apparatus 312 including: a user interface 320 for receiving at least one input data item, wherein the input data item is an image or text; and at least one processor 314 coupled to a memory 316, which is arranged to: analyze at least one received input data item using the trained visual-language ML model; and output at least one response to the input data item using the visual-language trained ML model.

[0144] Device 312 may be any of the following: a smartphone, a tablet, a laptop, a computer or computing device, a virtual assistant device, a robot or robotic device, a robotic assistant, an image capture system or device, an Internet of Things device, and a smart consumer device. It should be understood that this is a non-limiting and non-exhaustive list of devices.

[0145] At least one processor 314 may include one or more of the following: a microprocessor, a microcontroller, and an integrated circuit. For example, the memory 316 may include a volatile memory such as a random access memory (RAM) used as a temporary memory and / or a non-volatile memory such as a flash memory, a read-only memory (ROM), or an electrically erasable programmable ROM (EEPROM) for storing data, programs, or instructions.

[0146] Broadly speaking, embodiments of the present technology provide a method and apparatus for visual language understanding. Specifically, the present application provides a method for training a visual language model to classify images using novel categories, and a method for using the trained visual language model to classify images for a range of downstream tasks.

[0147] In a first method of the present technology, a computer-implemented method for training a visual-language machine learning ML model to classify images depicting novel or known categories is provided, the method comprising: obtaining a first training data set including a plurality of category names; and training the visual-language ML model by: generating at least one augmented text prompt for each category name in the first training data set to condition the visual-language ML model to output the category name of an object detected in the image; inputting at least one augmented soft prompt into a frozen text encoder of the visual-language ML model; outputting a first text embedding of each augmented text prompt from the frozen text encoder, the first text embedding representing the category name in the augmented text prompt; generating a plurality of first inputs by concatenating each of a plurality of learnable soft prompts to each category name in the first training data set; inputting the category names in the first training data set and the generated plurality of first inputs into the frozen text encoder of the visual-language ML model; outputting a second text embedding of each first input from the frozen text encoder of the visual-language ML model, the second text embedding representing the category name in each first input; and minimizing a cross-entropy text-to-text loss between the first text embedding and the second text embedding, thereby training the learnable soft prompt to be similar to the augmented text prompt. The method may be performed by a server.

[0148] The term "visual-language machine learning model" used in this article refers to a model that combines visual and language modalities. Specifically, the visual-language model has a visual / image module for processing images and a text module for processing text, as well as a technique for fusing information from the two modules together.

[0149] The term “textual hint” as used herein refers to a manually defined template sentence or phrase into which the category names can be inserted. Such textual hints enable the text encoder of the visual language model to create category-specific weights for the category names within the textual hints.

[0150] The term "learnable soft hint" used herein refers to a learnable vector. The learnable soft hint can be input into the text encoder of the visual language model together with the category name, and the learnable soft hint can be learned while keeping the text encoder frozen.

[0151] The term "frozen text encoder" as used herein refers to a pre-trained text encoder with weights and parameters that are fixed when performing the training process. In other words, the pre-trained text encoder is not updated or modified when performing the subsequent training process. This is so that the pre-training of the text encoder is frozen and will not be forgotten during the subsequent training process of the model.

[0152] Advantageously, the present technology provides a method for training a visual language ML model to be able to recognize novel objects (i.e., objects on which the ML model was not trained) using only a textual list of category names for those novel objects. That is, the trained ML model is able to recognize new objects based on being provided with only text, which may be a description of the object or only the category name of the object. This is advantageous because many ML models that perform object recognition and classification are trained using images or image-text pairs (e.g., labeled images), but labeled images are not always readily available. Here, the ML model is trained so that it can recognize novel objects based on textual input related to the novel objects without the need for accompanying images depicting the novel objects. This allows users to more easily personalize the trained ML model because the user can simply provide the trained ML model with novel category names, and the ML model is able to perform object recognition using those novel category names. Similarly, this also enables users to use the model to recognize objects based on a list of category names obtained from other sources (such as from food recipes, supermarkets, etc.) without having to provide example images of those category names.

[0153] The visual language ML model includes a text encoder and an image encoder. The text encoder is used to process text input and output a text embedding representing the text input, while the image encoder is used to process image input and output an image embedding representing the image input. The two embeddings represent the text input and the image input in a joint embedding / representation space of text and image. The text encoder and the image encoder are pre-trained to align similar text and image concepts in the joint embedding / representation space. Then, during the present training method for training the visual language ML model, the text encoder is frozen while training of the ML model occurs to train learnable soft prompts. As explained in more detail below, the image encoder can be further trained during the present training method for training the visual language ML model.

[0154] Minimizing the cross entropy text-to-text loss between the first text embedding and the second text embedding may include adjusting the learnable soft hint so that for each category name, the second text embedding is similar to the first text embedding. This is expected because the category names used to generate the first text embedding and the second text embedding are the same, so the embeddings should be similar. Learnable soft hints are learnable vectors that can be trained so that the entire text encoder does not need to be retrained. That is, although the fine-tuning model pre-trained on a large-scale dataset using self-supervised contrastive learning leads to a strongly supervised classifier, the fine-tuning process seriously negatively affects the ability to classify unseen categories. To alleviate this situation, the text encoder is kept frozen-it has been pre-trained-and only a small set of learnable parameters (called "soft hints") are trained. These parameters are passed into the text encoder as input along with the text category name in the data pair. (Whenever a text input that does not correspond to any image is received, these parameters are also input into the text encoder).

[0155] As described above, the goal of the present training method is to enable the trained ML model to recognize novel objects. To achieve this, the method includes training the ML model using a first training dataset that does not contain any visual examples of the category names therein.

[0156] Specifically, in a manner similar to image enhancement, the first training dataset is used to perform text enhancement, and the enhanced text is used to train the ML model. Here, the enhanced text is an enhanced text prompt, which is a sentence or phrase containing the category name of interest. For example, if the first training dataset includes the category name "bamboo forest", the enhanced text prompt generated using the category name may include "rotated image of bamboo forest", "enlarged image of bamboo forest", "sketch of bamboo forest", etc. In this way, the category name is enhanced by forming a sentence or short descriptor using the category name. Since the text encoder is part of the visual language model, the enhanced text prompt can be image-based because the model attempts to classify the image. These enhanced text prompts are then used with the original category names of the first training dataset to train the text encoder.

[0157] Generating at least one enhanced text prompt may include: selecting at least one manually defined enhancement template from a plurality of enhancement templates, each enhancement template being a text phrase into which a category name may be inserted; and inserting the category name from the first training data set into the selected at least one enhancement template, thereby generating at least one enhanced soft prompt. This is useful in the same way that having a set of image transformation techniques is useful - the enhancement template is a defined way of enhancing the training data to generate new training inputs.

[0158] Selecting at least one enhanced template may include selecting at least one group of enhanced templates. That is, the plurality of enhanced templates may be divided into groups, and each group may be associated with a certain sequence or set of prompt types that produce category-specific text features.

[0159] As described above, the visual language ML model includes a text encoder and an image encoder. The training method includes training the image encoder and a learnable soft prompt input into the text encoder. Therefore, the method may also include obtaining a second training data set including a plurality of data pairs, each data pair including an image depicting an object and a class name of the object. Therefore, the second training data set is different from the first training data set because the second training data set includes text and images. The class name in each data pair corresponds to the object in the image of the data pair. Therefore, the images in the second training data set are substantially labeled images.

[0160] In order to achieve the training of both the image encoder and the text encoder, the ML model is trained using data pairs, wherein each data pair includes an image of an object and a corresponding text category name. However, the input of the text encoder is not just the text category name in each data pair. Instead, the input of the text encoder is a text category name and a plurality of text inputs based on the text category name. The text input is generated using the text category name and a learnable soft prompt. Therefore, in the case where the second training data set is available, training the visual language ML model may also include: generating multiple second inputs by cascading each of the multiple learnable soft prompts to each category name in the data pair in the second training data set; inputting the category name in the second training data set and the generated multiple second inputs into the frozen text encoder of the visual language ML model; outputting the third text embedding of each second input from the frozen text encoder of the visual language ML model, the third text embedding representing the category name in each second input; inputting the image in each data pair of the second training data set into the image encoder of the visual language ML model; outputting the image embedding of the object in each input image in the data pair from the image encoder; and minimizing the cross entropy text-to-text loss between the third text embedding and the image embedding.

[0161] Training the visual language ML model by minimizing the cross entropy image-to-text loss between the third text embedding and the image embedding can include adjusting learnable parameters of an image encoder of the ML model. Specifically, fine-tuning layer normalization of the image encoder to train the image encoder to output an image embedding similar to the third text embedding for each data pair.

[0162] Training the visual language ML model may also include learning how to reduce the impact of data distribution shift. Data distribution shift may occur when the image data used to pre-train the image encoder is different from the images in the data pair of the second data set that can be used for a specific downstream task. Therefore, layer normalization fine-tuning is performed to combat data distribution shift, as described above. However, after performing fine-tuning, the visual encoder and the text encoder may not be aligned. Therefore, the method also includes learning an offset or deviation at the output of the text encoder, which can be used to realign the two encoders. Training the ML model may also include reducing the impact of data distribution shift by: learning an offset at the output of the text encoder for realigning the visual encoder and the text encoder; and adding the offset to the weight of the text encoder. In other words, the ML model is trained to correct the deviation. The fine-tuning and deviation correction steps can preferably be performed simultaneously.

[0163] Training the vision-language ML model by minimizing the cross entropy image-to-text loss between the third text embedding and the image embedding may include adjusting the learnable soft hints used to generate the second input into the frozen text encoder so that for each data pair, the third text embedding is similar to the image embedding. Thus, the learnable soft hints are adjusted via two training processes.

[0164] The image encoder can be trained using both the images in the data pair of the second training data set and the enhanced versions of those images. To this end, one or more possible enhancement techniques are used to enhance the images in the data pair. Enhancement can result in a synthetically generated version of the original image in the data pair. The image encoder then processes the enhanced image to output an image embedding of each object in the enhanced image. The text category name in the corresponding data pair provides a ground truth. The images in the data pair can be enhanced by applying at least one image transformation (such as rotation, tilt, flip around an axis, magnification, etc.). The images in the data pair can be enhanced by applying one or more of the following to change the appearance of the image: changes in color distribution, noise, Gaussian noise, blur, motion blur, zoom blur, simulated weather effects, simulated lighting changes, etc. The images in the data pair can be enhanced by inserting an object into the image, which object may not overlap or partially overlap with the object in the original image.

[0165] Training the visual-language ML model may include jointly training a visual encoder and a text encoder. Jointly training the image encoder and the text encoder may include alternating between optimizing a loss associated with the image encoder and a loss associated with the text encoder.

[0166] Training of vision-language ML models can include using zero-shot and / or hybrid few-shot training.

[0167] In a related method of the present technology, a server for training a visual language machine learning ML model to classify images using novel categories is provided, the server comprising at least one processor coupled to a memory, for implementing any of the methods described above with respect to the first method. Therefore, the features described above with respect to the first method are equally applicable here, and for the sake of brevity, they are not repeated.

[0168] In a second method of the present technology, a device is provided for classifying images depicting novel or known categories using a trained visual language machine learning (ML) model, the device comprising: an interface for receiving at least one input text data item containing at least one category name; a storage storing a list of category names and corresponding text embeddings; and at least one processor coupled to the storage, arranged to use the trained visual language ML model to: generate a text embedding representing each category name in the input text data item using a text encoder of the visual language ML model, compare the generated text embedding with the text embeddings of the category names in the stored list to determine whether the generated text embedding is novel, and when the generated text embedding of the category name is novel, add the category name and the corresponding text embedding to the list of category names in the storage.

[0169] Thus, once trained, the trained ML model is able to recognize known objects (i.e., previously seen during training) and novel objects (i.e., not seen during training) in an image. In the case where the trained ML model is provided with a text input containing a previously seen category, the trained ML model is able to process the text input in a straightforward manner. In other cases, the trained ML model is provided with a text input containing or representing a novel category. For example, the text input may be a single word or term representing the name of a novel category, or may be a sentence containing a novel category, or may be a list of category names, one or more of which may be novel. This is advantageous because a user of the model may be able to simply enter a novel category as a text input, and there is no need to also provide a corresponding example image depicting the novel object. This is particularly advantageous in certain scenarios where obtaining the corresponding image is very burdensome for the user. For example, if the user is using the trained ML model in conjunction with an AI cooking assistant, the user may want to enter a list of ingredients from a recipe (so that the AI ​​cooking assistant can help the user perform the recipe), where the list of ingredients may contain one or more novel categories. The present technology eliminates the need for the user to enter an image representing each ingredient in the list.

[0170] When the interface receives at least one image, at least one processor may be arranged to: identify at least one object in each input image using an image encoder of a visual language ML model, and generate an image embedding for each identified object using the image encoder; compare the image embedding of each identified object with the text embedding in storage to determine similarity; and using the comparison, output a category name for each identified object in the input image based on the most similar embedding. In other words, the ML model outputs a category name for each identified object in the input image, where the category name may be a category name that the model was trained to recognize, or a novel category name that the model has received after training.

[0171] The interface may receive at least one input text data item from a user or from an external source. That is, a user may input a text data item, or a text data item may be downloaded from an external source. For example, an external source may be an online food recipe that a user is following or from a supermarket or other store.

[0172] The at least one input text data item may be a single novel category name and / or at least one sentence containing the novel category name.

[0173] The apparatus may also include an AI assistant, wherein the category names output from the visual language ML model are input into the AI ​​assistant for processing. Thus, in some cases, the output of the visual language model may be output directly to a user, while in other cases the output may be used to control another entity or device, and in other cases the output may be input into another model (such as an AI assistant model).

[0174] In some cases, the AI ​​assistant may be an AI cooking assistant. In this case, the interface may be arranged to receive multiple input text data items including ingredients in a target recipe, and at least one image of a user following the target recipe. The processor of the device may use the trained visual language ML model to identify objects corresponding to ingredients in at least one image. The AI ​​assistant may then use the identified objects to provide instructions to the user. For example, at least one image may be a frame of a real-time video of a user performing steps of a recipe to make an apple pie. The user may be performing steps of a recipe that involve cutting apples for an apple pie. The visual language ML model may receive a list of ingredients as an input text data item. The model may also receive method steps as input text data items (which involve required ingredients and utensils / equipment). The visual language ML model may use only the input text data items to identify and classify objects in the video frame corresponding to the categories in the input text data items. For example, the ML model may identify apples and knives. The classified images may be provided to the AI ​​cooking assistant so that the AI ​​cooking assistant model may determine which step of the recipe the user is in, and provide the user with a prompt to perform the step. The AI ​​cooking assistant may also be able to provide the user with instructions on the next step to be performed.

[0175] In another example, the AI ​​assistant may be part of or used in conjunction with a smart refrigerator. The AI ​​assistant may be provided with a list of ingredients in a target recipe, and may analyze the list and image of the contents of the smart refrigerator to determine whether the smart refrigerator contains the ingredients in the recipe. The AI ​​assistant may then provide the user with instructions as to whether they need to purchase any ingredients for the recipe (or check for ingredients elsewhere in their kitchen).

[0176] The memory of the device may include a plurality of images. For example, the storage device may include a photo library. In this case, the interface may be arranged to receive an input text data item indicating an object or category to be recognized in a plurality of images, and the processor of the device may use the trained visual language ML model to analyze the images in storage and output at least one image of the plurality of images corresponding to the input text data item. Thus, the device enables fast photo library searches to be performed based solely on the input text data item (i.e., novel categories in text form).

[0177] More generally, the device can be used to perform any visual language task. The trained ML model can be used to perform the task, or can be used with another model (e.g., an AI assistant model) to perform the task.

[0178] For example, the visual language task may be a generation task, and when the input data item is an image, at least one response output by the ML model may be any of the following: an answer to a question based on the input image, and a description of the input image. For example, the description of the input image or the answer to the question may indicate whether a particular category of objects is present in the image.

[0179] The visual-linguistic task may be a generation task, and when the input data item is text, at least one response output by the ML model may be a synthetic image generated using the input text.

[0180] The visual-language task may be a classification task, and when the input data item is an image, at least one response output by the ML model may be a textual description of the image.

[0181] The visual-linguistic task may be a classification task, and when the input data item includes an image and text, at least one response output by the ML model may be an indication of whether the text correctly describes the image.

[0182] The visual-linguistic task may be a retrieval task, and when the input data item is text, at least one response output by the ML model may be at least one image extracted from an image database based on the text.

[0183] The apparatus may be a resource-constrained device, but it has minimal hardware capabilities to use the trained neural network / ML model. The apparatus may be any of the following: a smartphone, a tablet, a laptop, a computer or computing device, a virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, an image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an IoT device, or a smart consumer device (such as a smart refrigerator). It should be understood that this is a non-exhaustive and non-limiting list of example apparatuses.

[0184] In a related method of the present technology, there is a method for using a trained visual language machine learning ML model to classify images using novel categories. The features described above for the second method also apply to this method and are not repeated for the sake of brevity.

[0185] In a related method of the present technology, a computer-readable storage medium is provided that includes instructions that, when executed by a processor, cause the processor to perform any of the methods described herein.

[0186] As will be appreciated by those skilled in the art, the present technology can be embodied as a system, method or computer program product.Therefore, the present technology can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects.

[0187] Furthermore, the present technology may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The computer readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any suitable combination of the foregoing.

[0188] Computer program code for performing operations of the present technology may be written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. Code components may be embodied as procedures, methods, etc., and may include subcomponents that may take the form of instructions or sequences of instructions at any level of abstraction, from direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.

[0189] Embodiments of the present technology also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to perform any of the methods described herein.

[0190] The technology also provides processor control code to implement the above methods, for example, on a general-purpose computer system or a digital signal processor (DSP). The technology also provides a carrier carrying the processor control code, which implements any of the above methods when running, especially on a non-transitory data carrier. The code can be provided on a carrier such as a disk, a microprocessor, a CD-ROM or a DVD-ROM, a programming memory such as a non-volatile memory (e.g., flash memory) or a read-only memory (firmware), or provided on a data carrier such as an optical or electrical signal carrier. The code (and / or data) for implementing the embodiments of the technology described herein may include source, object or executable code in a conventional programming language (interpreted or compiled), such as Python, C or assembly code, code for setting or controlling an ASIC (application-specific integrated circuit) or an FPGA (field programmable gate array), or code for a hardware description language (such as Verilog (RTM) or VHDL (very high speed integrated circuit hardware description language)). As the technician will understand, such code and / or data can be distributed between multiple coupled components that communicate with each other. The technology may include a controller including a microprocessor, a working memory, and a program memory coupled to one or more components of the system.

[0191] It will also be clear to those skilled in the art that all or part of the logic methods according to embodiments of the present technology may be appropriately embodied in a logic device including logic elements to perform the steps of the above methods, and such logic elements may include components such as logic gates in a programmable logic array or an application-specific integrated circuit. Such a logic arrangement may also be embodied in an enabling element for temporarily or permanently establishing a logic structure in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using a fixed or transmittable carrier medium.

[0192] In an embodiment, the present technology can be implemented in the form of a data carrier having functional data thereon, wherein the functional data includes a functional computer data structure which, when loaded into a computer system or network and operated thereby, enables the computer system to perform all the steps of the above-described method.

[0193] The above method can be performed in whole or in part on a device (i.e., an electronic device) using machine learning or an artificial intelligence model. The model can be processed by a dedicated processor for artificial intelligence, which is designed with a hardware structure specified for artificial intelligence model processing. The artificial intelligence model can be obtained by training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform a desired feature (or purpose) by training a basic artificial intelligence model using a plurality of training data through a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and the neural network calculation is performed by calculation between the calculation result of the previous layer and the multiple weight values.

[0194] As described above, the present technology can be implemented using an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory, and processor. The processor may include one or more processors. At this time, one or more processors may be general-purpose processors such as a central processing unit (CPU), an application processor (AP), etc., a graphics processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor such as a neural processing unit (NPU). One or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. Predefined operating rules or artificial intelligence models are provided by training or learning. Here, providing by learning means that an AI model of predefined operating rules or desired characteristics is made by applying a learning algorithm to multiple learning data. Learning can be performed in the device itself in which the AI ​​according to the embodiment is performed, and / or it can be implemented by a separate server / system.

[0195] The AI ​​model can be composed of multiple neural network layers. Each layer has multiple weight values, and the layer operation is performed by the calculation of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0196] A learning algorithm is a method for training a predetermined target device (e.g., a robot) using multiple learning data to enable, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0197] References:

[0198] (1) Contrastive Language-Image Pre-training (CLIP) model - Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748-8763. PMLR, 2021.

[0199] (2) CoOp-Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9): 2337-2348, 2022b

[0200] (3) CoCoOp-Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp.16816-16825, 2022a

[0201] (4) Radford et al-Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019

[0202] (5) ImageNet-Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and LiFei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248-255. IEEE, 2009

[0203] (6) Caltech101-Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pages 178-178. IEEE, 2004

[0204] (7) Oxford-Pets-Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498-3505. IEEE, 2012

[0205] (8) Stanford Cars-Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3D object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554-561, 2013.

[0206] (9) Flowers102-Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722-729. IEEE, 2008.

[0207] (10) Food101-Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101-mining discriminative components with random forests. In European conference on computer vision, pages 446-461. Springer, 2014

[0208] (11) FGVC Aircraft-Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.

[0209] (12) SUN397-Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485-3492. IEEE, 2010.

[0210] (13) DTD-Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606-3613, 2014.

[0211] (14) EuroSAT - Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7): 2217-2226, 2019.

[0212] (15) UCF-101-Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.

[0213] (16) ProDA -Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian. Prompt distribution learning. In IEEE Conference on Computer Vision and Pattern Recognition, 2022.

[0214] Those skilled in the art will appreciate that, although what is considered to be the best mode has been described above, and other modes of carrying out the present technology where appropriate, the present technology should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognize that the present technology has a wide range of applications, and that the embodiments can be widely modified without departing from any inventive concept defined in the appended claims.

Claims

1. A computer-implemented method for training a visual-language machine learning (ML) model to classify images depicting novel or known categories, the method comprising: Obtain a first training data set including a plurality of category names; as well as Train the Vision-Language ML model by: generating at least one augmented text prompt for each category name in the first training dataset to condition the visual language ML model to output the category name of the object detected in the image; inputting the at least one augmented text prompt into a frozen text encoder of a vision-language ML model; Outputting a first text embedding for each augmented text prompt from the frozen text encoder, the first text embedding representing a category name in the augmented text prompt; generating a plurality of first inputs by concatenating each of the plurality of learnable soft prompts to each category name in the first training dataset; Inputting the category names in the first training dataset and the generated plurality of first inputs into a frozen text encoder of the vision-language ML model; outputting a second text embedding for each first input from a frozen text encoder of the vision-language ML model, the second text embedding representing a category name in each first input; and Minimize the cross entropy text-to-text loss between the first text embedding and the second text embedding.

2. The method according to claim 1, wherein: Minimizing a cross entropy text-to-text loss between the first text embedding and the second text embedding includes adjusting the learnable soft hints so that for each category name, the second text embedding is similar to the first text embedding.

3. The method according to claim 1 or 2, wherein: Generating at least one enhanced text prompt includes: selecting at least one manually defined enhancement template from a plurality of enhancement templates, each enhancement template being a text phrase into which a category name may be inserted; and The category names in the first training data set are inserted into the selected at least one enhanced template to generate at least one enhanced text prompt.

4. The method according to claim 3, wherein: Selecting at least one enhancement template includes selecting at least one group of enhancement templates.

5. The method according to any preceding claim, further comprising: obtaining a second training data set comprising a plurality of data pairs, each data pair comprising an image depicting an object and a class name of the object; Among them, training the visual language ML model also includes: generating a plurality of second inputs by concatenating each of the plurality of learnable soft prompts to each category name in a data pair in a second training dataset; Inputting the category names in the second training dataset and the generated plurality of second inputs into a frozen text encoder of the vision-language ML model; outputting a third text embedding for each second input from the frozen text encoder of the vision-language ML model, the third text embedding representing the category name in each second input; Inputting the image in each data pair of the second training dataset into the image encoder of the vision-language ML model; outputting, from the image encoder, an image embedding of the object in each input image; and Minimize the cross entropy image-to-text loss between the third text embedding and the image embedding.

6. The method according to claim 5, wherein: Training the vision-language ML model by minimizing the cross entropy image-to-text loss between the third text embedding and the image embedding includes: fine-tuning layer normalization of an image encoder of the ML model, thereby training the image encoder to output an image embedding similar to the third text embedding for each data pair.

7. The method according to claim 6, wherein: Training vision-language ML models also involves reducing the impact of data distribution shift by: learning an offset at the output of the text encoder for realigning the visual encoder and the text encoder; and Add offset to the weights of the frozen text encoder.

8. The method according to claim 5, 6 or 7, wherein: Training the vision-language ML model by minimizing a cross-entropy image-to-text loss between a third text embedding and an image embedding includes adjusting a learnable soft hint used to generate a second input into a frozen text encoder such that for each data pair, the third text embedding is similar to the image embedding.

9. An apparatus for classifying images depicting novel or known categories using a trained visual-language machine learning (ML) model, the trained visual-language machine learning (ML) model being trained according to the method of any one of claims 1 to 8, the apparatus comprising: An interface for receiving at least one input text data item including at least one category name; Storage, stores a list of category names and the corresponding text embeddings; and at least one processor, coupled to the memory, the at least one processor being arranged to use the trained visual-language ML model to: Use the text encoder of the vision-language ML model to generate text embeddings representing each category name in the input text data item, comparing the generated text embedding to the text embeddings of the category names in the stored list to determine whether the generated text embedding is novel, and When the generated text embedding of a category name is novel, the category name and the corresponding text embedding are added to the list of category names in storage.

10. The device according to claim 9, wherein: The interface receives at least one image, and the at least one processor is arranged to: using an image encoder of the visual-language ML model to identify at least one object in each input image, and generating an image embedding for each identified object using the image encoder; Compare the image embedding of each identified object with the stored text embeddings to determine similarity; as well as Using the comparison, the category name of each recognized object is output based on the most similar embedding.

11. The device according to claim 9 or 10, wherein: The interface receives the at least one input text data item from a user or from an external source.

12. The device according to claim 9, 10 or 11, wherein: The at least one input text data item is a single novel category name and / or at least one sentence containing a novel category name.

13. The device according to any one of claims 10 to 12, further comprising an AI assistant, wherein: The category names output from the visual-language ML model are fed into the AI ​​assistant for processing.

14. The device according to claim 13, wherein: AI Assistant is an AI cooking assistant; The interface is arranged to receive a plurality of input text data items, the plurality of input text data items comprising ingredients in a target recipe and at least one image of a user following the target recipe; The processor uses the trained visual language ML model to identify an object in the at least one image that corresponds to an ingredient; and The AI ​​assistant uses the recognized objects to provide instructions to the user.

15. The device according to any one of claims 9 to 14, wherein: The storage includes a plurality of images; The interface is arranged to receive an input text data item indicating an object or category to be identified in the plurality of images; as well as The processor uses the trained visual-language ML model to output at least one image of the plurality of images corresponding to the input text data item.