Classification using multi-modal large language models

By generating text descriptions and category predictions for images through a multimodal language model and combining visual and text features for image classification, the system solves the accuracy limitation problem caused by the single visual feature in existing technologies and achieves higher classification accuracy and flexibility.

CN120671011APending Publication Date: 2025-09-19GOOGLE LLC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510666180.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2025-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing technologies, neural networks only rely on visual features when classifying images, resulting in limited classification accuracy, especially poor performance in zero-shot classification of unseen categories.

Method used

A multimodal language model is used to process the input to generate text descriptions and category predictions for the image, and visual and text features are integrated through a text encoder neural network and a feature embedding fusion engine to generate query embeddings for classification.

Benefits of technology

Improved image classification accuracy, especially in the zero-shot setting, and enhanced classification flexibility and generalization by combining visual and textual features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671011A_ABST
    Figure CN120671011A_ABST
Patent Text Reader

Abstract

Methods, systems, and devices for classification. In one aspect, a method includes receiving an input and a request to classify the input into one of a plurality of categories, processing the input using a multi-modal model to generate (i) a description of the input and (ii) a category prediction, a description of the input and the category prediction are processed using a text encoder embedding neural network to generate (i) a text description feature embedding and (ii) a predictive feature embedding, a query feature embedding representing the input being generated from at least the description feature embedding and the predictive feature embedding, and classifying the input into one of a plurality of categories using the query embedding.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to using a neural network to process an input to generate an output sequence.

[0002] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict an output for a received input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer serves as the input to the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output from the received input based on the current value input of a corresponding set of parameters. Summary of the Invention

[0003] This specification describes a system implemented as a computer program on one or more computers in one or more locations that classifies input using a multimodal language model in response to receiving a request to classify the input into one (or more) of a plurality of categories.

[0004] To classify the input, the system can process the input using a multimodal language model to generate a description of the input and a category prediction for the input. The system can then use a text encoder neural network to process the description and category prediction of the input to generate corresponding text description feature embeddings and prediction feature embeddings.

[0005] The system can then generate a query embedding representing the input from at least the descriptive feature embedding and the predictive feature embedding, and the system can use the query embedding to classify the input into one of a plurality of categories.

[0006] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0007] The classification task may include classifying the input into one or more classes (e.g., categories) based on extracting features of the input using a pre-trained neural network. In some examples, the system can perform zero-shot classification by classifying inputs such as images into categories that were not explicitly presented during training. For example, the system can provide a request to a large language model (LLM) to generate a description of each category in the category. The system can process the input image using an image embedding neural network to generate an image embedding that represents the visual features of the image. The system matches the input image embedding with the most similar category embedding based on a similarity measure, and the system can then classify the input into the category corresponding to the most similar category embedding.

[0008] However, relying solely on extracting visual features from images may limit classification accuracy because the extracted visual features may not capture other descriptive characteristics of the image. In contrast, the described system utilizes multimodal LLM to perform zero-shot classification by generating a textual representation of the input, which leads to more accurate classification of the input.

[0009] Specifically, the system can use a multimodal language model to process the input to generate both a description of the input and a category prediction for the input. The system can then use a text encoder embedding neural network to process the description and category prediction of the input to generate corresponding embeddings, and the system can generate a query embedding from the corresponding embeddings to classify the input into one of the categories. Therefore, the system can more accurately classify the input because the system can utilize both the description and category prediction to generate the query embedding regardless of whether the initial category prediction is correct.

[0010] In this case, the system provides a first prompt to the multimodal language model, the first prompt requesting the generation of a description of the input, and the system can provide a category for classification and a second prompt requesting the generation of a category prediction for the input to the multimodal language model. The first prompt and the second prompt can be universally applicable to classification tasks, providing flexibility and adaptability of classification without requiring specific training data for each classification task.

[0011] Additionally, the system can classify inputs by processing category embeddings corresponding to multiple categories using a classifier. The system can generate category embeddings directly by using category labels, by processing text templates that include category labels, by processing category labels using a multimodal model to generate one or category description embeddings, or a combination thereof. In this case, the system can further utilize a multimodal LLM to generate category embeddings, which enables the system to more accurately match categories to inputs.

[0012] In some examples, if the input is an image, the system can generate an image feature embedding by processing the image using an image encoder neural network, and the system can also combine the image feature embedding with the corresponding embeddings of the input description and category prediction to generate a query embedding. Thus, the system can utilize features extracted from both modalities to improve classification accuracy.

[0013] Overall, the described techniques allow for classification of inputs with higher accuracy than simply extracting visual features by generating a textual description of the input to be classified using a multimodal LLM. Specifically, for image classification, the described system can use both textual and image features to classify an image by processing the image description and initial image category prediction using a multimodal LLM.

[0014] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 An example classification system is shown.

[0016] Figure 2 Example diagram showing the input and output of a classification system.

[0017] Figure 3 An example diagram illustrating an example process for classification is shown.

[0018] Figure 4 is a flowchart of an example process for classification using a classification system.

[0019] Figure 5 An example diagram illustrating an example process for classification is shown.

[0020] Figure 6 is a graph of the results of implementing a classification system for a classification task.

[0021] Figure 7 is another diagram of the results of implementing a classification system for a classification task.

[0022] Like reference numbers and designations in the various drawings refer to like elements. DETAILED DESCRIPTION

[0023] Figure 1 Illustrated is an example training system 100. System 100 is an example of a system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below may be implemented.

[0024] System 100 includes a classification system 102 and a user device 104. User device 104 may be a computer, and user device 104 may provide input 116 and a request 118 to classification system 102. Input 116 may be an image, text, audio, or video. Request 118 may include one or more prompts that include a request to classify input 116 into one of a plurality of categories. That is, request 118 may identify a plurality of categories. In some examples, different requests 118 may identify different categories. For example, as shown below: Figure 2 As discussed in further detail in , the input 116 may be an image of a cat, and the plurality of categories may be a plurality of different breeds of cats.

[0025] The system 100 is configured to enhance zero-shot classification by combining both textual and visual features of the input 116 using a multimodal model 106. Traditionally, classification systems rely solely on visual features extracted from images using pre-trained neural networks, which can limit accuracy. For example, existing systems may rely solely on neural networks trained using contrastive language-image pre-training (CLIP) to generate classifications for images.

[0026] In contrast, system 100 utilizes a multimodal language model to generate a textual representation of an input. Specifically, classification system 102 is configured to process input 116 and request 118 from user device 104 to generate a classification 130 for input 116 .

[0027] The classification system 102 includes a multimodal model 106 configured to process an input 116 and a request 118 to generate an initial class prediction for the input 116 and a description of the input 116. The multimodal model 106 can be pre-trained to perform classification such that the multimodal model 106 does not require task-specific fine-tuning before performing a given classification task.

[0028] The classification system 102 further includes a text encoder embedding neural network 108 configured to process multiple text inputs to generate corresponding embeddings and a feature embedding fusion engine 110 configured to combine (e.g., fuse) multiple feature embeddings to generate a query feature embedding.

[0029] The classification system 102 further includes a classifier 112 configured to process the query feature embedding 128 and the plurality of class embeddings to generate a classification 130 .

[0030] In some examples, the classification system 102 further includes an input encoder embedding neural network 114 configured to process the input 116 to generate an input feature embedding 132 .

[0031] Specifically, the classification system 102 processes the input 116 and the request 118 using the multimodal model 106 to generate a category prediction 120 and an input description 122 .

[0032] Class prediction 120 is an initial prediction of the classification of input 116. On the other hand, input description 122 includes text describing input 116. That is, class prediction 120 is a prediction of the class of input 116, while input description 122 is a visual description of input 116, such as Figure 2 and Figure 3 For example, if the multiple categories are types of cats (e.g., Abyssinian cat, American Bulldog, Birman cat, etc.), the category prediction 120 may be "Birman cat" and the input description 122 may be "I see a light gray cat."

[0033] To generate the category prediction 120 and the input description 122, the classification system 102 may provide the first prompt (e.g., image classification prompt) and the second prompt (e.g., image description prompt) to the multimodal model 106 from the request 118. Examples of the first prompt and the second prompt are shown in Table 1 below.

[0034]

[0035] Table 1

[0036] In some examples, request 118 may include a third prompt (eg, a category description prompt) prompting multimodal model 106 to generate a description for each of the plurality of categories. An example of the third prompt is shown in Table 2 below.

[0037]

[0038] Table 2

[0039] The multimodal model 106 can be a language model of any particular architecture that is configured to process inputs of different modalities, such as text and images, to generate outputs. In this case, the multimodal model 106 can be pre-trained on a large set of multimodal data including text and image pairs. That is, the multimodal model 106 is configured to generate text that is aligned with both textual and visual inputs, effectively integrating information from both modalities to generate category predictions 120 and input descriptions 122. In some examples, the multimodal model 106 can be a decoder-only Transformer, such as those used in models like PaLI (Path Language and Image), PaLI Gemma, Flamingo, or Gemini (e.g., Flamingo: A Visual Language Model for Few-Shot Learning (DeepMind, 2022), PaLI: A Jointly Scaled Multilingual Language-Image Model (Google Research, 2022), and Gemini: Google's Multimodal Base Model (Google DeepMind, 2023)).

[0040] The classification system 102 can then encode the category prediction 120 and the input description 122 into corresponding embeddings. Specifically, the classification system 102 can process the category prediction 120 using the text encoder embedding neural network 108 to generate a prediction feature embedding 124. Additionally, the classification system 102 can process the input description 122 using the text encoder embedding neural network 108 to generate a text description feature embedding 126.

[0041] The classification system 102 can then use the feature embedding fusion engine 110 to combine the prediction feature embedding 124 and the text description feature embedding 126 to generate a query feature embedding 128. As used in this specification, a feature embedding is an ordered set of numerical values ​​(e.g., a vector, a sequence of multiple vectors, or a matrix of floating point or other numerical values) that represents an input.

[0042] The classification system 102 can use the text encoder embedding neural network 108 to generate category feature embeddings 136 corresponding to multiple categories. Specifically, the system can generate the category feature embeddings 136 directly by using the category labels, by processing text templates including category labels, by processing the category descriptions 134, or a combination thereof, as described below with reference to Figure 3 described in further detail.

[0043] In some examples, if the input 116 is an image, the classification system 102 can process the input 116 using the input encoder embedding neural network 114 to generate an input feature embedding 132. In this case, the classification system 102 can then use the feature embedding fusion engine 110 to combine the input feature embedding 132 with the prediction feature embedding 124 and the text description feature embedding 126 to generate a query feature embedding 128.

[0044] The text encoder embedding neural network 108 and the input encoder embedding neural network 114 can be pre-trained to generate a joint embedding representation of text and image. For example, the text encoder embedding neural network 108 and the input encoder embedding neural network 114 can be trained to produce aligned embeddings using contrastive learning, wherein the training system can bring matching image-text pairs closer together in the embedding space while pushing mismatched pairs apart. In some other examples, the text encoder embedding neural network 108 and the input encoder embedding neural network 114 can be trained using an image-text matching objective, wherein the training system trains the encoders to classify the correspondence between a given text and image pair (e.g., using a binary classifier and a cross-entropy loss). That is, the text encoder embedding neural network 108 can be any suitable neural network capable of mapping text input to an embedding. For example, the text encoder embedding neural network 108 can be a Transformer, a convolutional neural network, a visual Transformer, or a recurrent neural network. That is, the input encoder embedding neural network 114 can be any suitable neural network capable of mapping image input to an embedding. For example, the input encoder embedding neural network 114 can be a Transformer, a convolutional neural network, a visual Transformer, or a recurrent neural network.

[0045] The classification system 102 may then use the classifier 112 to process the query feature embedding 128 to generate a classification 130. Specifically, the classification system 102 may compare the query feature embedding 128 with the category feature embedding 136, as described below with reference to Figure 3 described in further detail.

[0046] In some examples, the classifier 112 may be a classification engine that may identify the category feature embedding 136 that is closest to the query feature embedding 128 .

[0047] In some other examples, the classifier 112 can be a neural network having any suitable architecture that allows the classifier 112 to generate the classification 130 for the input 116. For example, the classifier 112 can be a convolutional neural network, such as a neural network or a Transformer neural network having a ResNet architecture, a multi-layer perceptron (MLP) architecture, etc. That is, the classifier 112 can be trained on a specific task of processing multiple feature embeddings to classify the input.

[0048] By integrating information from both text and image modalities, system 100 improves classification accuracy, particularly in the zero-shot setting where the class was not seen during training. Importantly, this approach provides flexibility and generalization across different classification tasks without requiring specialized training data.

[0049] Figure 2 For example, reference Figure 1 An example diagram of the input and output of the classification system is depicted in the classification system 102 .

[0050] The classification system 102 is configured to process the input 116 and the request 118 to generate the classification 130. For example, Figure 2 As shown, input 116 can be an image of a cat. In this case, the ground truth class (e.g., the "true label") of input 116 is the Abyssinian cat class.

[0051] The request 118 may include multiple prompts. Specifically, the request 118 may include a first prompt 202 requesting the system to classify the input 116 into a category from a plurality of categories. For example, Figure 2 As shown, the first prompt 202 may be: "Classify the image according to the given category label." The following category labels are: Abyssinian cat, British shorthair cat, Birman cat... " In addition, the request 118 may include a second prompt 204 requesting the generation of a description of the input 116. For example, Figure 2 As shown, the second prompt 204 may be: "What do you see?"

[0052] Classification system 102 can use multimodal model 106 to process input 116 and request 118 to generate input description 122 and category prediction 120. Input description 122 can include a description of input 116, such as "I see a cat...". Category prediction 120 can include an initial category prediction for input 116, such as "The image category is 'Birman cat'." That is, category prediction 120 may be incorrect.

[0053] The classification system 102 can then process the input description 122 and the category prediction 120 using the text encoder embedding neural network 108, the feature embedding fusion engine 110, and the classifier 112 to generate a classification 130 for the input 116, as described below with reference to Figure 3 In some examples, the classification system 102 may also process the input 116 using the input encoder embedding neural network 114 to generate the classification 130, as described below with reference to Figure 3 described in further detail.

[0054] Figure 3 An example diagram illustrating an example process for classification is shown.

[0055] The classification system 102 can process the input 116 and the request 118 to generate a classification 130. For example, Figure 3 As shown, input 116 may be an image of a plurality of rubber pencil erasers.

[0056] The classification system 102 can process the input 116 and the request 118 using the multimodal model 106 to generate a category prediction 120 (eg, “pencil”) and an input description 122 (eg, “There are five pencils in a row…”).

[0057] The classification system 102 can then process the category prediction 120 and the input description 122 using the text encoder embedding neural network 108 to generate a prediction feature embedding 124 and a text description feature embedding 126. In this case, the classification system 102 can also process the input 116 using the input encoder embedding neural network 114 to generate an input feature embedding 132. That is, the classification system 102 can generate a prediction feature embedding 124, a text description feature embedding 126, and an input feature embedding 132 (e.g., input features 124, 126, and 132).

[0058] Additionally, the classification system 102 can generate category feature embeddings 136 (eg, category features 136 ) corresponding to the plurality of categories to perform classification.

[0059] Specifically, the classification system 102 can use the class labels to directly generate the class feature embeddings 136. For example, the classification system 102 can process each class label (e.g., directly from a text label such as "pencil" or "eraser") using the text encoder embedding neural network 108 to generate a corresponding class feature embedding 136.

[0060] In another example, the classification system 102 can use the text encoder embedding neural network 108 to process a text template including a category label (e.g., “photo of {category label},” where {category label} refers to a text label for one of our categories) to generate a corresponding category feature embedding 136.

[0061] In another example, the classification system 102 can process one or more category descriptions using a text encoder embedding neural network to generate one or more category description embeddings, and the system can combine the one or more category description embeddings to generate corresponding category embeddings. In this case, the classification system 102 can use one or more prompts (as shown in Table 2 above) to generate category descriptions 134. The classification system 102 can then use a feature embedding fusion engine to generate a single embedded category feature embedding for each category by combining each of the category description embeddings. In another example, the classification system 102 can generate a category feature embedding 136 by combining (e.g., averaging) two or more of the category embeddings corresponding to the category label, text template, and category description 134.

[0062] The classification system 102 then uses the classifier 112 to generate a classification 130 by classifying the input into one of the categories based on the query feature embedding 128. Specifically, the classification system 102 uses the query feature embedding 128 to determine a corresponding similarity score for each category, and the classification system 102 can generate the classification 130 based on comparing the corresponding similarity scores, as shown in Formula 1:

[0063] (1)

[0064] Where W represents the similarity score, denotes the transpose of the query feature embedding 128, and M denotes the class feature embedding 136. The index of the final classification 130 is calculated as argmax(W) corresponding to the maximum similarity score.

[0065] In some examples, the classification system 102 may instead use the classifier 112 to process the query feature embedding 128 and the category feature embedding 136 to generate a corresponding classification score for each category. The classification system 102 may then select one or more categories based on the corresponding classification scores.

[0066] Figure 3 The graph shows how the query feature embedding 128 is better aligned with the ground truth features compared to the individual input features 124, 126, and 132. This demonstrates that the fusion of visual and textual cues not only improves semantic alignment but also improves classification performance by mitigating the shortcomings of any single modality. Therefore, combining features from different modalities leverages the complementary strengths of the multimodal model 106, the text encoder embedding neural network 108, and the input encoder embedding neural network 114.

[0067] Figure 4 is a flow chart of an example process for classifying using a classification system. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a suitably programmed system (e.g., Figure 1 The classification system 102) can perform process 400.

[0068] The system may receive input and a request to classify the input into one of a plurality of categories (402). The input may be an image. That is, the request 118 may include one or more prompts comprising a request to classify the input 116 into one of a plurality of categories. For example, the request 118 may identify a plurality of categories. In some examples, a user may provide a plurality of different requests 118 that may identify different categories.

[0069] The system can process the input using a multimodal model to generate (i) a description of the input and (ii) a category prediction (404). Specifically, the system can process the input and a first prompt using a multimodal model to generate a category prediction, the first prompt including a corresponding category label for each of a plurality of categories. The system can also process the input and a second prompt to generate a description. The second prompt includes a request to generate a description of the input. For example, the second prompt can be a question asking "What do you see?"

[0070] The system may process the input description and category prediction using a text encoder embedding neural network to generate (i) text description feature embeddings and (ii) prediction feature embeddings (406).

[0071] The system can generate a query feature embedding representing the input based on at least the descriptive feature embedding and the predicted feature embedding (508). In some examples, the system can generate the query feature embedding by processing the input using an image encoder embedding neural network to generate an image feature embedding, and combining the image feature embedding and the predicted feature embedding to generate the query feature embedding. The text encoder embedding neural network and the image encoder embedding neural network can be pre-trained to generate a joint embedding representation of text and image.

[0072] The system can use the query embedding to classify the input into one of a plurality of categories (410). In some examples, the system can use the query embedding to determine a corresponding similarity score for each of the plurality of categories and use the corresponding similarity scores to classify the input.

[0073] In some other examples, the system can use a classifier neural network to process the query embedding and the corresponding category embedding for each of the multiple categories to generate a corresponding classification score for each of the multiple categories, and the system can select one or more categories of the multiple categories based on the corresponding classification scores.

[0074] Specifically, the system can process query embeddings and corresponding category embeddings in the following way: Use a text encoder embedding neural network to process multiple category labels to generate corresponding category embeddings. For each category, the system can obtain a text template including the category label and use a text encoder neural network to process the text template to generate the corresponding category embedding.

[0075] In some examples, the system can process the category labels using a multimodal model to generate one or more category descriptions. The system can process the one or more category descriptions using a text encoder embedding neural network to generate one or more category description embeddings, and the system can combine the one or more category description embeddings to generate corresponding category embeddings.

[0076] In some examples, the system may process two or more of the following using a text encoder neural network to generate corresponding category embeddings: (i) category labels, (ii) text templates including category labels, or (iii) one or more category descriptions generated from category labels via multimodal model embeddings.

[0077] Figure 5 An example diagram illustrating an example process for classification is shown.

[0078] Figure 5 Two examples are shown that illustrate the interpretability of the described classification method by visualizing the contribution of each input (e.g., input image 116, class prediction 120, and input description 122) to the final classification decision (e.g., classification 130) of the classification system 102. For each example, the figure highlights the most influential regions or text tokens that contributed to the classification 130.

[0079] For each example, the original input image is shown, followed by a corresponding heat map indicating the spatial regions of the image that had the greatest impact on the final prediction identified using the sliding kernel masking method. Specifically, the system can measure the resulting change in prediction confidence based on masking patches of the image. Additionally, the system can measure the contribution of the text of the input description 122 at the word level using a similar masking technique based on masking one or more words at a time and observing the impact on the final prediction. The system can label words whose removal results in a significant change in the classification 130 as high contribution terms.

[0080] Thus, these examples illustrate how different input types can contribute differently depending on the context. For example, in some cases (e.g., images of catacombs), visual features dominate, while in other cases (e.g., images of ballpoint pens), the input description 122 can provide the decisive context. This demonstrates the complementary nature of input modalities and the value of fusing features in the proposed system to generate classification 130.

[0081] Figure 6 is a graph of the results of implementing a classification system for a classification task.

[0082] Figure 6 The example compares the classification performance of another approach using a zero-shot classification model (e.g., Comparative Language-Image Pretraining (CLIP)) and the described approach, which utilizes multiple feature representations derived from both visual and textual modalities.

[0083] Specifically, the input image depicts a living room scene. While another method incorrectly classified the image as a "clock tower," the described approach uses features derived from the image description and the class predictions generated by the LLM to correctly identify the primary object of interest (a "rocking chair"). In other words, the results suggest that the descriptions and predicted features derived from the LLM provide additional information that enhances classification.

[0084] This example shows that while alternative approaches are limited to interpreting visual signals, the described approach effectively combines the image description generated by the LLM with the initial category prediction, resulting in a more semantically aligned query representation. Consequently, the described approach improves classification accuracy, especially in situations where visual ambiguity or scene complexity impairs traditional image-only models.

[0085] Figure 7 is another diagram of the results of implementing a classification system for a classification task.

[0086] Figure 7 The graphs illustrate the performance of the described classification system compared to conventional systems using different multiple datasets.

[0087] Specifically, Figure 7 Each chart in the figure shows a confusion matrix illustrating how the performance of zero-shot image classification improves when using the proposed multimodal approach compared to other models (such as CLIP). Specifically, by leveraging features generated by the LLM (e.g., image descriptions and initial class predictions) into the fused feature representation, the system significantly reduces misclassifications, as evidenced by an increased concentration of correctly predicted classes along the matrix diagonal. This indicates that leveraging rich textual semantics alongside visual embeddings achieves more accurate alignment between input features and target classes, outperforming traditional approaches that rely solely on visual input from other pre-trained models.

[0088] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a particular operation or action, this means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when in operation, causes the system to perform that operation or action. For one or more computer programs to be configured to perform a particular operation or action, this means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform that operation or action.

[0089] Embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0090] The term "data processing apparatus" refers to data processing hardware and encompasses all types of equipment, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. An apparatus may also be or include special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, an apparatus may optionally include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0091] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages ​​or declarative or procedural languages, and it may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.

[0092] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and may be stored on a storage device in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which may be organized and accessed differently.

[0093] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed and run on the same computer or computers.

[0094] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.

[0095] A computer suitable for executing a computer program can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or random access memory, or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be supplemented by or incorporated into a dedicated logic circuit system. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.

[0096] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD ROM and DVD-ROM disks.

[0097] To provide for user interaction, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, as well as a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. In addition, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving responsive messages from the user in response.

[0098] A data processing device used to implement a machine learning model may also include, for example, dedicated hardware accelerator units for processing general-purpose and computationally intensive parts of machine learning training or production (i.e., inference, workloads).

[0099] Machine learning models can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework).

[0100] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0101] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from the user. Data generated at the user device, such as the results of the user interaction, may be received at the server from the device.

[0102] Although this specification contains many specific implementation details, these details should not be interpreted as limiting the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be unique to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may involve a subcombination or a variant of a subcombination.

[0103] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.

[0104] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method for classification, comprising: receiving input and a request to classify the input into one of a plurality of categories; processing the input using a multimodal model to generate (i) a description of the input and (ii) a class prediction; Processing the description and the category prediction of the input using a text encoder embedding neural network to generate (i) a text description feature embedding and (ii) a prediction feature embedding; generating a query feature embedding representing the input from at least the descriptive feature embedding and the predictive feature embedding; and The input is classified into one of the plurality of categories using the query embedding.

2. The computer-implemented method of claim 1, wherein the input is an image.

3. The computer-implemented method of claim 2, wherein generating the query feature embedding comprises: processing the input using an image encoder embedding neural network to generate an image feature embedding; as well as The image feature embedding, the description feature embedding, and the prediction feature embedding are combined to generate the query feature embedding.

4. The computer-implemented method of claim 1 , wherein processing the input using the multimodal model to generate the description and the category prediction for the input comprises: The input and a first prompt including a corresponding class label for each of the plurality of classes are processed using the multimodal model to generate the class prediction.

5. The computer-implemented method of claim 1 , wherein processing the input using the multimodal model to generate the description and the category prediction for the input comprises: The input and a second prompt are processed to generate the description, wherein the second prompt includes a request to generate the description of the input.

6. The computer-implemented method of claim 1 , wherein classifying the input into one of the plurality of categories using the query embedding comprises: determining a corresponding similarity score for each of the plurality of categories using the query embedding; as well as The inputs are classified using the corresponding similarity scores.

7. The computer-implemented method of claim 1 , wherein classifying the input into one of the plurality of categories using the query embedding comprises: processing the query embedding and a corresponding category embedding for each of the plurality of categories using a classifier to generate a corresponding classification score for each of the plurality of categories; as well as One or more categories of the plurality of categories are selected based on the corresponding classification scores.

8. The computer-implemented method of claim 3, wherein the text encoder embedding neural network and the image encoder embedding neural network are pre-trained to generate a joint embedding representation of text and image.

9. The computer-implemented method of claim 7, wherein processing the query embedding and the corresponding category embedding for each of the plurality of categories using a classifier to generate a corresponding classification score for each of the plurality of categories comprises: The plurality of category labels are processed using the text encoder embedding neural network to generate the corresponding category embeddings.

10. The computer-implemented method of claim 7, wherein processing the plurality of category labels using the text encoder embedding neural network to generate the corresponding category embeddings comprises, for each category: obtaining a text template including the category label; and The text template is processed using the text encoder neural network to generate the corresponding category embedding.

11. The computer-implemented method of claim 7, wherein processing the plurality of category labels using the text encoder embedding neural network to generate the corresponding category embeddings further comprises, for each category: processing the category labels using the multimodal model to generate one or more category descriptions; processing the one or more category descriptions using the text encoder embedding neural network to generate one or more category description embeddings; as well as The one or more category description embeddings are combined to generate the corresponding category embedding.

12. The computer-implemented method of claim 11 , wherein processing the plurality of category labels using the text encoder embedding neural network to generate the corresponding category embeddings comprises, for each category: The text encoder neural network is used to process two or more of the following to generate the corresponding category embedding: (i) the category label, (ii) a text template including the category label, or (iii) one or more category descriptions generated from the category label by embedding the multimodal model.

13. A system comprising: one or more computers; as well as One or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: receiving input and a request to classify the input into one of a plurality of categories; processing the input using a multimodal model to generate (i) a description of the input and (ii) a class prediction; Processing the description and the category prediction of the input using a text encoder embedding neural network to generate (i) a text description feature embedding and (ii) a prediction feature embedding; generating a query feature embedding representing the input from at least the descriptive feature embedding and the predictive feature embedding; and The input is classified into one of the plurality of categories using the query embedding. The system of claim 13 , wherein the input is an image.

15. The system of claim 14, wherein generating the query feature embedding comprises: processing the input using an image encoder embedding neural network to generate an image feature embedding; as well as The image feature embedding, the description feature embedding, and the prediction feature embedding are combined to generate the query feature embedding.

16. The system of claim 13, wherein processing the input using the multimodal model to generate the description and the category prediction for the input comprises: The input and a first prompt including a corresponding class label for each of the plurality of classes are processed using the multimodal model to generate the class prediction.

17. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising: receiving input and a request to classify the input into one of a plurality of categories; processing the input using a multimodal model to generate (i) a description of the input and (ii) a class prediction; Processing the description and the category prediction of the input using a text encoder embedding neural network to generate (i) a text description feature embedding and (ii) a prediction feature embedding; generating a query feature embedding representing the input from at least the descriptive feature embedding and the predictive feature embedding; and The input is classified into one of the plurality of categories using the query embedding.

18. The one or more non-transitory computer storage media of claim 17, wherein the input is an image.

19. The one or more non-transitory computer storage media of claim 18, wherein generating the query feature embedding comprises: processing the input using an image encoder embedding neural network to generate an image feature embedding; as well as The image feature embedding, the description feature embedding, and the prediction feature embedding are combined to generate the query feature embedding.

20. The one or more non-transitory computer storage media of claim 17, wherein processing the input using the multimodal model to generate the description and the category prediction for the input comprises: The input and a first prompt including a corresponding class label for each of the plurality of classes are processed using the multimodal model to generate the class prediction.

Citation Information

Cited By

  • Machine learning embeddings for evolving category sets

    US12711425B2

  • Machine learning embeddings for evolving category sets

    US20240420018A1