Semantic tagging of images using generative language models

By introducing the OmniScient model, using generative frameworks and large language models to predict class names, the challenge of object recognition in the open physical world is solved, and the ability to train across data sets without predefined vocabulary and without human intervention is achieved.

CN120014642APending Publication Date: 2025-05-16FACE CUTE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411616814.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-11-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing machine-aware technologies have difficulty accurately positioning and identifying objects in an open physical world, especially when dealing with new concepts beyond the scope of training, where there is a problem of class name predefined and tag definition conflict.

Method used

The OmniScient Model (OSM) is introduced. This model adopts a generative framework to predict class names through large language models, without predefined vocabulary, can be trained across datasets without manual intervention, and has robust generalization capabilities.

Benefits of technology

OSM shows strong generalization capabilities when dealing with new concepts and training across datasets, can present promising results on various benchmarks, and effectively solves the problem of class name predefined and tag definition conflicts in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014642A_ABST
    Figure CN120014642A_ABST
Patent Text Reader

Abstract

Embodiments of the invention relate to semantic tagging of an image using a generative language model. A computing system includes one or more processing devices configured to receive an image. The processing device is further configured to calculate a segmentation mask identifying a region of interest included in the image. At the feature extractor, the processing device is further configured to compute an encoded image feature based on the image. The processing device is also configured to receive a textual instruction. At the visual resampler, the processing device is further configured to compute a mask query based on the segmented mask, the encoded image features, and the textual instructions. At the generative language model, the processing device is further configured to receive a natural language query including a mask query and a textual instruction. Based on the natural language query, at the generative language model, the processing device is further configured to generate and output semantic tags associated with the region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] With the development of deep learning methods, machine perception has become an area that has widely used machine learning techniques in recent years. One such field of machine perception is computer vision, which includes tasks such as image recognition and tagging. In applications such as automatic captioning and optical character recognition, as well as in applications such as autonomous driving that involve extracting semantic understanding from images of the device's physical environment, machine learning models are often used to perform computer vision tasks. Summary of the invention

[0002] According to one aspect of the present disclosure, a computing system is provided, the computing system including one or more processing devices, the one or more processing devices being configured to receive an image. The one or more processing devices are also configured to calculate a segmentation mask of a region of interest included in the identified image. At the feature extractor, the one or more processing devices are also configured to calculate a plurality of encoded image features based at least in part on the image. The one or more processing devices are also configured to receive text instructions. At the visual resampler, the one or more processing devices are also configured to calculate a mask query based at least in part on the segmentation mask, the plurality of encoded image features and the text instructions, the mask query comprising a plurality of text symbols. At the generative language model, the one or more processing devices are also configured to receive a natural language query including a mask query and a text instruction. At the generative language model, the one or more processing devices are also configured to generate a semantic label associated with the region of interest based at least in part on the natural language query. The one or more processing devices are also configured to output a semantic label.

[0003] This summary is provided to introduce some concepts in a simplified form, which will be further described in the following detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. In addition, the claimed subject matter is not limited to embodiments that address any or all of the disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 Open recognition performed on an image according to an example embodiment is illustrated.

[0005] Figure 2 Examples of a closed vocabulary recognition setting, an open vocabulary recognition setting, and an open recognition setting are schematically illustrated according to an example embodiment.

[0006] Figure 3A Schematically shows the Figure 1 An example of a computing system configured to execute a model (OmniScient Model (OSM)) during inference time.

[0007] Figure 3B Schematically illustrates the time during training when a visual resampler included in OSM is trained according to one example Figure 3A computing system.

[0008] Figure 4 Schematically shows the Figure 3A The example visual resampler has more details on its architecture.

[0009] Figure 5 Shown according to Figure 3A Table of values ​​for classification accuracy and not-in-vocabulary (NIV) rate for OSM for an example.

[0010] Figure 6 Including according to Figure 3A Table of examples showing the results of ablation experiments on OSM.

[0011] Figure 7 Shown according to Figure 3A An example table comparing the performance of OSM with that of other common segmentation models.

[0012] Figure 8 Shown according to Figure 3A A table of example instruction templates that may be received at a generative language model.

[0013] Fig. 9 Shown according to Figure 3B Example plots of NIV and accuracy of OSM with different numbers of training masks.

[0014] Fig.10 Shown according to Figure 3A Example images labeled with part-level and box-level data.

[0015] Fig.11A Shown according to Figure 3A Example images of examples of , and the corresponding segmentation masks and semantic labels generated for these images using SAM as a segmentor.

[0016] Fig. 11B Shown according to Figure 3A Example images of examples of , and the corresponding segmentation masks and semantic labels generated for these images using kMax-DeepLab as the segmentor.

[0017] Fig.12 Shown according to Figure 3A An example of a labeled image providing a qualitative comparison between the relevant multimodal macro model and OSM.

[0018] Fig.13 Shown according to Figure 3A Table of examples comparing OSM to traditional open vocabulary tagging methods.

[0019] FIG. 14A to FIG. 14B Shown according to Figure 3A An example of where OSM assigns NIV tags to an image.

[0020] Fig.15A Schematically shows the Figure 3A A flowchart of an example method for use with a computing system to assign semantic tags to an image.

[0021] FIG. 15B to FIG. 15D shows that can be performed in some examples Fig.15A Additional steps of the method.

[0022] Fig.16 A schematic diagram of an example computing environment is shown in which a FIG. 3A to FIG. 3B computing system. DETAILED DESCRIPTION

[0023] Localizing and recognizing objects in an open physical world poses a long-standing challenge in the field of machine perception. Recent approaches strive to address the problem by adopting class-agnostic mask (or box) proposal models, supplemented by open vocabulary classifiers trained using pre-extracted text embeddings (e.g., CLIP). But it is worth noting that these open vocabulary recognition models still show limitations in practical applications. On the one hand, they rely on providing class names during testing, and the recognition performance depends heavily on this set of semantic classes predefined by the user. On the other hand, when using multiple datasets for training, manual intervention is required to alleviate the label definition conflicts between them.

[0024] In this disclosure, a large language model (LLM)-based mask classifier, referred to herein as the OmniScient Model (OSM), is introduced to address the above challenges. Specifically, OSM predicts class labels in a generative manner, eliminating the need to provide class names during training and testing. It is also capable of cross-dataset training without human intervention and exhibits robust generalization capabilities thanks to the world knowledge acquired from the LLM. By combining OSM with a mask proposal model, promising results are presented on various benchmarks, and its effectiveness in processing out-of-domain images is demonstrated.

[0025] 1. Introduction

[0026] A technical challenge in the field of machine perception involves accurately localizing and identifying objects in the real world. Despite significant progress on various standard benchmarks, existing methods still struggle to cope with the complexity of real-life scenarios, where new concepts outside the training dataset often appear. To address this issue and improve the practical utility of models, a common strategy is to decompose the problem into two components: class-agnostic mask / box proposal and mask / box classification. It has been found that when trained on datasets such as COCO, mask / box proposal models can still effectively generalize to previously unseen concepts. In addition, taking the Segment Anything Model (SAM) as an example, recent progress has expanded the training dataset to a very large scale, covering 1.1 billion class-agnostic masks from 11 million images. This extension produces a mask proposal model that is characterized by robust zero-shot segmentation capabilities and generalization to new images and concepts.

[0027] Despite such progress in the development of general proposal models, addressing the challenge of classifying novel concepts in real-world scenarios remains an unsolved problem. Many existing methods leverage visual language models (VLMs), such as CLIP and ALIGN, which have been pre-trained on a large number of Internet datasets and have been shown to excel in aligning images and text within a shared embedding space. Specifically, these existing methods train open vocabulary classifiers that rely on pre-computed text embeddings derived from VLMs, rather than learning label embeddings directly from training datasets. The reliance on VLM text embeddings highlights the inherent power and generalization ability of VLMs, which to some extent guarantees the classifier's ability to generalize to novel concepts.

[0028] Although the methods discussed above have shown promise, they still face several challenges that hinder their practical application. First, these models usually operate under the assumption that the class names are predefined during testing, which is rarely the case in real life. In addition, when leveraging multiple different datasets, complications arise if there are different label definitions or label space conflicts between them, such as differences in the level of abstraction or specificity between label spaces. Therefore, many current multi-dataset frameworks address this issue by using separate decoders or classifiers trained on each dataset, or manually merge the label spaces, which increases the complexity of the process.

[0029] To address these challenges, a machine learning model is introduced, referred to in this paper as the OmniScient Model (OSM). OSM includes a generative framework that can be applied to open recognition tasks. Instead of training the model to "select" the correct class from a predefined vocabulary, the approach focuses on training the model to generate the desired class names. This paradigm shift means that the model no longer needs to be provided with prior knowledge of all possible class names by the user, thereby eliminating the need for an explicitly defined vocabulary during both the training and testing phases. As a result, this approach is naturally suitable for training and testing on datasets with different label spaces, eliminating the need for human intervention to reconcile the differences. In addition, by building on a pre-trained large language model (LLM), OSM leverages the implicitly learned world knowledge encoded in the LLM, enhancing its ability to effectively generalize to new concepts, further improving its practicality and reliability.

[0030] The experiments used to evaluate the appropriateness of employing generative models for discriminative tasks are discussed below. The investigation consists of evaluating the ability of generative models to effectively capture and adapt to the characteristics of a given training dataset and its associated vocabulary. Its performance is compared with that of discriminative models, focusing mainly on classification accuracy. In addition, a pattern query mechanism is introduced that enables the model to make predictions in a predefined vocabulary (called vocabulary-specific predictions) or provide open-ended predictions that are not restricted by the vocabulary (called vocabulary-independent predictions). Finally, OSM is integrated with various off-the-shelf segmenters (i.e., mask proposal models) such as kMax-DeepLab and SAM, and its effectiveness is verified on multiple benchmarks.

[0031] 2. Related Work

[0032] Open vocabulary recognition

[0033] Recently, open vocabulary detection methods have achieved promising results, exemplified by CLIP and ALIGN. These methods involve pre-training dual encoder models (for images and text) using a contrastive objective on a large collection of noisy image-text pairs. The feature representations produced by this pre-training process are cross-model capable, demonstrating robust performance on zero-shot downstream tasks. Inspired by these advances, significant breakthroughs have also been achieved in the fields of open vocabulary detection and segmentation, where the class names provided during testing may not have been encountered during the training phase. Most of these state-of-the-art techniques address the problem by decomposing it into class-agnostic proposals and perform open vocabulary proposal classification by leveraging the pre-trained CLIP model. However, while these open vocabulary methods have achieved impressive results in recognizing unseen classes beyond the training dataset, they all rely on a strong but fragile assumption that the semantic classes (i.e., vocabulary) are known in advance and remain static, an assumption that can be easily broken in real applications. In parallel with the research work, vocabulary-free image classification attempts to address this challenge by dynamically generating vocabulary through processes such as parsing captions or retrieving them from external databases. In contrast, the methods discussed below reformulate the open classification problem as text generation, which naturally does not require a user-defined vocabulary.

[0034] Figure 1 An example of open recognition is shown in the figure. Figure 1 In the example of , an input image 10 is shown, as well as three different indicated locations 12 within the input image 10. The open recognition task is decomposed into two subtasks: class-independent mask proposal and open mask classification. To solve this task, an open mask classifier OSM 30 works in conjunction with a class-independent mask proposal model 20 (e.g., SAM) at which a segmentation mask 22 is calculated. Unlike existing open vocabulary recognition models, OSM 30 does not require any user-predefined vocabulary, but can directly use an unconstrained vocabulary in a generative manner to predict the class 40 of each proposal. Therefore, OSM 30 exhibits strong generalization capabilities. Figure 1 In the example of , the image predicts that it is a dog, and predictions for new parts such as tail and ears are observed, even though OSM 30 has never seen such masks 22 or labels 40 during training. Furthermore, by obtaining the mask 22 from a class-agnostic segmenter 20, a wide range of cue types including points, boxes, and masks can be utilized.

[0035] Large Language Model

[0036] In recent years, the research community has witnessed a significant growth in the development of large language models (LLMs). These models have demonstrated impressive emerging capabilities, including in-context learning, instruction following, and thought chain reasoning. However, a significant limitation of these LLMs is their inherent "blindness" to other modalities (such as visual input). Recently, multimodal LLMs such as the related multimodal large model have been developed. Pioneering research has provided a promising approach to bridge the gap between language and visual modalities. This approach involves building modular models that typically consist of a frozen CLIP visual encoder, a trainable bridge module, and a frozen LLM. In addition, the ability to reference or locate grounding can be added to multimodal LLMs by taking bounding boxes as input or output. The proposed OSM can be classified as a modular multimodal LLM with reference capabilities. However, previous efforts have mainly aimed to enhance multimodal LLMs with bounding boxes for dialogue applications (because bounding boxes can be naturally represented in text by referencing their coordinates), which also requires providing a vocabulary in the input prompt. The techniques discussed in this paper highlight the value of enabling multimodal LLM to identify segmentation masks and serve as a standalone tool.

[0037] 3. Methods

[0038] This section introduces the Open Classifier OSM (OmniScient Model). The traditional classification task is transformed into a text generation task, in keeping with the principles outlined in (Section 3.1). The construction of OSM is explained, which follows the previous modular vision-language model (Section 3.2). A comprehensive overview of the training and evaluation protocol is also provided (Section 3.3).

[0039] 3.1. Classification Problem Statement

[0040] Without loss of generality, this discussion focuses on mask classification. Given an input image and a series of M segmentation masks In the case of masks (from a pre-trained segmenter such as SAM), the goal is to predict the semantic class for each of these masks:

[0041]

[0042] Among them, m i is the i-th mask in M, and c i is its predicted class, which belongs to a predefined set of semantic classes C, which is assumed to be known during both training and testing. In the closed vocabulary setting, the model only focuses on the target class, which means that the predefined set of semantic classes is the same during training and testing (i.e., C 训练 =C 测试, where the subscripts denote the training or testing phase). In contrast, in the open vocabulary setting, this assumption is relaxed, i.e., allowing C 测试 can include new categories not seen during training (i.e., C 测试 ≠C 训练 ). Nevertheless, in both cases, C is used during both the training and testing phases. 训练 and C 测试 Therefore, recognition performance depends largely on the C 训练 and C 测试 of careful design.

[0043] The above assumptions (i.e., for C 训练 and c 训练 ) plays a key role in contemporary recognition frameworks, whether operating in closed or open vocabulary settings. These frameworks typically rely on computing similarity logits between semantic class candidates and selecting the candidate with the highest probability as the final prediction. While these recognition methods have demonstrated their effectiveness and success on various tasks and benchmarks over the past few decades, they are not without severe limitations. First, it is practically impossible to predefine and cover all latent semantic classes that exist in the real world. This limitation poses a significant challenge to previous open vocabulary recognition methods as it requires a priori definitions of new concepts in the vocabulary. Furthermore, many of these methods are built around handcrafted and carefully designed label spaces in the hope of covering common concepts, which ideally should have well-defined definitions. However, manual curation of label spaces may not be scalable, especially when researchers aim to scale up their models to cover all available datasets from various sources. This process may require labor-intensive work such as meticulous manual merging or separate training.

[0044] To address these challenges, this paper proposes a paradigm called open visual recognition, in which the vocabulary C is considered unknown both during training and testing. Figure 2 This shift in perspective is illustrated to provide an overall comparison of the different paradigms. Figure 2Examples of a closed vocabulary recognition setting 100, an open vocabulary recognition setting 120, and an open recognition setting 130 are schematically shown. In the closed vocabulary recognition setting 100, the set of semantic classes 104 forming a user-defined vocabulary is fixed during both training and testing. One learnable predictor 108 (e.g., a 1×1 convolutional layer) is used for each training data set. The closed vocabulary recognition setting 100 also utilizes an image encoder 106 that receives an image 102 and outputs encoded image data to the learnable predictor. The learnable predictor 108 that receives the encoded image data corresponds to the data set to which the image 102 belongs. The learnable predictor 108 outputs logits 110, from which semantic labels 112 are selected.

[0045] Figure 2 Further shown is an open vocabulary recognition setting 120. In the open vocabulary recognition setting 120, the semantic clusters can be different during training and testing, allowing new concepts to be detected during testing by leveraging a pre-trained visual transformer backbone (e.g., CLIP). The image encoder 106 receives the image 102, and the text encoder 122 receives the semantic cluster 104. The text-based predictor 124 (i.e., the text embedding of the pre-defined semantic cluster 104) is different for each dataset.

[0046] Figure 2 Further illustrated is an open recognition setting 130. In the open recognition setting 130, the LLM-based predictor 132 directly predicts the class name 112 in a generative manner, thereby eliminating the need to predefine the semantic classes 104 during training and testing. In addition, the open recognition setting 130 allows for easier cross-dataset training (e.g., without manual involvement to resolve label definition conflicts between datasets).

[0047] Instead of selecting the predicted class from a predefined vocabulary, the approach used in OSM directly predicts the class name of the target object. This direct prediction reformulates the recognition task as a text generation problem. Mathematically, open recognition is viewed as maximizing the conditional likelihood of the class name under a forward autoregressive factorization:

[0048]

[0049] Among them, c i,j Corresponds to the class name c i The jth text symbol in .

[0050] Model Architecture

[0051] According to an example embodiment, FIG. 3A to FIG. 3B An overview of the OSM 30 architecture is presented in Figure 3A is shown at inference time and at Figure 3B 200 is shown as being at training time. OSM 30 is configured to execute on a computing system 200 including one or more processing devices 202 and a memory 204. Figure 3A , OSM 30 is depicted at inference time. OSM 30 includes three main components: a frozen feature extractor 210 (e.g., an open vocabulary classifier such as CLIP-ViT visual transformer), a trainable visual resampler 220 (e.g., MaskQ-Former), and a frozen generative language model 240 (e.g., a large language model (LLM)).

[0052] like Figure 3A As shown, one or more processing devices 202 are configured to receive an image 10. In addition, at a pre-trained segmenter 20 (e.g., SAM), the one or more processing devices 202 are further configured to calculate a segmentation mask 22 identifying a region of interest included in the image 10. In some examples, such as Figure 3A As shown, the one or more processing devices 202 are configured to compute a plurality of different segmentation masks 22 corresponding to different regions of interest within the same image 10 .

[0053] High-resolution feature extraction using frozen CLIP-ViT

[0054] At feature extractor 210, one or more processing devices 202 are further configured to compute a plurality of encoded image features 212 based at least in part on image 10. For example, encoded image features 212 may be pixel embeddings. As described above, feature extractor 210 may be a CLIP-ViT visual transformer or some other open vocabulary classifier.

[0055] The frozen Visual Transformer (ViT) backbone pre-trained in the CLIP style has become the standard choice for existing multimodal LLM designs. The appeal of CLIP-ViT lies in its dual advantages: it provides a robust and adaptable feature representation for input images, and its feature space is well suited for seamless conversion to language symbols that can be understood by LLMs as input.

[0056] Nevertheless, while CLIP-ViT has been used successfully in many multimodal LLM applications such as image captioning and visual question answering, it also has limitations. It was originally pre-trained on a lower resolution, typically 224×224. This lower resolution may affect its performance, especially when performing object-level recognition tasks. In addition, previous studies have found that frozen ViT exhibits weak generalization capabilities across different input resolutions.

[0057] Although the frozen ViT backbone is widely used in multimodal LLM models, it is clear that the input resolution of 224×224 is insufficient, especially for object-level recognition in larger images. Typical adaptation methods (such as windowed attention as seen in ViTDet) may not be applicable to the fully frozen ViT backbone. To address this limitation, the following strategy is introduced to extract more effective features using frozen ViT at higher resolutions (e.g., 896×896). Specifically, a sliding window feature extraction method is adopted at the input level, where each window size matches the pre-trained image size of ViT. Therefore, as Figure 3A As shown, one or more processing devices 202 are configured to compute encoded image features 212 at least in part by sampling multiple windows 214 of the image 10. These windows 214 are spatial windows of the image data contained in the image 10. In such an example, the multiple windows 214 each have a window size that is smaller than the total size of the image 10. These sampled windows 214 are used as input to the feature extractor 210. Thereafter, a global position embedding is added to compensate for the missing position information across the windows 214. This strategy significantly improves the performance of feature extraction compared to using a high-resolution input.

[0058] MaskQ-Former

[0059] A visual resampler 220 (e.g., a Q-Former or a perceptual resampler) is used to bridge the gap between the encoded image features 212 and the input suitable for the LLM 240. Therefore, the term "visual resampler" as used herein refers to a machine learning model that converts image feature data into LLM queries. The visual resampler 220 may include a stack of transformer decoders that transform image symbols into a reduced set of query symbols, which are typically much smaller in number than the image symbols. However, existing visual resamplers employ a set of queries that globally focus on image features without considering segmentation mask priors.

[0060] To address this limitation, a novel variant of the visual resampler 220, called MaskQ-Former, is introduced. MaskQ-Former takes a segmentation mask as input and performs mask cross attention. At the MaskQ-Former visual resampler 220, one or more processing devices 202 are also configured to compute a mask query 222 based at least in part on the segmentation mask 22, the plurality of encoded image features 212. The mask query 222 includes a plurality of text tokens, and the one or more processing devices 202 are configured to input these text tokens into a generative language model 240.

[0061] The input to the MaskQ-Former visual resampler 220 also includes text instructions 230. The text instructions 230 are natural language input received at the one or more processing devices 202 in the form of a plurality of text symbols 232. The text instructions 230 instruct the visual resampler 220 to identify objects depicted in the region of interest defined by the segmentation mask 22. Figure 3A In the example of , the text instruction is “What’s in the segmentation mask?” The visual resampler 220 is configured to generate a mask query 222 based at least in part on the text instruction 230 as well as the segmentation mask 22 and the encoded image features 212 .

[0062] exist Figure 3A In the example of , the input to MaskQ-Former includes two sets of learnable queries: mask queries 222 and context queries 226. The mask queries 222 perform mask cross-attention, whose focus is limited to the focus region associated with the segmentation mask 22, while the context queries 226 focus on a wider area derived from the segmentation mask 22, such as a bounding box area, to provide supplementary context information. The one or more processing devices 202 are also configured to calculate the context query 226 associated with the bounding box 228 surrounding the focus region, and the visual resampler 220 is also configured to receive the context query 226 as input. By incorporating context information related to the area surrounding the focus region, the context information can be used to make object recognition more accurate and unbiased.

[0063] Figure 4 Schematically illustrates more details of the architecture of the MaskQ-Former visual resampler 220 according to one example. Figure 4 In the example of , the visual resampler 220 has a transformer architecture that includes a plurality of transformer layers 250. Each transformer layer 250 includes a self-attention layer 252, a mask cross-attention layer 254, a context cross-attention layer 256, a first feed-forward layer 258, and a second feed-forward layer 260. The self-attention layer 252 is configured to receive the mask query 222, the context query 226, and the text symbol 232 included in the text instruction 230.

[0064] The self-attention layer 252 is also configured to transmit its output vector to the mask cross-attention layer 254, the context cross-attention layer 256, and the second feed-forward layer 260. The mask cross-attention layer 254 and the context cross-attention layer 256 are also configured to receive the encoded image features 212 and the segmentation mask 22 as input. The mask cross-attention layer 254 and the context cross-attention layer 256 are also configured to transmit their respective output vectors to the first feed-forward layer 258.

[0065] The first feed-forward layer 258 is configured to recalculate the mask query 222 and the context query 226 in response to receiving the output vectors of the mask cross-attention layer 254 and the context cross-attention layer 256. In addition, the second feed-forward layer 260 is configured to recalculate the text symbol 232 in response to receiving the output of the self-attention layer 252. These recalculated mask queries 222, context queries 226, and text symbols 232 can be used as inputs to the subsequent transformer layer 250.

[0066] The parameters of the mask cross-attention layer 254 and the context cross-attention layer 256 can be shared. In addition, the mask query 222 focuses on the mask area in the mask cross-attention layer 254, while the context query can focus on a larger area around the segmentation mask 22. These queries and symbols communicate with each other in the self-attention layer 252.

[0067] The MaskQ-Former summarizes the attention region while retaining access to the contextual content. The information exchange between the mask query 222 and the context query 226 is facilitated by the self-attention layer 252. In addition to the learnable query initialization, parameters are shared between the mask query 222 and the context query 226, resulting in negligible additional cost. The mask query 222 output by the final transformer layer 250 is included in the input of the generative language model 240.

[0068] Mode Query

[0069] In some examples, such as Figure 3A As shown, pattern queries 224 are used to align the output of the MaskQ-Former visual resampler 220 with specific vocabulary. These pattern queries 224 take advantage of the powerful instruction-following capabilities of the LLM, thereby enhancing the adaptability of the OSM 30 across different scenarios. Specifically, for both the MaskQ-Former and LLM inputs, a dedicated learnable query is appended for each vocabulary. The visual resampler 220 is configured to receive the pattern query 224 indicating a vocabulary-specific pattern 234, and to compute the mask query 222 based at least in part on the pattern query 224. The pattern query 224 can be appended to the context query 226, such as Figure 4 The vocabulary-specific pattern 234 may be a vocabulary-specific pattern, where the pattern query 224 includes a plurality of predefined classification labels. Alternatively, the vocabulary-specific pattern 234 may be a vocabulary-independent pattern, where the pattern query 224 does not include a predefined classification label. Thus, the MaskQ-Former visual resampler 220 may be used with an open or closed vocabulary.

[0070] Generative Language Models

[0071] The generative language model 240 is configured to receive a natural language query 242 including the mask query 222 and the text instruction 230. Figure 3A In the example of , the natural language query 242 also includes a context query 226 and a pattern query 224. The generative language model 240 is also configured to generate a semantic label 112 associated with the region of interest for which the segmentation mask 22 is calculated based at least in part on the natural language query. Therefore, the one or more processing devices 202 are configured to generate a semantic label 112 for the region of the image 10 corresponding to the segmentation mask 22.

[0072] 3.3. Training and Evaluation Protocol

[0073] Dataset

[0074] Figure 3B The OSM 30 is schematically shown during training, as discussed above. The visual resampler 220 is trained using a training corpus 270 including a plurality of training images 272. In addition, the training corpus 270 also includes a plurality of ground truth masks 274 associated with corresponding training regions of interest within the training images 272. The training corpus 270 also includes a plurality of ground truth labels 276 associated with the ground truth masks 274.

[0075] In the experiments discussed below, in order to create a robust training and evaluation framework, six publicly available segmentation datasets are integrated together to form a training corpus 270. These segmentation datasets cover a variety of image distributions, domains, and segmentation tasks. These datasets include COCO panoptic segmentation, ADE20K panoptic segmentation, Cityscapes panoptic segmentation, LVIS instance segmentation, ADE-847 semantic segmentation, and PC-459 semantic segmentation.

[0076] Training Protocol

[0077] During training, the visual resampler 220 is trained via instruction fine-tuning. This instruction fine-tuning approach facilitates the integration of the visual resampler 220 with the generative language model 240 during training. Specifically, for each training iteration, a training image 272 and its corresponding ground truth mask 274 are randomly selected from the training corpus 270. In the experiments, an instruction template is randomly selected and the ground truth label 276 is inserted into the instruction template. This approach allows the visual resampler 220 to be trained using the next symbol prediction loss function 280. The template "What's in the segmentation mask?" is used as the default template, and greedy search decoding is used during testing.

[0078] The choice of training batch size varies across datasets, with a batch size of 32 for COCO, 64 for LVIS, 16 for ADE-847, 8 for PC-459, 16 for ADE20K, and 8 for Cityscapes. In each training batch, half of the inputs activated vocabulary-specific queries corresponding to their corresponding datasets, while the other half activated vocabulary-independent queries. The AdamW optimizer was used with a learning rate of 4×10-5 and a weight decay of 0.05. The learning rate follows a cosine decay schedule. Training is performed until the model has processed a total of 6 million masks.

[0079] Back to Figure 3B In some examples, the training corpus 270 is a union of multiple training data subsets 278, and the corresponding reference truth labels 276 in the multiple training data subsets have different corresponding label spaces. The label space of the training data subset 278 is the set of all different reference truth labels 276 included in the training data subset 278. Therefore, the label space defines a codomain of potential labels that can be assigned to the reference truth masks 274 included in the training data subset 278. The one or more processing devices 202 are also configured to calculate multiple pattern queries 224, which are respectively associated with the training data subsets 278 and indicate corresponding label spaces. In such an example, the visual resampler 220 is also configured to receive the pattern queries 224 during training.

[0080] During training, when utilizing datasets from various sources (which form the training data subset 278 of the training corpus 270), corresponding vocabulary-specific queries for each dataset are activated, allowing the visual resampler 220 to effectively "memorize" the associated vocabulary of each dataset, thereby improving alignment with that vocabulary during prediction. Additionally, to maintain open-ended recognition capabilities, a general vocabulary-independent query is included that is activated during training on each dataset. This approach provides flexibility during testing. Vocabulary-specific queries can be activated to better align OSM's predictions with the desired vocabulary, or vocabulary-independent queries can be activated to facilitate open-ended predictions. This adaptability enhances the usefulness of OSM across a range of real-world scenarios, making it a versatile tool for a wide range of applications.

[0081] Evaluation Protocol

[0082] On the validation set of each dataset, the model is evaluated using the following two types of masks: 1) the ground truth mask 274, and 2) the segmentation mask 22 produced by the pre-trained segmenter 20. When using the ground truth mask 274 as input, the prediction is considered correct only when the predicted class name exactly matches the class name in the ground truth annotation. To enhance the reliability of the metric, the ground truth class names are augmented with synonyms. In addition, the plural and singular forms of the class names are also considered. These synonyms are not used during model training because they are not always semantically aligned (e.g., "person", "man", and "woman" are synonyms in COCO and LVIS). As a result, two metrics are reported: accuracy (Acc) and not in vocabulary (NIV), which represent the percentage of predictions that correctly match the ground truth class, or predictions that are not in the vocabulary of the dataset, respectively. The metric Acc directly evaluates the classification ability of the model, while NIV reflects the generalization ability or overfitting degree of the model on the training corpus 270.

[0083] In addition, a more practical application is considered, where OSM 30 is connected to a pre-trained mask proposal model (e.g., kMax-DeepLab or SAM). The performance of the model is directly evaluated on established academic benchmarks including panoptic segmentation and semantic segmentation, using panoptic quality (PQ) and mean intersection over union (mIoU), respectively.

[0084] 4. Experimental Results

[0085] In this section, the settings used for ablation experiments and the final model are provided. In Section 4.1, OSM is evaluated using ground truth masks and ablation experiments are also performed. In Section 4.2, OSM is evaluated in a configuration using a pre-trained mask proposal model.

[0086] Default settings for ablation

[0087] Unless otherwise specified, ablation experiments use the following default settings: During training, both images and masks are resized until the longer side reaches a length of 896 pixels, and then the shorter side is padded to match this length. Minimal data augmentation is performed, limited to random flips. Context queries in MaskQ-Former focus on the entire image. OSM is initialized with weights pre-trained with InstructBLIP, which uses EVA-ViT-g / 224 as the visual encoder and Vicuna-7B as the LLM. 32 mask queries, 32 context queries, and 1 pattern query are used. The pattern query is randomly selected between vocabulary-independent queries (shared across datasets) and vocabulary-specific queries (one for each dataset).

[0088] Settings used for the final model

[0089] Based on findings from ablation experiments (described in detail later in the Results), for the final model, the image resolution was increased to 1120, and the context query focused on a bounding box region 0.5 times larger than the mask region constrained by the box. Random scale jittering in the range of [0.5, 1.5] was used.

[0090] 4.1. Mask Classification Using Ground Truth Masks

[0091] Generative models for discriminative tasks

[0092] Figure 5 Table 1 shows the mask classification accuracy on six segmentation datasets using ground truth masks. OSM(vocabulary independent) and OSM(vocabulary specific) are obtained from the same model and weights, but activated for vocabulary independent or vocabulary specific queries respectively during inference. NIV stands for out of vocabulary. Notation Represents the final model settings.

[0093] In Table 1, it is shown that the generative model can effectively capture the training corpus 270, thereby producing predictions that are well aligned with the training vocabulary. Specifically, as shown in the first few rows of the table ("Single Dataset"), OSM 30 is first trained on each of the six segmented datasets separately, and its mask classification accuracy is evaluated using the ground truth mask 274. It is worth noting that although the model is tasked with generating class names without restriction, it consistently delivers predictions within the vocabulary range of its corresponding training data subset 278. This consistency is reflected in the high percentage of accurate predictions (i.e., high Acc scores) and the very low percentage of predictions that fall out of the vocabulary (i.e., low NIV scores), demonstrating the ability of the generative model to perform discriminative tasks.

[0094] Next, we explore the example of training with all six datasets (“Multiple Datasets” in the table). In this example, even in the presence of potential label conflicts, OSM 30 still maintains high accuracy for each individual dataset. Specifically, the proposed pattern query scheme effectively alleviates label conflicts between datasets, where vocabulary-specific queries (“Vocabulary-Specific” in the table) learn the associated vocabulary more accurately for each dataset, while vocabulary-independent queries (“Vocabulary-Independent”) maintain open recognition ability (indicated by higher NIV scores). These results highlight the value of pattern query 224.

[0095] In addition, two discriminative baselines are established for comparison. The first baseline (denoted as “Learnable Embedding”) replaces the frozen LLM with six learnable linear layers, each of which is tailored to a specific dataset. The second baseline (named “Text Embedding”) initializes the classification layer with pre-extracted text embeddings and applies the classification layer to each dataset individually. As shown in Table 1, the average performance of the generative model OSM is comparable to that of the “Learnable Embedding” baseline (78.7% vs. 78.9% Acc) and outperforms the “Text Embedding” baseline (which has an Acc of 78.1%). Therefore, even in the discriminative task, the generative model achieves similar accuracy to the discriminative model, highlighting its versatility and effectiveness.

[0096] Finally, as shown in the last two rows of Table 1 (expressed as OSM ), using the final model settings (e.g., larger input size) can further significantly improve the performance in both the vocabulary-independent and vocabulary-specific settings.

[0097] Adaptation to higher input resolutions

[0098] In contrast to many multimodal LLM methods that directly adopt frozen CLIP-ViT, higher input resolutions can lead to more accurate object-level recognition. However, frozen ViT often exhibits poor performance when adapting to input resolutions larger than its pre-training resolution. To address this limitation, the sliding window approach discussed above is introduced to obtain enhanced features from frozen ViT when processing higher resolution inputs.

[0099] Figure 6 Tables 2A, 2B, 2C, and 2D show the results of ablation experiments on OSM. As shown in Table 2A, the experiments consistently show improved performance as the input resolution increases, especially when increasing from 224×224 to 448×448, reflecting an impressive improvement of +12.8% in Acc. This highlights the role of larger input resolution in achieving higher object-level recognition performance. This benefit persists until the input resolution reaches 1120×1120, while larger input resolutions lead to a drop in performance, probably because each sliding window fails to capture meaningful semantic features. It is worth noting that the “Average NIV” metric remains relatively stable across all experiments, indicating that the performance improvement mainly stems from improved mask classification rather than better overfitting to the corresponding vocabulary.

[0100] Sliding window stride

[0101] Table 2B validates the sliding window design, where directly applying frozen ViT (“global”) in the case of high-resolution inputs leads to a significant performance degradation (-7.2% Acc). In addition, Table 2B reflects that adopting a sliding window approach with overlapping windows further enhances the results, but the incremental gain decreases with increasing overlap. Given the significant additional computational cost associated with using overlapping windows, they were not used in the final setting.

[0102] Impact of Pattern Queries

[0103] Table 2C shows the effect of pattern query 224 on performance. As shown in Table 2C, training OSM 30 across multiple datasets without using pattern query 224 can produce better generalization capabilities, but at the expense of alignment with specific datasets. These effects can be seen from the lower "Average Acc" and higher "Average NIV". However, with the incorporation of pattern query, OSM 30 demonstrates the ability to operate in both "closed" mode (vocabulary specific) and "open" mode (vocabulary independent). This allows OSM 30 to strike a balance between generalization and alignment while retaining both of these essential capabilities.

[0104] Context is important for recognition

[0105] Table 2D shows the impact of context expansion on performance. Here, "global" means that the context attention covers the entire image, while "0.0×" refers to a tightly restricted bounding box that tightly surrounds the segmentation mask. The notation "k×" means that the bounding box is expanded by a factor of "k" on each side. The results in the table highlight the importance of context. Even tightly defined bounding boxes provide a significant improvement over global context (+0.8%). It is worth noting that transitioning to looser bounding boxes gradually improves the gains, with the largest gain (+2.6%) occurring at "0.5×" compared to the global context configuration.

[0106] 4.2. Mask Classification Using Pre-trained Mask Proposal Model

[0107] Benchmarking with other common models

[0108] In addition to evaluating OSM30 using baseline truth masks274, a practical evaluation is provided by integrating OSM30 with a pre-trained mask proposal model. Mask proposals generated by kMax-DeepLab are employed and OSM30 is applied to classify these mask proposals. Comparisons with other general segmentation models that are jointly trained using multiple segmentation datasets, similar to the current setting, are also provided. Specifically, text embedding based methods such as LMSeg and DaTaSeg are compared on various datasets including COCO Panoramic, ADE20K Panoramic and Semantic, Cityscapes Panoramic and Semantic Segmentation. The results of these comparisons are presented in Figure 7 As shown in Table 3 in . As shown in Table 3, OSM consistently achieves higher panorama quality (PQ) and mean intersection over union (mIoU) scores compared to the discriminative methods. Specifically, with the R-50 proposal model backbone, OSM 30 outperforms LMSeg by +14.7, +8.4, and +4.7 PQ on COCO, ADE20K, and Cityscapes, respectively. Compared to DaTaSeg, OSM also improves COCO PQ by +4.3, +2.6 for R50 and large backbone variants, and improves ADE20K mIoU by +1.9, +1.2. OSM 30 also shows comparable performance to the dedicated model Mask2Former.

[0109] 5. Conclusion

[0110] In the above discussion, the open-ended visual recognition task was introduced. To address this challenge, a generative framework, OSM, is proposed. OSM processes segmentation masks as input and generates semantic class predictions in a generative manner without restricting these semantic class predictions to a predefined vocabulary. Experiments using OSM show that this generative model produces guaranteed recognition accuracy and shows great potential for real-world applications, especially in dealing with new concepts beyond the scope of the predefined vocabulary.

[0111] Supplementary Materials

[0112] In the supplementary material, more technical details of OSM are provided. In addition, more visualizations and comparisons with related multimodal large models are included. Furthermore, the results show that OSM can be easily extended with part-level and box-level datasets, further unlocking the potential of OSM.

[0113] Instruction templates

[0114] In such Figure 8The instruction templates used for OSM training are summarized in Table 4. During training, an instruction template is randomly selected and the ground truth class name is inserted. Only the first template “What’s in the segmentation mask?” is used during testing.

[0115] The trade-off between accuracy and generalization

[0116] OSM is trained with different numbers of observed masks (i.e., 1 million, 3 million, 6 million, 9 million, respectively). Fig. 9 A plot 300 of NIV and Acc for OSM with these different numbers of training masks is shown. As a rule of thumb, Acc is considered a measure of how accurately the model recognizes an object. NIV is considered a measure of how well the model generalizes. Plot 300 shows a tradeoff between accuracy and generalization, such that as the number of observed masks increases, the model achieves higher accuracy while increasing overfitting to the training vocabulary and predicting in a more conservative manner. The improvement in accuracy from 6M to 9M is primarily due to a reduction in NIV.

[0117] Integrating part-level and box-level datasets

[0118] OSM seamlessly adapts to part-level and box-level datasets, further enhancing its versatility. To enhance OSM’s part-level and box-level recognition capabilities (note that, e.g. Figure 1 As shown in

[15] , OSM has shown the ability to recognize new parts, but the introduction of such datasets can further improve its part recognition ability), PartImageNet, Pascal-Part and V3Det datasets were introduced into the training data. For part data, the object name is prepended to the part name in case multiple parts share the same name (for example, in PartImageNet, many different classes may have the same part named "head"). In addition, overly ambiguous class names are removed (for example, left side of train in Pascal-Part, upper side of bus).

[0119] For detection data, bounding boxes are treated as box-shaped binary masks and are therefore easily unified into OSM. Additionally, panoptic / instance segmentation data (e.g., COCO, LVIS) are augmented by randomly converting each segmentation mask into its corresponding bounding box. When bounding boxes are given as input, the text instructions are appropriately adjusted by replacing the term “segmentation mask” with “bounding box”. Image-level data (e.g., ImageNet) are not included at this stage because semantic labels may introduce bias when multiple objects share a single label.

[0120] Fig.10 An example image labeled with part-level and box-level data is shown. Fig.10In the example of , SAM and DETA are used as the proposed models respectively. Fig.10 A first image 400 and a second image 410 are depicted, each of which is labeled using part-level data. A plurality of part-level labels 402 are shown within the first image 400 and the second image 410. In addition, Fig.10 A third image 420 and a fourth image 430 are depicted labeled with box-level data. A plurality of bounding boxes 422 and corresponding box-level labels 424 are depicted in the third image 420 and the fourth image 430.

[0121] Qualitative results

[0122] FIG. 11A to FIG. 11B Qualitative results are provided when OSM is used with SAM and kMax-DeepLab as segmenters, respectively. These results demonstrate the ability of OSM to perform open recognition in real-world scenarios when fine-grained masks are used. Fig.11A Example images 500 are shown, along with corresponding segmentation masks 502 and semantic labels 504 generated for these images 500 when using SAM as a segmenter. When obtaining mask proposals from SAM, the SAM variant has a ViT-H backbone, 32 points per side, an IoU threshold of 0.95, a stability threshold of 0.95, and a minimum mask size of 800. These settings avoid outputting a large number of unrecognizable small masks (e.g., superpixel-level masks).

[0123] Fig. 11B Example images 510 are shown, along with corresponding segmentation masks 512 and semantic labels 514 generated for these images 510 when using kMax-DeepLab as a segmenter. When obtaining mask proposals from kMax-DeepLab, a model trained on the COCO panoramic dataset with ConvNeXt-L as the backbone was used, and the "foreground things" and "background things" thresholds were set to 0.1. After OSM processes the mask proposals, mask-level post-processing is applied to the output of OSM.

[0124] Comparison with related multimodal large models

[0125] The following provides a qualitative comparison between relevant multimodal large models and OSM, such as Fig.12 As shown in the example of . The relevant multimodal large model is used for mask recognition. Fig.12 , the mask boundaries are highlighted in the example image 600 as auxiliary clues, and each mask center is annotated with a numeric ID 602. Together with the text hint "I have marked the center of each visual object in the image with a bright numeric ID. Please list their names (i.e., semantic classes) in one to three words.", each hint image 600 is fed into the relevant multimodal large model.

[0126] Fig.12 The results of semantic labeling using the relevant multimodal large model and OSM are further shown. Fig.12 , the first column of images 600 shows the image after performing mask hinting and inputting the masked image into the relevant multimodal large model. The second column shows the image 610 labeled by the relevant multimodal large model, where the semantic labels predicted by the relevant multimodal large model are depicted at the corresponding interest regions. The third column shows the OSM labeled image 620, where the semantic labels 622 calculated using OSM are shown at the interest regions.

[0127] from Fig.12 The results shown show that OSM's predictions are more accurate than the relevant multimodal large model (for example, in the first row, OSM correctly predicts masks 5 and 11 as "bench" and "fence", while the relevant multimodal large model incorrectly predicts both of them as "streetlight"), and the relevant multimodal large model is often confused by the context (for example, in the first row, for mask 10, the relevant multimodal large model predicts "building" instead of "mountain", which may be confused by the building below). However, compared with the relevant multimodal large model that can predict more specific words, OSM's predictions are more conservative. For example, in the second row, the relevant multimodal large model predicts the man in the image as "man in armor", while OSM predicts it as "person" in a safer way.

[0128] Fig.12 The method for generating labeled images shown in was also used to compare OSM with an open source multimodal LLM. However, the open source multimodal LLM failed to generate reasonable output.

[0129] Evaluation using the open vocabulary benchmark

[0130] like Fig.13As shown, OSM is also evaluated against state-of-the-art open vocabulary tagging methods in Table 5. To provide an open vocabulary setting (i.e., a setting where the target dataset has never been seen during training), OSM is trained using only COCO and LVIS data and evaluated on the ADE20K dataset in a zero-shot manner. During testing, OSM’s predictions are mapped to the target vocabulary using the text embedding similarity between the predicted class names and the target vocabulary class names. Geometric ensemble is also applied to enhance the tagging results predicted using frozen CLIP. Table 5 reports the results with and without geometric ensemble, where the results with geometric ensemble are indicated by asterisks. As shown in Table 5, when the geometric ensemble method is not used, OSM achieves higher PQ, AP, and mIoU scores compared to the state-of-the-art open vocabulary methods. When geometric ensemble is used against another frozen CLIP, OSM still shows comparable performance to other state-of-the-art methods such as ODISE and FC-CLIP.

[0131] Visualization of NIV situations

[0132] To demonstrate the NIV (not in vocabulary) situation discussed above, FIG. 14A to FIG. 14B The NIV situations are described in Fig.14A and Fig. 14B The ground truth masks are compared with the ground truth annotations for the COCO validation set and the ADE20K validation set. FIG. 14A to FIG. 14B In the example of FIG. 7 , an example image 700 is shown with the corresponding segmentation mask 702 and bounding box 704 highlighted.

[0133] Even ground truth annotations often have biases and limitations using predefined vocabularies, where annotators have to select the most similar class in a given vocabulary (e.g., in COCO, all monitors are labeled as “tv”). These biases may be learned and inherited in existing closed-vocabulary and open-vocabulary models. However, OSM can predict more appropriate class names without being restricted to a given vocabulary, demonstrating the effectiveness of moving away from predefined vocabularies and pursuing open-ended visual recognition.

[0134] Fig.15A A flow chart of a method 800 for assigning semantic tags to images for use with a computing system is schematically shown. For example, the method 800 may be performed in FIG. 3A to FIG. 3B The processing steps are executed at one or more processing devices 202 included in the computing system 200 .

[0135] At step 802, the method 800 includes receiving an image. At step 804, the method 800 also includes computing a segmentation mask identifying a region of interest included in the image. The segmentation mask may be computed on a pre-trained segmenter (eg, SAM or kMax-DeepLab).

[0136] In some examples, at step 806, method 800 further includes sampling a plurality of windows of the image. The windows may be sampled at the feature extractor. In such examples, the plurality of windows each have a window size that is less than the total size of the image. For example, each window having a size of 224×224 may be sampled from a higher resolution image. In some examples of performing step 806, overlapping windows are sampled from the image.

[0137] At step 808, method 800 also includes computing a plurality of encoded image features based at least in part on the image. The encoded image features may be pixel embeddings. Step 808 may also be performed at a feature extractor. In an example of performing step 806, the feature extractor is configured to receive a window as input. The feature extractor may be an open vocabulary classifier, such as CLIP-ViT.

[0138] At step 810 , the method 800 further includes receiving a text instruction, the text instruction including a plurality of text symbols.

[0139] At step 812, method 800 also includes computing a mask query. The mask query is computed based at least in part on the segmentation mask, a plurality of encoded image features, and a text instruction, the mask query comprising a plurality of text symbols. The mask query may be computed at a visual resampler. In some examples where the mask query is computed at a visual resampler, the input to the visual resampler may also include a pattern query and / or a context query. The visual resampler may have a transformer architecture comprising a plurality of transformer layers. In such an example, each transformer layer may include a self-attention layer, a mask cross-attention layer, a context cross-attention layer, a first feed-forward layer, and a second feed-forward layer.

[0140] Steps 814 and 816 of method 800 may be performed at a generative language model (which may be an LLM). At step 814, method 800 further includes receiving a natural language query including a mask query and a text instruction. In an example using a pattern query, the natural language query may also include a pattern query. At step 816, method 800 further includes generating a semantic label associated with the focus area based at least in part on the natural language query.

[0141] At step 818, method 800 also includes outputting the semantic tag. For example, the semantic tag can be output to a graphical user interface (GUI) that displays the semantic tag. In such an example, the image and the focus area can also be displayed on the GUI.

[0142] FIG. 15B to FIG. 15D Additional steps of method 800 are shown that may be performed in some examples. Fig. 15B The steps may be performed at the visual resampler. In step 820, Fig. 15B As shown, method 800 may also include receiving a pattern query indicating a vocabulary-specific mode. The vocabulary-specific mode is a vocabulary-specific mode or a vocabulary-independent mode, in which the pattern query includes multiple predefined classification labels, and in which the pattern query does not include predefined classification labels. The pattern query can be a text input including multiple text symbols. At step 822, method 800 may also include calculating a mask query based at least in part on the pattern query. Therefore, the visual resampler can switch between the vocabulary-specific mode and the vocabulary-independent mode.

[0143] Fig. 15C Additional steps of method 800 that may be performed when using context queries are shown. At step 824, method 800 may also include computing a context query. The context query is associated with a bounding box surrounding the region of interest. At step 826, method 800 may also include receiving the context query as input at a visual resampler. The context query may be iteratively recalculated on each of a plurality of layers of the visual resampler. Thus, image data in the region within the bounding box may be used to provide additional context used by the visual resampler to compute the mask query.

[0144] Fig.15D Additional steps of method 800 that may be performed when training a visual resampler prior to performing step 802 are shown. At step 828, method 800 may also include training the visual resampler using a training corpus, the training corpus comprising a plurality of training images. The training corpus may also include a plurality of ground truth masks associated with corresponding training regions of interest within the training images. Additionally, the training corpus may also include a plurality of ground truth labels associated with the ground truth masks.

[0145] In execution Fig. 15BIn an example of the steps of, training the visual resampler at step 828 may include performing steps 830, 832, and 834. In such an example, at step 830, step 828 may include computing the training corpus as a union of multiple training data subsets, the corresponding reference truth labels in the multiple training data subsets having different corresponding label spaces. In step 832, step 828 may also include computing multiple pattern queries, the pattern queries being respectively associated with the training data subsets and indicating corresponding label spaces. In step 834, step 828 may also include receiving the pattern query as input at the visual resampler during training.

[0146] Additionally or alternatively to steps 830, 832, and 834, in some examples, step 828 may include step 836. In step 828, step 836 may include training the visual resampler via instruction fine-tuning.

[0147] In some embodiments, the methods and processes described herein may be associated with a computing system of one or more computing devices. Specifically, such methods and processes may be implemented as a computer application or service, an application programming interface (API), a library, and / or other computer program products.

[0148] Fig.16 A non-limiting embodiment of a computing system 900 is schematically shown that can implement one or more of the above methods and processes. The computing system 900 is shown in simplified form. The computing system 900 can embody the above description and FIG. 3A to FIG. 3B 900 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).

[0149] The computing system 900 includes a logic processor 902, a volatile memory 904, and a non-volatile storage device 906. The computing system 900 may optionally include a display subsystem 908, an input subsystem 910, a communication subsystem 912, and / or Fig.16 Other components not shown.

[0150] Logical processor 902 includes one or more physical devices configured to execute instructions. For example, a logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical structures. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.

[0151] The logical processor may include one or more physical processors configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processor of the logical processor 902 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel and / or distributed processing. The individual components of the logical processor may optionally be distributed in two or more separate devices, which may be located remotely and / or configured to coordinate processing. Various aspects of the logical processor may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration. It will be understood that in this case, these virtualized aspects run on different physical logical processors of various different machines.

[0152] The non-volatile storage device 906 includes one or more physical devices that are configured to store instructions executable by a logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 906 can be transformed, for example, to store different data.

[0153] The non-volatile storage device 906 may include a removable and / or built-in physical device. The non-volatile storage device 906 may include an optical memory, a semiconductor memory, and / or a magnetic memory, or include other mass storage device technologies. The non-volatile storage device 906 may include a non-volatile, dynamic, static, read / write, read-only, sequential access, location addressable, file addressable, and / or content addressable device. It should be understood that the non-volatile storage device 906 is configured to save instructions even when the power to the non-volatile storage device 906 is cut off.

[0154] The volatile memory 904 may include a physical device that includes random access memory. The volatile memory 904 is typically used by the logical processor 902 to temporarily store information during the processing of software instructions. It should be understood that when the power to the volatile memory 904 is cut off, the volatile memory 904 will generally not continue to store instructions.

[0155] Aspects of the logic processor 902, volatile memory 904, and non-volatile storage device 906 may be integrated together into one or more hardware logic components. Such hardware logic components may include field programmable gate arrays (FPGAs), program and application specific integrated circuits (PASIC / ASIC), program and application specific standard products (PSSP / ASSP), systems on chips (SOCs), and complex programmable logic devices (CPLDs), among others.

[0156] The terms "module", "program", and "engine" may be used to describe aspects of the computing system 900 that are typically implemented in software by a processor to perform a specific function using portions of volatile memory, which involves a transformation process that specifically configures the processor to perform the function. Thus, a module, program, or engine may be instantiated by executing instructions stored by a non-volatile storage device 906 via a logical processor 902 using portions of volatile memory 904. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, and the like. Likewise, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, and the like. The terms "module", "program", and "engine" may cover individual or multiple groups of executable files, data files, libraries, drivers, scripts, database records, and the like.

[0157] The display subsystem 908, when included, can be used to present a visual representation of the data stored by the non-volatile storage device 906. The visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data stored by the non-volatile storage device, and thus transform the state of the non-volatile storage device, the state of the display subsystem 908 can also be similarly transformed to visually represent the changes in the underlying data. The display subsystem 908 may include one or more display devices utilizing almost any type of technology. Such a display device can be combined with the logical processor 902, the volatile memory 904, and / or the non-volatile storage device 906 in a shared housing, or such a display device can be a peripheral display device.

[0158] The input subsystem 910, when included, may include or interface with one or more user input devices, such as a keyboard, mouse, touch screen, or game controller. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. These components may be integrated components or peripheral components, and the transduction and / or processing of input actions may be performed on-board or off-board. Example NUI components may include microphones for speech and / or voice recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition; and electric field sensing components for assessing brain activity; and / or any other suitable sensors.

[0159] The communication subsystem 912, when included, can be configured to communicatively couple the various computing devices described herein to each other, as well as to communicatively couple to other devices. The communication subsystem 912 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network, or a wired or wireless local area network or wide area network. In some embodiments, the communication subsystem can allow the computing system 900 to send messages to other devices and / or receive messages from other devices via networks such as the Internet.

[0160] The following paragraphs provide additional descriptions of the subject matter of the present disclosure. According to one aspect of the present disclosure, a computing system is provided, the computing system including one or more processing devices, the one or more processing devices being configured to receive an image. The one or more processing devices are also configured to calculate a segmentation mask that identifies a region of interest included in the image. At the feature extractor, the one or more processing devices are also configured to calculate a plurality of encoded image features based at least in part on the image. The one or more processing devices are also configured to receive text instructions. At the visual resampler, the one or more processing devices are also configured to calculate a mask query based at least in part on the segmentation mask, the plurality of encoded image features, and the text instructions, the mask query comprising a plurality of text symbols. At the generative language model, the one or more processing devices are also configured to receive a natural language query comprising a mask query and a text instruction. At the generative language model, the one or more processing devices are also configured to generate a semantic label associated with the region of interest based at least in part on the natural language query. The one or more processing devices are also configured to output the semantic label. The above features may have the following technical effect: regions of interest in an image are labeled in a generative manner, thereby allowing for increased diversity among labels while also maintaining a consistent level of label specificity.

[0161] According to this aspect, the visual resampler may be further configured to receive a pattern query indicating a vocabulary specificity pattern. The visual resampler may be further configured to compute a mask query based at least in part on the pattern query. The above features may have the following technical effects: giving the visual resampler a runtime customizable vocabulary specificity level.

[0162] According to this aspect, the natural language query may also include a pattern query.The above features may have the technical effect of providing a level of vocabulary specificity to a generative language model when generating semantic tags.

[0163] According to this aspect, the vocabulary specific mode can be a vocabulary specific mode, in which the pattern query includes a plurality of predefined classification labels, or a vocabulary independent mode, in which the pattern query does not include a predefined classification label. The above features can have the following technical effects: allowing a user to define a set of predefined classification labels from which semantic labels are selected, or alternatively not restricting semantic labels to members of a predefined set.

[0164] According to this aspect, the one or more processing devices can be configured to compute the encoded image features at least in part by sampling multiple windows of the image. The multiple windows can each have a window size that is smaller than the total size of the image. The above features can have the following technical effects: even if the size of the image is different from the input size of the pre-trained feature extractor, the pre-trained feature extractor can be used.

[0165] According to this aspect, the one or more processing devices may also be configured to compute a context query associated with a bounding box surrounding the region of interest. The visual resampler may also be configured to receive the context query as input. The above features may have the following technical effects: when the visual resampler computes the mask query, context information from the image region surrounding the region of interest is incorporated.

[0166] According to this aspect, the visual resampler can have a transformer architecture, the transformer architecture comprising a plurality of transformer layers, wherein each transformer layer comprises a self-attention layer, a mask cross-attention layer, a context cross-attention layer, a first feed-forward layer, and a second feed-forward layer. The above features can have the following technical effects: allowing the visual resampler to generalize the attention region while retaining access to the context content.

[0167] According to this aspect, a visual resampler may be trained using a training corpus, the training corpus comprising a plurality of training images. The training corpus may also include a plurality of ground truth masks associated with corresponding training regions of interest within the training images. The training corpus may also include a plurality of ground truth labels associated with the ground truth masks. The above features may have the following technical effects: the visual resampler is trained to generate mask queries in a manner that matches the distribution of the training corpus.

[0168] According to this aspect, the visual resampler may be trained via instruction fine-tuning.The above features may have the following technical effect: facilitating the integration of the visual resampler with a generative language model during training.

[0169] According to this aspect, the training corpus is a union of multiple training data subsets, and the corresponding reference truth labels in the multiple training data subsets have different corresponding label spaces. The one or more processing devices can also be configured to calculate multiple pattern queries, which are respectively associated with the training data subsets and indicate the corresponding label spaces. The visual resampler can also be configured to receive pattern queries during training. The above features can have the following technical effects: the training visual resampler recognizes multiple different sets of predefined classification labels when used in a vocabulary-specific mode at runtime.

[0170] According to another aspect of the present disclosure, a method for image processing is provided. The method includes receiving an image and calculating a segmentation mask, which identifies a region of interest included in the image. The method also includes calculating a plurality of encoded image features based at least in part on the image. The method also includes receiving a text instruction. The method also includes calculating a mask query based at least in part on the segmentation mask, the plurality of encoded image features and the text instruction, the mask query including a plurality of text symbols. The method also includes receiving a natural language query including the mask query and the text instruction. The method also includes generating a semantic label associated with the region of interest based at least in part on the natural language query. The method also includes outputting the semantic label. The above features can have the following technical effects: marking the region of interest in the image in a generative manner, thereby allowing for an increase in diversity among the labels while also maintaining consistency in the level of label specificity.

[0171] According to this aspect, the method may further include receiving a pattern query indicating a vocabulary specificity pattern. The method may further include computing a mask query based at least in part on the pattern query. The above features may have the following technical effect: giving the visual resampler a runtime customizable vocabulary specificity level.

[0172] According to this aspect, the natural language query may also include a pattern query. The vocabulary-specific pattern may be a vocabulary-specific pattern, in which the pattern query includes multiple predefined classification labels, or a vocabulary-independent pattern, in which the pattern query does not include predefined classification labels. The above features may have the following technical effects: providing a degree of vocabulary specificity to a generative language model when generating semantic labels. The above features may also have the following technical effects: allowing a user to define a set of predefined classification labels from which to select semantic labels, or alternatively not restricting semantic labels to members of a predefined set.

[0173] According to this aspect, calculating the encoded image features may include sampling multiple windows of the image. The multiple windows may each have a window size that is smaller than the total size of the image. The above features may have the following technical effects: even if the size of the image is different from the input size of the pre-trained feature extractor, the pre-trained feature extractor can be used.

[0174] According to this aspect, the method may further include computing a context query associated with a bounding box surrounding the region of interest. The method may further include receiving the context query as input. The above features may have the following technical effect: when the visual resampler computes the mask query, context information from the image region surrounding the region of interest is combined.

[0175] According to this aspect, a mask query may be computed at a visual resampler. The visual resampler may have a transformer architecture comprising a plurality of transformer layers, wherein each transformer layer comprises a self-attention layer, a mask cross-attention layer, a context cross-attention layer, a first feed-forward layer, and a second feed-forward layer. The above features may have the following technical effects: allowing the visual resampler to generalize a region of interest while retaining access to contextual content.

[0176] According to this aspect, the method may further include training the visual resampler using a training corpus, the training corpus comprising a plurality of training images, a plurality of ground truth masks associated with respective training regions of interest within the training images, and a plurality of ground truth labels associated with the ground truth masks. The above features may have the following technical effects: the visual resampler is trained to generate mask queries in a manner that matches the distribution of the training corpus.

[0177] According to this aspect, the visual resampler can be trained via instruction fine-tuning. The above features can have the following technical effects: there is a facilitation of the integration of the visual resampler with the generative language model during training.

[0178] According to this aspect, the training corpus can be a union of multiple training data subsets, and the corresponding reference truth labels in the multiple training data subsets have different corresponding label spaces. The method may also include calculating multiple pattern queries, the pattern queries are respectively associated with the training data subsets and indicate the corresponding label spaces. The method may also include receiving the pattern queries at the visual resampler during training. The above features can have the following technical effects: the training visual resampler recognizes multiple different sets of predefined classification labels when used in a vocabulary-specific mode at runtime.

[0179] According to another aspect of the present disclosure, a computing system is provided, the computing system including one or more processing devices, the one or more processing devices being configured to receive an image. The one or more processing devices are also configured to calculate a segmentation mask identifying a region of interest included in the image. The one or more processing devices are also configured to calculate a plurality of encoded image features based at least in part on the image. The one or more processing devices are also configured to calculate a context query associated with a bounding box surrounding the region of interest. The one or more processing devices are also configured to receive a text instruction. The one or more processing devices are also configured to receive a pattern query indicating a vocabulary-specific pattern. The one or more processing devices are also configured to calculate a mask query based at least in part on the segmentation mask, the plurality of encoded image features, the context query, the text instruction, and the pattern query, the mask query comprising a plurality of text symbols. The one or more processing devices are also configured to receive a natural language query including the mask query and the text instruction. The one or more processing devices are also configured to generate a semantic tag associated with the region of interest based at least in part on the natural language query. The one or more processing devices are also configured to output the semantic tag. The above features may have the following technical effect: regions of interest in an image are labeled in a generative manner, thereby allowing for increased diversity among labels while also maintaining a consistent level of label specificity.

[0180] As used herein, "and / or" is defined as an inclusive OR, as specified by the following truth table:

[0181] A B A∨B real real real real Fake real Fake real real Fake Fake Fake

[0182] It should be understood that the configuration and / or method described herein is exemplary in nature, and these specific embodiments or examples should not be considered restrictive because there may be multiple variations. The specific routines or methods described herein can represent one or more of any number of processing strategies. Therefore, the various actions illustrated and / or described can be performed in the illustrated and / or described order, in other orders, in parallel, or omitted. Similarly, the order of the above process can also be changed.

[0183] The subject matter of the present disclosure includes all novel and nonobvious combinations and subcombinations of the various processes, systems and configurations, and other features, functions, acts, and / or properties disclosed herein, as well as any and all equivalents thereof.

Claims

1. A computing system comprising: One or more processing devices, the one or more processing devices being configured to: receiving an image; calculating a segmentation mask, the segmentation mask identifying a region of interest included in the image; At a feature extractor, computing a plurality of encoded image features based at least in part on the image; Receive text instructions; At a visual resampler, computing a mask query based at least in part on the segmentation mask, the plurality of encoded image features, and the textual instructions, the mask query comprising a plurality of textual symbols; At the generative language model: receiving a natural language query including the mask query and the text instruction; as well as generating a semantic tag associated with the region of interest based at least in part on the natural language query; and The semantic label is output.

2. The computing system of claim 1, wherein the visual resampler is further configured to: receiving a pattern query indicating a vocabulary-specific pattern; and The mask query is computed based at least in part on the pattern query.

3. The computing system of claim 2, wherein the natural language query further comprises the pattern query.

4. The computing system of claim 3, wherein the vocabulary-specific pattern is: A vocabulary-specific pattern, wherein the pattern query includes a plurality of predefined classification tags; or A vocabulary-independent mode, where the mode query does not include predefined classification labels.

5. The computing system of claim 1, wherein: The one or more processing devices are configured to compute the encoded image features at least in part by sampling a plurality of windows of the image; and Each of the plurality of windows has a window size that is smaller than a total size of the image.

6. The computing system of claim 1, wherein: The one or more processing devices are further configured to compute a context query associated with a bounding box surrounding the region of interest; and The visual resampler is further configured to receive the context query as input.

7. The computing system of claim 1, wherein the visual resampler has a transformer architecture, the transformer architecture comprising a plurality of transformer layers, each transformer layer of the plurality of transformer layers comprising: Self-attention layer; Masked cross attention layer; Contextual cross-attention layer; The first feed-forward layer; as well as The second feed-forward layer.

8. The computing system of claim 1, wherein the visual resampler is trained using a training corpus, the training corpus comprising: Multiple training images; a plurality of reference truth masks, the plurality of reference truth masks being associated with respective training regions of interest within the training image; as well as A plurality of ground truth labels are associated with the ground truth mask.

9. The computing system of claim 8, wherein the vision sampler is trained via instruction fine-tuning.

10. The computing system of claim 8, wherein: The training corpus is a union of a plurality of training data subsets, and corresponding reference truth labels in the plurality of training data subsets have different corresponding label spaces; The one or more processing devices are further configured to compute a plurality of pattern queries, the pattern queries being respectively associated with the training data subsets and indicating the corresponding label spaces; and The visual resampler is also configured to receive the pattern query during training.

11. A method for image processing, the method comprising: receiving an image; calculating a segmentation mask, the segmentation mask identifying a region of interest included in the image; computing a plurality of encoded image features based at least in part on the image; Receive text instructions; computing a mask query based at least in part on the segmentation mask, the plurality of encoded image features, and the textual instructions, the mask query comprising a plurality of textual symbols; receiving a natural language query including the mask query and the text instruction; as well as generating a semantic tag associated with the region of interest based at least in part on the natural language query; and The semantic label is output.

12. The method of claim 11, further comprising: receiving a pattern query indicating a lexically specific pattern; as well as The mask query is computed based at least in part on the pattern query.

13. The method of claim 12, wherein: The natural language query also includes the pattern query; and The vocabulary-specific patterns are: A vocabulary-specific pattern, wherein the pattern query includes a plurality of predefined classification tags; or A vocabulary-independent mode, where the mode query does not include predefined classification labels.

14. The method of claim 11, wherein: Computing the encoded image features includes sampling a plurality of windows of the image; and Each of the plurality of windows has a window size that is smaller than a total size of the image.

15. The method of claim 11, further comprising: computing a context query associated with a bounding box surrounding the region of interest; as well as The context query is received as input.

16. The method of claim 11, wherein: The mask query is computed at a visual resampler; and The visual resampler has a transformer architecture, the transformer architecture comprising a plurality of transformer layers, each transformer layer of the plurality of transformer layers comprising: Self-attention layer; Masked cross attention layer; Contextual cross-attention layer; a first feed-forward layer; and The second feed-forward layer.

17. The method of claim 16, further comprising training the visual resampler using a training corpus, the training corpus comprising: Multiple training images; a plurality of reference truth masks associated with respective training regions of interest within the training image; as well as A plurality of ground truth labels are associated with the ground truth mask. The method of claim 17 , wherein the visual resampler is trained via instruction fine-tuning.

19. The method of claim 17, wherein: The training corpus is a union of a plurality of training data subsets, and corresponding reference truth labels in the plurality of training data subsets have different corresponding label spaces; and The method further comprises: computing a plurality of pattern queries, the pattern queries being respectively associated with the training data subsets and indicating the corresponding label spaces; as well as At the visual resampler, the pattern query is received during training.

20. A computing system comprising: One or more processing devices, the one or more processing devices being configured to: receiving an image; calculating a segmentation mask, the segmentation mask identifying a region of interest included in the image; computing a plurality of encoded image features based at least in part on the image; computing a context query associated with a bounding box surrounding the region of interest; Receive text instructions; receiving a pattern query indicating a lexically specific pattern; computing a mask query based at least in part on the segmentation mask, the plurality of encoded image features, the context query, the text instruction, and the pattern query, the mask query comprising a plurality of text symbols; receiving a natural language query including the mask query and the text instruction; generating a semantic tag associated with the region of interest based at least in part on the natural language query; as well as The semantic label is output.

Citation Information

Cited By

  • Landslide disaster early warning method and device based on deep learning, medium and product

    CN120299220A