Multi-criteria dense system with open vocabulary and methods for multi-criteria recording of images with open vocabulary

The multi-criteria dense system with a summarizing CLIP head and pseudo-labels improves image retrieval accuracy by maintaining open vocabulary capabilities and utilizing unlabeled data, addressing the limitations of existing techniques.

DE102024117300B4Active Publication Date: 2026-02-12GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024117300
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2024-06-19
Publication Date
2026-02-12
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

Existing fixed-pretrained open vocabulary techniques perform poorly on target datasets and suffer from slow, complex two-step embedding processes, while supervised techniques lose open vocabulary capabilities when new categories are introduced.

Method used

A multi-criteria dense system with an open vocabulary using a summarizing CLIP head trained on supervised and pseudo-label losses, coupled with a classifier to identify target images, and leveraging unlabeled data for improved accuracy.

Benefits of technology

Enhances image retrieval accuracy for both trained and untrained categories with faster inference times and preserved open vocabulary capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Multi-criteria dense system with an open vocabulary, which features the following: an image encoder (326) with a summarizing contrast speech image pretraining head (summarizing CLIP head, 230) that can be connected to an originating device, wherein: the summary CLIP head (230) is trained using monitored losses from unlabeled image data (206a) and labeled image data (206b); The summary CLIP header (230) loses the ability to use an open vocabulary as its capacity increases; the summary CLIP head (230) is trained on pseudo-label losses from a variety of pseudo-labels (228) that compensate for the loss of open vocabulary skills; the multitude of pseudo-labels (228) is generated from a multitude of text embeddings (218) based on similarities with a multitude of average semantics, which are generated by a dense CLIP head (210); and The summary clip header (230) is ready for use at: Receiving a large number of captured images (325) from the originating device; and Generating a multitude of image embeddings based on the multitude of captured images (325); and a classifier (328) coupled to the image encoder (326), which can be coupled to a text encoder (102) and which can be coupled to a targeting device (330a-330b), wherein the classifier (328) is ready for operation to: Receiving one or more destinations from the text encoder (102); Receiving the multitude of image embeddings from the summary CLIP head (230); Classifying the multitude of image embeddings to identify one or more output images (338) that contain the one or more targets; and Presenting one or more output images (338) to the targeting device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present description refers to a system and a method for multi-criteria dense image acquisition with an open vocabulary.

[0002] Existing fixed-pretrained open vocabulary techniques perform worse on target datasets on which the technique was not trained. These fixed-pretrained open vocabulary techniques also employ two-step embedding processes, which tend to be slow and complex. Existing supervised open vocabulary techniques outperform fixed-pretrained open vocabulary techniques in annotated categories and can employ a single-step embedding process. However, supervised open vocabulary techniques tend to deviate from the original training when new categories are introduced as learned capacity grows.

[0003] Peize Sun et al. describe in “Going denser with open-vocabulary part segmentation, IEEE Xplore, 2023, 15453 - 15465” possibilities for equipping object detectors with the fine-grained detection function of open vocabulary segmentation.

[0004] Size Wu et al. describe in “CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction”, arXiv, 2024, 1-20, an analysis of a density representation of ViT-based CLIP models.

[0005] Sheng Liu et al. describe in “OVIS: Open-Vocabulary Visual Instance Search via Visial-Semantic Aligned Representation Learning”, AAAI-22, 2022, 1773 - 1781, a visually-semantically aligned representation learning method for an open-vocabulary circumstance search.

[0006] Hila Levi et al. describe in “Object-Centric Open-Vocabulary Image Retrieval with Aggregated Features”, arXiv, 2023, the use of local features for an object-centric image query task, so that a CLIP view-language association is preserved.

[0007] Zafran Khan et al. describe a multi-model-based image retrieval system in “DenseBert4Ret: Deep bi.modal for image retrieval”, Information Science, 2022, Vol 612, 1171 - 1186.

[0008] The task can be considered to be to specify an improved multi-criteria dense system with an open vocabulary and a method for image recording.

[0009] The problem is solved by a multi-criteria dense system with an open vocabulary according to claim 1 and a method for multi-criteria recording of images with an open vocabulary according to claim 10. Furthermore, a vehicle is described.

[0010] A system with a dense, open vocabulary and multiple objectives is provided. The system includes an image encoder and a classifier. The image encoder has a summarizing contrastive speech image pre-training head (CLIP) and is coupled to an originating device. The summarizing CLIP head is trained using supervised losses from unlabeled and labeled image data. The summarizing CLIP head loses open vocabulary capabilities as its capacity increases. The summarizing CLIP head is trained on pseudo-label losses from a variety of pseudo-labels that compensate for the loss of open vocabulary capabilities. The variety of pseudo-labels is generated from a variety of text embeddings based on similarities to a variety of average semantics created by a dense CLIP head.The summarizing CLIP head is capable of receiving a multitude of captured images from the source device and generating a multitude of image embeddings based on these images. The classifier is coupled to the image encoder, can be coupled to a text encoder, and can be coupled to a target device. The classifier is capable of receiving one or more targets from the text encoder, receiving the multitude of image embeddings from the summarizing CLIP, classifying the multitude of image embeddings to identify one or more output images containing the one or more targets, and presenting the one or more output images to the target device.

[0011] In one or more embodiments of the system, the summarizing CLIP head includes a backbone that serves to extract a variety of finely tuned features from the unlabeled image data and the labeled image data.

[0012] In one or more embodiments of the system, the summarizing CLIP head further includes a recognition transformer decoding layer that predicts a multitude of objects based on a multitude of finely tuned features and a multitude of adaptive queries.

[0013] In one or more embodiments of the system, the summarizing CLIP head also includes a multi-head attention layer that generates the multitude of image embeddings in response to the multitude of finely tuned features and the multitude of objects.

[0014] In one or more embodiments of the system, the dense CLIP head includes a backbone that serves to extract a variety of fixed features from the unlabeled image data and the labeled image data.

[0015] In one or more embodiments of the system, the dense CLIP head also includes a clustering module that serves to cluster the multiple fixed features in order to generate the multiple average semantics.

[0016] In one or more embodiments of the system, the dense CLIP head also includes an embedding system that generates the multiple pseudo-labels from the multiple text embeddings based on the multiple average semantics.

[0017] In one or more embodiments of the system, the original device is a camera that is operated in such a way as to generate the multitude of captured images.

[0018] In one or more embodiments of the system, the targeting device is a memory used to record one or more output images.

[0019] In one or more embodiments of the system, the targeting device is a display device that serves to optically display one or more output images.

[0020] A method for multi-criteria open-vocabulary image capture is presented here. The method involves receiving a multitude of captured images from a source device at an image encoder. The image encoder has a summarizing head for contrastive speech image pretraining (CLIP). The summarizing CLIP head is trained using supervised loss from unlabeled and labeled image data. The summarizing CLIP head loses open-vocabulary capabilities as its capacity increases. The summarizing CLIP head is trained on pseudo-label loss from a multitude of pseudo-labels, compensating for the open-vocabulary capability loss. The multitude of pseudo-labels is generated from a multitude of text embeddings based on similarities to a multitude of average semantics produced by a dense CLIP head.The procedure comprises generating a multitude of image embeddings with the summarizing CLIP head based on the multitude of captured images, receiving one or more targets from a text encoder at a classifier, receiving the multitude of image embeddings from the summarizing CLIP head at the classifier, classifying the multitude of image embeddings to identify one or more output images containing the one or more targets, and presenting the one or more output images to a target device.

[0021] In one or more embodiments, the method comprises extracting a multitude of finely tuned features from the unlabeled image data and the labeled image data using the summarization CLIP head.

[0022] In one or more embodiments, the method comprises predicting a multitude of objects based on a multitude of finely tuned features and a multitude of machine learning queries using the summary CLIP header.

[0023] In one or more embodiments, the method comprises generating the multitude of image embeddings in response to the multitude of finely tuned features and the multitude of objects using the summarizing CLIP head.

[0024] In one or more embodiments, the method comprises extracting a plurality of fixed features from the unlabeled image data and the labeled image data using the dense CLIP head.

[0025] In one or more embodiments, the method includes clustering the multiple fixed features to generate the multiple average semantics using the dense CLIP head.

[0026] In one or more embodiments, the method comprises generating the plurality of pseudo-labels from the plurality of text embeddings based on the plurality of average semantics using the dense CLIP head.

[0027] In one or more embodiments, the method comprises generating the multiple recorded images with a camera.

[0028] In one or more embodiments, the method comprises recording the one or more output images and optically displaying the one or more output images.

[0029] Furthermore, a vehicle is described. The vehicle comprises a camera, a CLIP (Contrastive Language Image Pre-Training) text encoder, a targeting device, and a dense, open-vocabulary, multi-target system. The camera is in operation to generate a multitude of captured images. The CLIP text encoder is capable of generating one or more targets. The targeting device is used to (i) record one or more output images and (ii) optically display the one or more output images. The dense, open-vocabulary, multi-target system has an image encoder and a classifier. The image encoder has a summarizing CLIP head and is coupled to the camera. The summarizing CLIP head is trained based on supervised loss from unlabeled and labeled image data. The summarizing CLIP head loses open-vocabulary capabilities as its capacity increases.The summary CLIP head is trained on pseudo-label loss from a multitude of pseudo-labels, compensating for the loss of open vocabulary capabilities. This multitude of pseudo-labels is generated from a multitude of text embeddings based on similarities with a multitude of average semantics, produced by a dense CLIP head. The dense CLIP head receives the multitude of images captured by the camera and generates a multitude of image embeddings based on this multitude of captured images. The classifier is connected to the image encoder, the CLIP text encoder, and the targeting device.The classifier is able to receive the one or more targets from the CLIP text encoder, receive the multitude of image embeddings from the summarizing CLIP head, classify the multitude of image embeddings to identify the one or more output images containing the one or more targets, and present the one or more output images to the target device. Fig. Figure 1 is a schematic diagram of the text processing in accordance with one or more exemplary embodiments. Fig. Figure 2 is a schematic diagram of a training framework of the system in accordance with one or more exemplary embodiments. Fig. Figure 3 is a schematic diagram of a dense CLIP (Dense-Contrastive Language-Image Pre-Training) head in accordance with one or more exemplary embodiments. Fig. Figure 4 is a schematic representation of a summary CLIP head in accordance with one or more exemplary embodiments. Fig. Figure 5 is a schematic diagram of a semi-supervised setup in accordance with one or more exemplary embodiments. Fig. Figure 6 is a schematic diagram of an online-triggered recording system in accordance with one or more exemplary embodiments. Fig. Figure 7 is an image of an initial image with detected targets in accordance with one or more exemplary embodiments. Fig. Figure 8 is a schematic representation of a control unit in accordance with one or more exemplary embodiments. Fig. Figure 9 is a schematic diagram of a first stage of an offline system in accordance with one or more exemplary embodiments. Fig. Figure 10 is a schematic diagram of a second stage of the offline system in accordance with one or more exemplary embodiments.

[0030] Embodiments of the description provide a system and / or method for multi-criteria dense open-vocabulary image retrieval. Dense-open-vocabulary image retrieval (D-OVIR) systems are commonly used in a wide range of applications and enable dense text querying. In various embodiments, the system / method includes both a fixed pre-trained technique and a supervised fine-tuning technique. The fixed pre-trained technique uses a pre-trained open-vocabulary head that maintains an original image-language association between images and text. The supervised fine-tuning technique is directly optimized for searching within the categories of the target dataset but tends to lose the capabilities of the open vocabulary as capacity increases.Therefore, the fine-tuning scheme of the supervised method is extended with auxiliary targets from the fixed scheme, enabling learning without forgetting the open vocabulary, which can be further improved by using unlabeled data. The combination of both schemes leads to improved query results in a target dataset for both trained and untrained categories.

[0031] Fig. Figure 1 shows a schematic diagram of an exemplary implementation of the text processing system in accordance with one or more exemplary embodiments. The text processing system 100 generally includes a text encoder 102. In various embodiments, the text encoder may implement a CLIP (Contrastive Language-Image Pre-Training) text encoder 102. The CLIP text encoder 102 is used to generate embeddings for text words and / or strings in various category lists 103. The category lists 103 may contain a basic category list 104, a new category list 106, and a pseudo-category list 108. The basic category list 104 contains basic terms such as chair 104a, bird 104b, bicycle 104c, etc. The new category list 106 generally contains other items such as scissors 106a, cake 106b, cow 106c, and so on.The CLIP text encoder 102 generates and presents the embeddings in corresponding groups of target embeddings 110. The target embeddings 110 can contain basic text codings 112, novel text codings 114, and pseudo-codings 116. The basic text codings 112 include embeddings for training, basic truth validation, and baseline assessment. The novel text codings 114 can include embeddings for validation and novel assessments. The pseudo-codings 116 can include embeddings for training and pseudo-labels.

[0032] In Fig. 2 is with reference to Fig. Figure 1 shows a schematic diagram of an exemplary training framework of a system in accordance with one or more exemplary embodiments. The training framework 170 generally comprises an unattended open vocabulary system 180, a supervised open vocabulary system 190, a visual CLIP backbone 202, and category lists (e.g., the basic category lists 104). The visual CLIP backbone 202 can receive images as unlabeled image data 206a and labeled image data 206b.

[0033] The unsupervised open vocabulary system 180 implements a pretrained open vocabulary model system. In some embodiments, the unsupervised open vocabulary system 180 comprises a dense CLIP head 210, a cluster module 212, and an embedding system 214. The embedding system 214 generally comprises first image embeddings 216, a first text embedding 218, and an embedding space 224.

[0034] The supervised open vocabulary system 190 implements a finely tuned open vocabulary model system. In various embodiments, the supervised open vocabulary system 190 includes a summary (SUM) clip head 230. The supervised open vocabulary system 190 trains the SUM clip head 230 based on the unlabeled image data 206a, the labeled image data 206b, and the pseudo-labels 228 received from the unsupervised open vocabulary system 180. The tuning generally helps to compensate for the loss of open vocabulary capabilities as the model's capacity grows.

[0035] The training framework 170 is based on a pre-trained dual-encoder vision-language model with separate processing pipelines for vision and text. Text processing is performed using the pre-trained CLIP text encoder 102 ( Fig. 1) on the three text lists: the list of basic categories 104, the list of novel categories 106 (both derived from the annotation space of the target dataset) and the list of pseudocategories 108.

[0036] Implementations of the description aim to retrieve images containing objects from the list of new categories (106) that extend beyond the list of base categories (104) on which an image coding model is trained. For a target dataset, the image coding model is trained on a training evaluation where both the list of base categories 104 (CB) and the list of new categories 106 (CN) are not seen by the training (e.g., intersection of lists CB ∩ CN = Ø). The training framework 170 is based on a pre-trained dual-encoder view-language model with separate processing pipelines for view and text. The text processing (represented in Fig. 1) is achieved by applying the pre-trained CLIP text encoder 102 to the three category lists 103 (e.g., 104, 106, and 108) of text. The visual processing comprises a frozen residual neural network (ResNet) 202, followed by two heads: (i) the dense CLIP head 210 in the unsupervised open vocabulary system 180 and (ii) the SUM CLIP head 230 in the supervised open vocabulary system 190, arranged in parallel streams.

[0037] The training follows a semi-supervised paradigm, in which the trainable SUM-CLIP head 230 is instructed both by a supervised loss 252 and by the outputs of the dense CLIP head 210 via a pseudo-label loss. For an input image with adjacent basic categories, processing proceeds through an initial execution by the visual CLIP backbone 202.

[0038] The generated intermediate feature maps are then processed by the dense CLIP head 210, which generates pseudo-labels 228, and by the SUM CLIP head 230, which summarizes the image content and generates multiple (e.g., N) second image embeddings 234. During training, the second image embeddings 234 are monitored by comparison with two sets of speech embeddings (e.g., CB and CP) using prediction losses. Positive results can be defined by image labels (monitored loss) and by the pseudo-labels 228 (unmonitored loss) generated by the dense CLIP branch. The results can be further improved by using unlabeled data from the target dataset and focusing solely on the unmonitored loss. Experiments show that effective results can be achieved even when only a small portion of the data is labeled.

[0039] During inference, an image processing encoder, comprising the visual CLIP backbone 202 and the finely tuned SUM-CLIP head 230, is applied to each image in the dataset, generating a set of embeddings for each image. Evaluation is performed by ranking the cosine similarity between the text embedding of each category in a union CB ∪ CN and the secondary image embeddings 234 generated by the SUM-CLIP head 230.

[0040] In a CLIP head, the final CLIP layer is implemented as a pooling multi-head attention layer, where the query is pooled by averaging from the input tensor itself. The CLIP head sums the information of all pixels in the input tensor, weighted according to their similarity to the query vector, and projects a linear output layer. The CLIP attention layer produces a single global embedding per image.

[0041] The SUM-CLIP head 230 aims to represent the "average" semantics in images with a single query. The SUM-CLIP head 230 captures multiple objects by employing additional adaptive queries and decoder layers prior to the CLIP head.

[0042] The dense CLIP head 210 focuses on the local semantics induced by the original CLIP weights. The dense CLIP head 210 aims to utilize the local semantics already captured by the spatial locations at the entrance to the attention layer.

[0043] The visual CLIP backbone 202 serves to encode a visual dataset. The visual CLIP backbone 202 can be referred to as a first visual backbone. The visual dataset generally comprises the unlabeled image data 206a and the labeled image data 206b. The resulting encoded data 208 are presented to the unsupervised open vocabulary system 180 and the supervised open vocabulary system 190.

[0044] The dense CLIP head 210 generates fixed, pre-trained embeddings of the open vocabulary from the encoded data 208 generated by the visual CLIP backbone 202. The embeddings provide the pseudo-labels 228, which are used to train the trainable part of the supervised open vocabulary system 190.

[0045] Cluster module 212 generally implements a fixed cluster module. Cluster module 212 groups similar image-vector-text-vector pairs into multiple clusters.

[0046] The first image embedding 216 can represent a cluster CLIP image embedding. The image embeddings generally provide numerical representations of images that capture semantic meaning and visual features as numerical vectors. The first image embeddings 216 provide initial image vector representations 220 of the associated unlabeled image data 206a and the labeled image data 206b, as processed by the dense CLIP head 210.

[0047] The first text embedding 218 can implement WordNet text embeddings. WordNet is a database of English words developed by Princeton University (Princeton, New Jersey). Text embeddings are generally neurolinguistic (NLP) techniques that convert text data into numerical vectors. The first text embedding 218 provides text vector representations 222 of the associated text strings.

[0048] Pairs of image vector representations 220 and text vector representations 222 fill the embedding space 224. After contrastive pretraining, diagonal pairs generally have high cosine similarities, while off-diagonal pairs have lower cosine similarities. Pairs with similarities above a threshold 226 are selected as pseudo-labels 228 and presented to the supervised open vocabulary system 190.

[0049] In addition to the SUM-CLIP head 230, the supervised open vocabulary system 190 can also contain a list of machine learning queries 232. The summary head 230 generally receives the coded data 208 from the visual CLIP backbone 202 and generates the second image embeddings 234.

[0050] The SUM-CLIP head 230 is generally able to generate dense video labels and create summaries 236 by selecting key frames from the video based on the encoded data 208 and the machine learning queries 232.

[0051] The machine learning queries 232 implement a set of several (e.g. N) queries that are used to train the supervised open vocabulary system 190.

[0052] The second image embeddings 234 generally provide second image vector representations 238 for the dense video labels and form summaries 236 that are processed by the SUM-CLIP head 230.

[0053] The basic categories 104 implement the sets of text strings 104a-104n. The text strings 104a-104n are processed by the CLIP text encoder 102 ( Fig. 1) used to create text embeddings 244 that correspond to the basic truth.

[0054] The grounded truth text embeddings 244 can be paired with suitable second image vector representations 238. The first cross-entropy losses 250 are a metric used during training to measure how well the resulting classification model performs. As the open vocabulary capacity of the grounded text embeddings 244 increases, the supervised open vocabulary system 190 tends to lose its open vocabulary capabilities.

[0055] The pseudo-labels 228 can be paired with matching generated second image vector representations 238. Second cross-entropy losses 252 are a metric used in training to measure how well the resulting classification model performs. The supervised system 190 with open vocabulary, as it is trained, has the advantage of learning from both the grounded truth text embeddings 244 and the pseudo-labels 228, thus avoiding losing some or most of the open vocabulary's capabilities.

[0056] Fig. Figure 3 shows a schematic diagram of an example implementation of the dense CLIP head 210 in accordance with one or more exemplary embodiments. The dense CLIP head 210 generally comprises a second backbone 260 and a first head network 262. The second backbone 260 can receive the encoded data 208. The first head network 262 generates first image embeddings 264.

[0057] The dense CLIP head 210 utilizes the local semantics already captured by the spatial locations at the entrance to the attention layer. The formulation is as follows: y i = c(z i ), where z i = v(x i The output embedding Y ∈ R K×Co is a tensor, y i is the representation of an i-th spatial pixel: Y={yi}Ki=1,yi∈Rl×Co K is determined by the size of the input image and the model stride (e.g., K = 196 for an image size of 448×448 and stride = 32). The formulation, implemented by removing the linear query and key layers and replacing the linear value and output layers with 1×1 convolutional layers (initialized with CLIP weights), essentially creates dense patch embeddings aligned to a CLIP output space.

[0058] The second backbone 260 can be a ResNet backbone. This second backbone extracts fixed features from the input images to create feature maps at different resolution levels. The feature maps at the lower resolution levels contain precise spatial information. The feature maps at the higher resolution level contain more nuanced semantic information due to a larger receptive field.

[0059] The first head network 262 serves to merge the features of all layers in order to improve the features with both higher accuracy and greater semantic meaning. Based on the improved features, the first main network 262 can further perform object detection, object class classification, and object bounding box regression. A linear layer 266 (e.g., V) and a first concatenation layer 268 can feed into an output layer 270. The first image embeddings 264 are presented by the output layer 270 of the dense CLIP head 210.

[0060] Fig. Figure 4 shows a schematic diagram of an exemplary implantation of the SUM-CLIP head 230 in accordance with one or more exemplary embodiments. The SUM-CLIP head 230 generally comprises a third backbone 272, a decoder layer 274, and a multi-head attention layer 276. The third backbone 272 can receive the encoded data 208. The decoder layer 274 can receive the machine-learning queries 232. The multi-head attention layer 276 generates the second image embeddings 234.

[0061] The SUM-CLIP head 230 is implemented as a multi-headed attention layer. The linear layers 278 (e.g., q, k, and v) of the SUM-CLIP head 230 are initialized with CLIP weights, and the trainable queries 232, the number of which is predefined, can be trained as follows: y i = c(z), and z i - softmax(q(Q)-k(X) T)(X). In contrast to CLIP, which aims to represent "average" semantics in images and treats x as a single query, SUM-CLIP is designed to capture multiple objects by incorporating additional learnable embeddings Q ∈ R N×Ci They are set as queries. The output embedding Y ∈ R N×Co is a tensor, and y i This is the representation of the i-th spatial pixel: Y={yi∈R}l×CoNi=1. In various embodiments, the architecture is a decoder variant of a CLIP attention layer with the adaptive queries 232 and can be extended as such by additional (e.g., L) decoder layers 274. The additional adaptive queries 232 generally achieve two goals: First, the output dimension is limited to a small number of representatives, which is suitable for large-scale retrieval frameworks without extensive post-processing. Second, the linear layers can be initialized with CLIP weights, enabling training focused on the original CLIP vision-language association.

[0062] The third backbone 272 implements another visual ResNet backbone. This backbone extracts finely tuned features from the input images to create feature maps at various resolutions. The feature maps at the lower resolutions contain precise spatial information. The feature maps at the higher resolutions contain more nuanced semantic information due to a larger receptive field.

[0063] Decoder layer 274 implements one or more (e.g., L) DETR (Detection Transformer) decoder layers. Decoder layer 274 performs set-based object detection using a transformer on a convolutional framework. Decoder layer 274 includes a self-attention layer 280, a cross-attention layer 282, and a forward feedback layer (FFW layer 284).

[0064] The multi-head attention layer 276 serves to execute an attention mechanism multiple times in parallel. The parallel attention outputs are then chained together and linearly transformed into an expected dimension. The multi-head attention layer 276 comprises the linear layers 278 (e.g., q, k, and v), a scaled dot-product attention layer 286, a second chaining layer 288, and an output layer 290. The output layer 290 presents the second image embedding 234.

[0065] In Fig. Figure 5 is a schematic diagram of an example implementation of a semi-supervised setup 300 in accordance with one or more exemplary embodiments. The semi-supervised setup 300 generally comprises the unsupervised open vocabulary system 180, the supervised open vocabulary system 190, and the image data 301. The image data 301 comprises a large unlabeled data set 302, a first smaller labeled data set 304a, and a second smaller labeled data set 304b. In some embodiments, the first smaller annotated data set 304a may be identical to the second smaller annotated data set 304b. In other embodiments, the first smaller annotated data set 304a may be different from the second smaller annotated data set 304b.

[0066] As in Fig. As shown in Figure 2, the unsupervised open vocabulary system 180 comprises the dense CLIP model 210, the cluster module 212, which generates the first image embeddings 216 that form the pseudo-labels 228. The supervised open vocabulary system 190 comprises the SUM-CLIP head 230, which generates the second image embeddings 234.

[0067] The large unlabeled dataset 302 and the first smaller labeled dataset 304a provide initial image data 132a for training the unsupervised open vocabulary system 180. The second smaller annotated dataset 304b provides further image data 132b for training the supervised open vocabulary system 190. The unsupervised open vocabulary system 180 also provides the pseudo-labels 228 to further train the supervised open vocabulary system 190 and to account for forgotten pairings.

[0068] The setup 300 generally uses an output (e.g., the pseudo-labels 228) of the unsupervised open vocabulary system 180 as auxiliary targets for training the supervised open vocabulary system 190. The lossy terms recovered by the pseudo-labels 228 aid in training for targets 306 in the unlabeled dataset 302, so that the supervised open vocabulary system 190 is trained on more than just the labeled images in the first smaller labeled dataset 304a and / or the second smaller labeled dataset 304b. Unlabeled data, which are usually more abundant, can also be used to further improve the accuracy in the semi-supervised system 300.

[0069] In Fig. 6 refers to Fig. Figure 2 shows a schematic diagram of an exemplary implementation of a recording system 320 according to one or more exemplary embodiments. The recording system 320 can be implemented in a vehicle 322. In various embodiments, the vehicle 322 can be an automobile. Other embodiments of the vehicle 322 can be a boat, construction equipment, an aircraft, or other devices that can use images to record features in an operational environment.

[0070] The vehicle 322 can include a text encoder 102, a camera 324, an image encoder 326, a classifier 328, and one or more targeting devices 330a-330b. The targeting devices can include a recording device 330a and an optional display device 330b. The text encoder 102 receives the text strings 332.

[0071] While the vehicle 322 moves in the operating environment, the camera 324 generates recorded images 325 of the local environment. A supervised open vocabulary system 190 ( Fig. 2) In the image encoder 326, the captured images 325 are encoded to generate the second image embeds 234. The second image embeds 234 are presented to the classifier 328. A text string 332 of a target (e.g., "A photo of a bird") can be entered into the text encoder 102. The text encoder 102 generates a requested target 334 (e.g., a text embed) for the text string 332. The target 334 is presented to the classifier 328. The classifier 328 can search the second image embeds 234 for an object identified by the target 334. If the search finds the object (e.g., the bird), the classifier 328 can generate an output image 338 containing the target object. The recording target device 330a generally stores the output image 338 for later analysis and / or additional training.Optionally, the output image 338 can be presented to a person traveling in vehicle 322 on the display device 330b. Although only a single text string 332 is requested in the described example, multiple text strings 332 can be requested successively.

[0072] In Fig. Figure 7 shows an example image 340 with recognized targets according to one or more exemplary embodiments. In the example, "passenger car" 342, "truck" 344, and "cone" 346 may be present as annotated categories in many autonomous vehicle datasets. In contrast, "car transporter trailer" 348, "toll roads" 350, and "turning mark" 352 may not be included in the datasets of autonomous vehicles.

[0073] Using both labeled and unlabeled datasets in a multi-criteria framework offers the advantages of both approaches. The multi-criteria system of the Recording System 320 demonstrates higher retrieval accuracy in the target dataset through fine-tuning: both for finely tuned categories (e.g., bicycles, motorcycles, priority signs) and for zero-shot categories that are not annotated in the target dataset.

[0074] In Fig. 8 refers to Fig. Figure 7 shows a schematic diagram of an exemplary implementation of a control unit 360 in accordance with one or more exemplary embodiments. The control unit 360 can be installed in the vehicle 322 and coupled with the camera 324, the recording targeting device 330a, and the display device 330b. The control unit 360 comprises a computer-aided processing device 362, a communication device 364, an input / output coordination device 366, and a storage device 368. In various embodiments, the control unit 360 may include further components, and some of the components are not present in some embodiments.

[0075] The processing device 362 can include a memory (e.g., a read-only memory (ROM) and a random access memory (RAM)) that stores processor-executable instructions and one or more processors that execute the processor-executable instructions. In embodiments in which the processing device 362 includes two or more processors, the processors can operate in parallel or in a distributed manner. The processing device 362 can run the operating system of the control unit 360. The processing device 362 can include one or more modules that execute programmed code or computer-aided processes or methods with executable steps. The modules shown can comprise a single physical device or functionality that extends over several physical devices.The processing device 362 may also contain programming modules, including the unattended open vocabulary system 180, the supervised open vocabulary system 190 and the text encoder 102.

[0076] The communication device 364 may include a communication / data link with a bus device configured to transmit data to various components of the vehicle, and may include one or more wireless transceivers to perform wireless communication.

[0077] The input / output coordination device 366 comprises hardware and / or software configured to enable the processing device 362 to receive and / or exchange data with onboard sensors of the vehicle 322, such as the camera 324, the recording targeting device 328a, and the display 328b. The input / output coordination device 366 can also enable the control of switches, modules, and processes throughout the vehicle 322 based on the findings made by the processing device 362.

[0078] The storage device 368 is a device that stores data generated or received by the control unit 360. The storage device 368 may include, but is not limited to, a hard disk drive, an optical drive, and / or a flash memory drive.

[0079] The control unit 360 is an exemplary computer-aided device capable of executing programmed code to operate the disclosed method. A number of different embodiments of the control unit 360 and the modules operating therein are conceivable, and the description is not intended to be limited to the examples listed here.

[0080] Fig. Figure 9 shows a schematic diagram of an exemplary first stage 382 of an offline system 380 in accordance with one or more exemplary embodiments. The first stage 382 generally comprises a list of input images 384, an image processing encoder 386, an image processing backbone and fine-tuning head 388, a database indexing module 390, and a database 392. The first stage 382 of the offline system 380 generally searches the list of images 384 for relevant concepts in large datasets and annotates the images for further use.

[0081] The first stage 382 serves to collect and index image embeddings for each image in the database 392. The image processing encoder 386 encodes the image data obtained from the list of images 384.

[0082] The backbone and the fine-tuning head 388 can implement a CLIP backbone and a SUM-CLIP head. The backbone and the fine-tuning head 388 are used to add image embeddings to the encoded images.

[0083] The database indexing module 390 indexes the encoded image data along with the image embeddings to enable faster access and retrieval later. The indexed image data is then transferred to database 392.

[0084] Database 392 consists of one or more storage media. Database 392 serves to store the indexed, encoded image data with the image embeddings.

[0085] Fig. Figure 10 shows a schematic diagram of an exemplary second stage 402 of the offline system 380 in accordance with one or more exemplary embodiments. The second stage 402 generally comprises the database 392 from the first stage 382 ( Fig.9), a text encoder 404, an image encoder 406 and a search module 408. The second stage 402 of the offline system 380 generally searches the database 392 for stored images 410 with corresponding versions of input text strings 412 and / or input images 414.

[0086] The text encoder 404 encodes one or more input texts 412 (e.g., "wheelchair"). The resulting target text embeddings are presented to the search module 408.

[0087] The image encoder 406 encodes one or more input images 414. The resulting requested target image embeddings are presented to the search module 408.

[0088] The search module 408 searches for relevant embeddings in the indexed images stored in database 392. Up to a predefined number (e.g., K images) of matching images 410 with the best correlation to the relevant embeddings can be copied from database 392 and presented as offline results.

[0089] Various implementations of the system and / or method provide a framework that overcomes the limitations of existing supervised, fine-tuned, and fixed pre-trained techniques. Compared to existing supervised techniques, the framework delivers improved results for zero-shot categories (e.g., categories of targets not seen during training). Compared to existing fixed procedures, the framework generally allows for faster inference time and higher recall accuracy for both trained and new categories. Furthermore, multi-criteria training helps mitigate open-ended vocabulary forgetting and provides improvement through training on unlabeled data.

[0090] The framework combines two open vocabulary retrieval approaches, leveraging their respective advantages. Similar to supervised fine-tuning techniques, the framework is fine-tuned in a supervised manner on a target dataset, compensating for potential domain shifts and increasing the accuracy of the trained categories. Regarding fixed pretraining, the system uses both the pre-trained model backbones (for model transfer) and the pre-trained model outputs (as auxiliary targets) for training, ensuring that the pre-trained vision-language association is preserved even for untrained categories. Compared to existing fixed pretraining techniques, a summary header is fine-tuned, simplifying and accelerating the coding process by eliminating the need for post-processing.The finely tuned network is widely applicable and can be integrated into online and / or offline applications.

[0091] Embodiments of the description generally provide a multi-target, dense, open vocabulary system comprising an image encoder and a classifier. The image encoder has a summarizing contrastive speech image pre-training head (CLIP) and is coupling to an originating device. The summarizing CLIP head is trained on supervised losses of unlabeled and labeled image data, loses open vocabulary capabilities as capacity increases, and is trained on pseudo-label losses from a variety of pseudo-labels that compensate for the loss of open vocabulary capabilities. The variety of pseudo-labels is generated from a variety of text embeddings based on similarities with a variety of average semantics produced by a dense CLIP head.The summarizing CLIP head is capable of receiving a multitude of captured images from a source device and generating a multitude of image embeddings based on the multitude of captured images.

[0092] The classifier is coupled to the image encoder, can be coupled to a text encoder, and can be connected to a targeting device. The classifier is capable of receiving one or more targets from the text encoder, receiving the multitude of image embeds from the summary CLIP, classifying the multitude of image embeds to identify one or more output images containing the one or more targets, and presenting the one or more output images to the targeting device.

Claims

[1] Multi-criteria dense system with an open vocabulary that includes: an image encoder (326) with a summarizing contrast speech image pretraining head (summarizing CLIP head, 230) that can be connected to an originating device, wherein: the summary CLIP head (230) is trained using monitored losses from unlabeled image data (206a) and labeled image data (206b); The summary CLIP header (230) loses the ability to use an open vocabulary as its capacity increases; the summary CLIP head (230) is trained on pseudo-label losses from a variety of pseudo-labels (228) that compensate for the loss of open vocabulary skills; the multitude of pseudo-labels (228) is generated from a multitude of text embeddings (218) based on similarities with a multitude of average semantics, which are generated by a dense CLIP head (210); and The summary clip header (230) is ready for use at: Receiving a large number of captured images (325) from the originating device; and Generating a multitude of image embeddings based on the multitude of captured images (325); and a classifier (328) coupled to the image encoder (326), which can be coupled to a text encoder (102) and which can be coupled to a targeting device (330a-330b), wherein the classifier (328) is ready for operation to: Receiving one or more destinations from the text encoder (102); Receiving the multitude of image embeddings from the summary CLIP head (230); Classifying the multitude of image embeddings to identify one or more output images (338) that contain the one or more targets; and Presenting one or more output images (338) to the targeting device. [2] Multi-criteria dense open vocabulary system according to claim 1, wherein the summary CLIP head (230) comprises: a backbone operation for extracting a large number of finely tuned features from the unlabeled image data (206a) and the labeled image data (206b). [3] Multi-criteria dense open vocabulary system according to claim 2, wherein the summary CLIP head (230) further comprises: a recognition transformer decoder layer (274) that predicts a variety of objects based on a variety of finely tuned features and a variety of learnable queries (232). [4] Multi-criteria dense open vocabulary system according to claim 3, wherein the summary CLIP head (230) further comprises: a multi-head attention layer (276) that operates in such a way as to generate the multitude of image embeddings in response to the multitude of finely tuned features and the multitude of objects. [5] Multi-criteria dense open vocabulary system according to claim 1, wherein the dense CLIP head (210) comprises: a backbone (202) which serves to extract a variety of fixed features from the unlabeled image data (206a) and the labeled image data (206b). [6] Multi-criteria dense open vocabulary system according to claim 5, wherein the dense CLIP head (210) further comprises: a cluster module (212) that clusters the multitude of fixed features to generate the multitude of average semantics. [7] Multi-criteria dense open vocabulary system according to claim 6, wherein the dense CLIP head (210) further comprises: an embedding system (214) that is able to generate the majority of pseudo-labels (228) from the majority of text embeddings (218) based on the majority of average semantics. [8] Multi-criteria dense system with open vocabulary according to claim 1, wherein the originating device is a camera (324) that generates the plurality of recorded images (325). [9] Multi-criteria dense open vocabulary system according to claim 1, wherein the targeting device (330a-330b) is a memory for recording one or more output images (338). [10] Method for multi-criteria image recording with an open vocabulary, which includes the following: Receiving a plurality of captured images (325) at an image encoder (326) from a source device, wherein: the image encoder (326) has a summary contrast speech image pre-training head (summary CLIP head, 230); The summary CLIP head (230) is trained using a controlled loss from unlabeled image data (206a) and labeled image data (206b); The summary CLIP header (230) loses the ability to use an open vocabulary as its capacity increases; the summary CLIP head (230) is trained on pseudo-label loss from a variety of pseudo-labels (228), which compensates for the loss of open vocabulary skills; and the multitude of pseudo-labels (228) is generated from a multitude of text embeddings (218) based on similarities with a multitude of average semantics generated by a dense CLIP head (210); Generating a multitude of image embeddings using the summarise CLIP head (230) based on the multitude of captured images (325); Receiving one or more destinations from a text encoder (102) at a classifier (328); Receiving the multitude of image embeddings from the summary CLIP head (230) at the classifier (328); Classifying the multitude of image embeddings to identify one or more output images (338) that contain the one or more targets; and Presenting one or more source images on a targeting device (330a-330b).