System and method for transferring semantic segmentation via base model learnable image cue
By combining the unlearning multimodal basic model and the learnable task-specific model, the confidence bias and high computational cost problems in cross-domain transfer learning in semantic segmentation tasks are solved, and the optimal combination of cross-domain generalization capabilities and intra-domain performance is achieved.
Patent Information
- Application Number
- CN202411848789.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-17
AI Technical Summary
In semantic segmentation tasks, prior art is difficult to effectively solve the confidence bias and high computational cost problems in cross-domain transfer learning, especially when the expected use of the basic model is not fully aligned with the pre-task.
The combination of cross-domain generalization capabilities and intra-domain performance is achieved by leveraging non-learning multimodal basic models and coupling them with learnable image prompt models, task-specific encoder models, decoder models, and learnable fusion models.
This method can improve downstream tasks specific performance, reduce computational costs, and avoid confidence bias while maintaining cross-domain generalization capabilities.
Smart Images

Figure CN120163974A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to neural network pipelines and improving the performance, transferability, and generalizability of such pipelines, such as those including a base model. Background Art
[0002] Semantic segmentation remains a popular task in the computer vision and machine learning communities, largely due to its importance for scene understanding in autonomous driving. However, this task presents many challenges, as the effort involved in labeling every pixel of every image across a variety of weather conditions in various geographical regions precludes large-scale dataset creation and the deployment of learning-based models in autonomous vehicles. Due to this labeling cost, it has become popular to generate a dataset of synthetic images from 3D simulators, train a model on these synthetic source domains to predict semantic segmentation, and then transfer those trained models to the real-world context to predict semantic segmentation on an unlabeled photo-realistic target dataset. To avoid transferring spurious knowledge in the target domain, various transfer learning methods have been proposed. Traditionally, popular transfer methods involve some form of self-training, for example, where pseudo-labels are generated from the most confident predictions of the source model on target domain observations. However, these methods suffer from some form of confidence bias, as they are vulnerable to disproportionate contributions from the majority classes in the source training set. Additionally, standard transfer methods exhibit high computational costs, as the pseudo-labeling process must be repeated in multiple stages in order to obtain an optimal one that produces a competitive inference model.
[0003] A base model can be a model with large data representation capabilities (e.g., through a large number of layer sizes and internal weight and bias parameters, as in large language models or "LLMs" or vision-language models or "VLMs") that has been additionally pre-trained on multiple large-scale datasets. These datasets may consist of millions of paired data samples - for example, images with their captions - and the base model can be trained with one of several objectives. One objective may be to learn to score the alignment (similarity) between inputs (e.g., an arbitrary image and an arbitrary text caption). However, challenges arise when the intended use ("downstream task") of the base model is not perfectly aligned with the pretext task of the base model: in these cases, the representations produced by the base model may not be useful; in the worst case, they may actually cause the entire task-specific framework to perform worse on the downstream task and not generalize better compared to the case where the base model was not included at all. The systems and methods described below discuss how to achieve the best of both worlds: leveraging the representation and generalization capabilities of the base model while maintaining or improving task-specific performance. Summary of the Invention
[0004] According to a first illustrative embodiment, a computer-implemented method includes receiving one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images; in response to utilizing the fixed text prompts and the one or more images at a base model associated with a machine learning network, outputting an intermediate representation from a sequence of generated objects and tasks; decoding the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with the fixed text prompts and images; and in response to identifying a highest probability associated with the matrix using label selection, outputting a final label associated with a vision-based prediction task.
[0005] According to a second illustrative embodiment, a method includes receiving one or more fixed text prompts at a base model; receiving one or more images at a learnable image prompt network, wherein the fixed text prompts are associated with the one or more images; generating a fixed-dimensional continuous latent vector at the learnable image prompt network using the one or more images; in response to utilizing the fixed text prompts and the fixed-dimensional continuous latent vector at a base model associated with a machine learning network, outputting an intermediate representation from a sequence of generated objects and tasks; decoding the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with the fixed text prompts and images; and in response to identifying a highest probability associated with the matrix using label selection, outputting a final label associated with a vision-based prediction task.
[0006] According to a third illustrative embodiment, a system includes a controller configured to receive one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images. The controller is further configured to, in response to utilizing the fixed text prompts and the one or more images at a base model associated with a machine learning network, output an intermediate representation from a sequence of generated objects and tasks, decode the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with the fixed text prompts and images, and in response to identifying a highest probability associated with the matrix using label selection, output a final label associated with a vision-based prediction task. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Shows a system for training a neural network according to an embodiment.
[0008] Figure 2 Shows a computer-implemented method for training and utilizing a neural network according to an embodiment.
[0009] Figure 3A Illustrates the architecture of a system configured to perform transfer semantic segmentation via self-supervision and prompting of a multimodal foundation model.
[0010] Figure 3B Illustrates the architecture of a system configured to perform transfer semantic segmentation via self-supervision from a multimodal foundation model.
[0011] Figure 3C Illustrates the architecture of a system configured to perform transfer semantic segmentation via self-supervision from a multimodal foundation model.
[0012] Figure 4 Illustrates an overview of a system for an application scenario according to one embodiment.
[0013] Figure 5 Depicts a schematic diagram of the interaction between a computer-controlled machine and a control system according to an embodiment.
[0014] Figure 6 Depicts a Figure 5 control system configured to control a vehicle, which can be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot.
[0015] Figure 7 Depicts Figure 5 a control system configured to control manufacturing machines (such as part of a production line) of a manufacturing system, such as a stamping machine, a cutting machine, or a gun drill.
[0016] Figure 8 Depicts Figure 5 a control system configured to control a power tool with at least a partially autonomous mode, such as a drill or a driver.
[0017] Figure 9 Depicts a Figure 5 control system configured to control an automated personal assistant.
[0018] Figure 10 Depicts Figure 5 a control system configured to control a surveillance system, such as an access control system or a monitoring system.
[0019] Figure 11 Depicts Figure 5 a control system configured to control an imaging system, such as an MRI device, an x-ray imaging device, or an ultrasound device. Detailed Description
[0020] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The figures are not necessarily to scale; some features may be enlarged or reduced to show details of particular components. Accordingly, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching one skilled in the art to employ the embodiments in various ways. As will be understood by one of ordinary skill in the art, the various features illustrated and described with reference to any one figure may be combined with features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. Combinations of the illustrated features provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired.
[0021] In the present disclosure, the system may utilize a multimodal foundation model in a machine learning training and inference pipeline. The foundation model may be a model with large data representation capabilities (e.g., through a large number of layer sizes and internal weight and bias parameters, as in large language models or "LLMs"), which has been additionally pre-trained on multiple large-scale datasets. These datasets may consist of millions of paired data samples - e.g., images with their captions - and the LLM may be trained with one of several objectives. One objective may be to learn to score the alignment (similarity) between inputs (e.g., any image and any text caption). Another objective may be to reconstruct an image given a natural language text caption plus the corresponding image with random tiles missing or deleted. In addition to these training objectives, some intermediate continuous-valued vector representations from the foundation model may be used to perform (pre-) tasks such as image classification, image captioning, object segmentation, and semantic segmentation. An implicit assumption may be that the intended use (downstream) of the foundation model is also to perform one of the tasks in its set of pre-tasks. Through this extensive pre-training (using large datasets with challenging training objectives on various pre-tasks), the LLM will accumulate sufficient experience in multiple domains to serve as a basis for task-specific architectures built on top of the foundation model. In operation, after their pre-training, the foundation models may become untrainable ("frozen") and simply be used in "inference mode" on various downstream tasks. In this way, the foundation model, through its experience in modeling several tasks and domains, enables cross-domain generalization capabilities for downstream task-specific frameworks. A novel aspect of the present disclosure may include the training and development of additional modules that perform non-linear transformations on the outputs of the frozen foundation model. Accordingly, the system can extract the best of both worlds: cross-domain generalization capabilities (preserved from the foundation model), and competitive in-domain performance (from the task-specific backbone).
[0022] The present disclosure provides a method and system for a learnable and reusable perception model that can be attached to any larger system that needs to operate in a variety of real-world environments. Across different operating environments, there are data domain distribution shifts (e.g., differences in the visual appearance of objects, differences in the physical dynamics of the larger system, differences in the sensor models used by the larger system to sense the world, etc.), which pose challenges to any non-generalizable components in the larger system. Examples of larger systems that would significantly benefit from the generalizable perception module detailed in this disclosure include (but are not limited to) mobile robots, autonomous vehicles, and smart cameras. One novelty of our generalizable perception module stems from the fact that it utilizes an unlearnable (“frozen”) multimodal base model in order to enhance the cross-domain performance of the larger system; however, directly using the base model is challenging because it requires strong assumptions that the intended use (“downstream task”) is exactly aligned with how the base model was initially trained (“(one or more) upstream tasks”). If there is no alignment between the (one or more) upstream tasks and the downstream task, then the in-domain performance of the larger system will be severely affected. Thus, the second source of novelty in our work is that the base model can be coupled with task-specific modules, which in a way that preserves the cross-domain generalization ability (obtained from the multimodal base model), while also achieving competitive in-domain performance on the downstream tasks of interest. The ways in which this coupling can occur are through multiple components: (i) a learnable image prompting model, (ii) a task-specific encoder model, (iii) a task-specific decoder model, and (iv) a learnable fusion model. These component models can be trainable neural networks or any other type of model with learnable function parameters. For illustrative purposes, the present disclosure will specifically describe the foregoing components in the context of the transfer semantic segmentation downstream perception task.
[0023] The present invention provides a method and system for a learnable and reusable perception model that can be attached to any larger system that needs to operate in a variety of real-world environments. Across different operating environments, there are data domain distribution shifts (e.g., differences in the visual appearance of objects, differences in the physical dynamics of the larger system, differences in the sensor models used by the larger system to sense the world, etc.), which poses challenges to any non-generalizable components in the larger system. Examples of larger systems that will significantly benefit from the generalizable perception module detailed in this disclosure include (but are not limited to) mobile robots, automated vehicles, and smart cameras. One of the novelties of our generalizable perception module stems from the fact that it utilizes a non-learnable ("frozen") multimodal base model in order to enhance the cross-domain performance of the larger system; however, using the base model directly is challenging because it requires a strong assumption that the intended use ("downstream task") is fully aligned with how the base model was originally trained ("precursor task(s)"). If there is no alignment between the precursor task(s) and the downstream task, the in-domain performance of the larger system will be drastically affected. Therefore, the second source of novelty in our work is that the base model can be coupled with task-specific modules in a way that maintains cross-domain generalization capabilities (obtained from the multimodal base model) while also achieving competitive intra-domain performance on the downstream tasks of interest. The way this coupling may occur is through multiple components: (i) a learnable image cue model, (ii) a task-specific encoder model, (iii) a task-specific decoder model, and (iv) a learnable fusion model. These component models can be trainable neural networks or any other type of model with learnable function parameters. For the sake of illustration, this disclosure will specifically describe the aforementioned components in the context of the downstream perception task of transfer semantic segmentation.
[0024] Reference is now made to the embodiments illustrated in the accompanying figures, which may apply these teachings to a machine learning model or a neural network. Figure 1 A system 100 for training a neural network (e.g., a deep neural network) is shown. The system 100 may include an input interface for accessing training data 102 for the neural network. For example, Figure 1As illustrated, the input interface may be constituted by a data storage interface 104, and the data storage interface 104 may access training data 102 from a data storage device 106. For example, the data storage interface 104 may be a memory interface or a permanent storage interface, such as a hard disk or SSD interface, but may also be a personal, local area, or wide area network interface, such as a Bluetooth, Zigbee, or Wi-Fi interface or an Ethernet or fiber optic interface. The data storage device 106 may be an internal data storage device of the system 100, such as a hard disk drive or SSD, but may also be an external data storage device, such as a network-accessible data storage device.
[0025] In some embodiments, the data storage device 106 may further include a data representation 108 of an untrained version of the neural network, which may be accessed by the system 100 from the data storage device 106. However, it will be appreciated that the training data 102 and the data representation 108 of the untrained neural network may also each be accessed from different data storage devices, for example, via different subsystems of the data storage interface 104. Each subsystem may have the types as described above for the data storage interface 104. In other embodiments, the data representation 108 of the untrained neural network may be internally generated by the system 100 based on the design parameters of the neural network and may thus not be explicitly stored on the data storage device 106. The system 100 may further include a processor subsystem 110, which may be configured to provide an iterative function as a replacement for the stack of the neural network to be trained during the operation of the system 100. Here, the respective layers of the replaced stack may have weights shared with each other and may receive the output of the previous layer as input, or for the first layer of the stack, receive the initial activation and a part of the input of the stack. The processor subsystem 110 may also be configured to iteratively train the neural network using the training data 102. Here, the training iteration of the processor subsystem 110 may include a forward propagation part and a backward propagation part. The processor subsystem 110 may be configured to perform the forward propagation part by, in addition to other operations that may execute to define the forward propagation part, determining the equilibrium point of the iterative function at which the iterative function converges to a fixed point, where determining the equilibrium point includes using a numerical root-finding algorithm to find the root solution of the iterative function minus its input, and by providing the equilibrium point as a replacement for the output of the stack in the neural network. The system 100 may further include an output interface for outputting a data representation 112 of the trained neural network, which may also be referred to as trained model data 112. For example, also as Figure 1As illustrated, the output interface may be constituted by the data storage interface 104, where in these embodiments, the interface is an input / output ("IO") interface via which the trained model data 112 may be stored in the data storage device 106. For example, the data representation 108 defining the "untrained" neural network may be at least partially replaced by the data representation 112 of the trained neural network during or after training because the parameters of the neural network (such as weights, hyperparameters, and other types of parameters of the neural network) may be adapted to reflect the training on the training data 102. This is also illustrated in Figure 1 by reference numerals 108, 112, where reference numerals 108, 112 refer to the same data record on the data storage device 106. In other embodiments, the data representation 112 may be stored separately from the data representation 108 defining the "untrained" neural network. In some embodiments, the output interface may be separate from the data storage interface 104 but generally may be of the type described above for the data storage interface 104.
[0026] The structure of system 100 is an example of a system that can be used to train the image-to-image machine learning model and the mixer machine learning model described herein. Figure 2 Additional structures for operating and training machine learning models are shown in
[0027] Figure 2 System 200 is depicted that implements the machine learning models described herein (such as the image-to-image machine learning model, the mixer machine learning model, and the pre-trained reference model described herein). System 200 may be implemented to perform the image quantization process described herein. System 200 may include at least one computing system 202. The computing system 202 may include at least one processor 204 operatively connected to a memory unit 208. The processor 204 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. The CPU 206 may be a commercially available processing unit that implements an instruction set such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, the CPU 206 may execute stored program instructions retrieved from the memory unit 208. The stored program instructions may include software that controls the operation of the CPU 206 to perform the operations described herein. In some examples, the processor 204 may be a system-on-chip (SoC) that integrates the functionality of the CPU 206, the memory unit 208, the network interface, and the input / output interface into a single integrated device. The computing system 202 may implement an operating system for managing various aspects of the operation. Although one processor 204, one CPU 206, and one memory 208 are shown in Figure 2 FIG., more than one of each may of course be utilized throughout the system.
[0028] The memory unit 208 may include volatile and non-volatile memories for storing instructions and data. The non-volatile memory may include solid-state memory such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or powered off. The volatile memory may include static and dynamic random access memories (RAMs) that store program instructions and data. For example, the memory unit 208 may store the machine learning model 210 or algorithm, the training data set 212 of the machine learning model 210, and the original source data set 216.
[0029] The computing system 202 may include a network interface device 222 configured to provide communication with external systems and devices. For example, the network interface device 222 may include wired and / or wireless Ethernet interfaces defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard family. The network interface device 222 may include a cellular communication interface for communicating with cellular networks (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or cloud.
[0030] The external network 224 may be referred to as the World Wide Web or the Internet. The external network 224 may establish standard communication protocols between computing devices. The external network 224 may allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 may communicate with the external network 224.
[0031] The computing system 202 may include an input / output (I / O) interface 220 that may be configured to provide digital and / or analog inputs and outputs. The I / O interface 220 is used to transfer information between the internal storage device and external input and / or output devices (e.g., HMI devices). The I / O 220 interface may include associated circuitry or bus networks to transfer information to and from the (one or more) processors and storage device and between them. For example, the I / O interface 220 may include digital I / O logic lines that may be read or set by the (one or more) processors, handshake lines for supervising data transfer via the I / O lines, timing and counting facilities, and other structures known to provide such functionality. Examples of input devices include keyboards, mice, sensors, etc. Examples of output devices include monitors, printers, speakers, etc. The I / O interface 220 may include additional serial interfaces (e.g., Universal Serial Bus (USB) interfaces) for communicating with external devices.
[0032] The computing system 202 may include a human-machine interface (HMI) device 218, and the human-machine interface device 218 may include any device that enables the system 200 to receive control inputs. Examples of input devices may include human-machine interface inputs such as keyboards, mice, touchscreens, voice input devices, and other similar devices. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphical and text information to the display device 232. The display device 232 may include an electronic display screen, a projector, a printer, or other suitable devices for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with remote HMIs and remote display devices via a network interface device 222.
[0033] The system 200 may be implemented using one or more computing systems. Although the example depicts a single computing system 202 implementing all the described features, it is intended that various features and functions may be separated and implemented by multiple computing units that communicate with each other. The particular system architecture chosen may depend on a variety of factors.
[0034] The system 200 may implement a machine learning algorithm 210 configured to analyze a raw source data set 216. The raw source data set 216 may include raw or unprocessed sensor data, which may represent the input data set for the machine learning system. The raw source data set 216 may include videos, video clips, images, text-based information, audio or human speech, time series data (e.g., pressure sensor signals over time), and raw or partially processed sensor data (e.g., radar maps of objects). Refer to Figures 5 - 11 Several different examples of inputs are shown and described. In some examples, the machine learning algorithm 210 may be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm may be configured in an automotive application to identify street signs or pedestrians in an image. The (one or more) machine learning algorithms 210 may include algorithms configured to operate the image-to-image machine learning model, the mixer machine learning model, and the pre-trained reference model described herein.
[0035] The computer system 200 may store a training data set 212 for the machine learning algorithm 210. The training data set 212 may represent a previously constructed data set for training the machine learning algorithm 210. The machine learning algorithm 210 may use the training data set 212 to learn the weighting factors associated with the neural network algorithm. The training data set 212 may include a source data set that has corresponding outcomes or results that the machine learning algorithm 210 attempts to replicate via the learning process. In this example, the training data set 212 may include input images containing objects (e.g., street signs). The input images may include various scenarios in which the objects are identified.
[0036] The machine learning algorithm 210 can operate in a learning mode using the training data set 212 as input. The machine learning algorithm 210 can perform multiple iterations using the data from the training data set 212. With each iteration, the machine learning algorithm 210 can update internal weighting factors based on the results obtained. For example, the machine learning algorithm 210 can compare the output results (e.g., a reconstructed or supplemented image in the case where image data is the input) with those included in the training data set 212. Since the training data set 212 includes the expected results, the machine learning algorithm 210 can determine when the performance is acceptable. After the machine learning algorithm 210 reaches a predetermined performance level (e.g., 100% consistent with the results associated with the training data set 212) or converges, the machine learning algorithm 210 can be executed using data not in the training data set 212. It should be understood that in the present disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations have occurred, or the residuals are small enough (e.g., the change in the approximate probability in the iteration is changing at less than a threshold), or other convergence conditions. The trained machine learning algorithm 210 can be applied to a new data set to generate annotated data.
[0037] The machine learning algorithm 210 can be configured to identify specific features in the original source data 216. The original source data 216 can include multiple instances or input data sets for which supplementary results are desired. For example, the machine learning algorithm 210 can be configured to identify the presence of road signs in video images and annotate the occurrences. The machine learning algorithm 210 can be programmed to process the original source data 216 to identify the presence of specific features. The machine learning algorithm 210 can be configured to identify the features in the original source data 216 as predetermined features (e.g., road signs). The original source data 216 can be derived from various sources. For example, the original source data 216 can be actual input data collected by a machine learning system. The original source data 216 can be machine-generated for testing the system. As an example, the original source data 216 can include raw video images from a camera.
[0038] In an example, the original source data 216 can include image data representing an image. Applying the machine learning algorithms described herein (e.g., an image-to-image machine learning model, a mixer machine learning model, and a pre-trained reference model), the output can be a quantized version of the input image.
[0039] Figure 3AIllustrated is the architecture of a system configured to perform transfer semantic segmentation via self-supervision and prompting of a multimodal base model. The system can include a novel framework that utilizes base models such as CLIP, UnifiedIO, GroupViT, etc. for transfer learning in the context of semantic segmentation for autonomous driving. Compared with traditional unsupervised domain adaptation methods, such a system can have competitive performance with this method but with relatively minimal computational cost. Due to the improved domain generalization brought by this method, although the target performance is competitive, the system can show that the source performance is not affected. In addition, such an embodiment can exceed the state-of-the-art (SOTA) task performance, even when such a system can be configured for zero-shot inference. Such a system can propose a novel framework for learning continuous prompts in semantic segmentation, thus allowing our method to avoid full fine-tuning of the base model while also avoiding the expensive multi-stage self-training process in traditional domain adaptation methods.
[0040] In one embodiment, such a framework includes a multimodal base model 311, a learnable image-based prompt pattern 305, a task-specific encoder 307, a task-specific decoder 315, a learnable fusion module 313. An illustration of our method is included in the Figure 3A below. Alternative configurations of similar methods are included in Figure 3B and 3C below. The input X can include a fixed text prompt 301. The fixed text prompt 301 can include the task and the associated object. For example, the task can include "semantic segmentation" or another task related to image recognition. The object can refer to the associated object in the image. The input X can also include an RGB image 303. Of course, any type of image can be utilized, such as sound images, pictures, videos, ultrasounds, radars, etc. The image prompt network D θ is a learnable mapping function - from the input X to a continuous latent vector space Z of a certain fixed dimension prompt :
[0041] D θ : X → Z prompt .
[0042] One job of the image prompt network 305 can be to generate the learned image prompts. In order to be an effective prompt input to the multimodal base model 311, the statistical distribution of Z prompt can have certain properties. First, the latent vector space Z promptThe statistical distribution such that its sub-domains (statistical patterns) match the sub-domains of the following: (i) the source input X distribution, (ii) the relevant sub-domains supported by the base model from its pre-training objectives and pre-tasks, and (iii) the sub-domains of the source label Y distribution. The system can include further regularizing Z by using domain knowledge prompt of the distribution to enable matching with the target input / label distribution. The form and effectiveness of this domain knowledge (knowledge graph, statistical priors, constraints, logical rules, expert demonstrations, or labels, etc.) vary depending on the downstream application. To encourage these properties, relative to the source dataset that has known meaningful sub-domains in its data distribution, the system can make D θ obey a meta-learning objective. Second, Z prompt representation must have a clearly expressed statistical pattern (high density within the pattern and low density between patterns), which enables effective reasoning. To pursue these properties, relative to the initial statistical distribution represented by Z prompt (e.g., following from the above meta-learning stage) and the similarity function (e.g., cosine distance metric function), the system can make D θ obey a contrastive learning objective. The specific formulation of the contrastive objective that will be proven effective can vary by domain and depends on the type of input. In the context of image-based prompt learning, in the transfer semantic segmentation task, the system can enforce that prompts from images that look "similar" (e.g., similar in road layout, object arrangement) - although belonging to different sub-domains (e.g., bright summer weather vs. dark winter conditions) - should be close to each other in the projected representation space.
[0043] The task-specific encoder E φ 307 can be a learnable function mapping - from the input X to a continuous latent vector space Z of a fixed dimension enc :
[0044] E φ : X → Z enc
[0045] One of its jobs can be to produce a latent vector that serves as a concise and dense summary of the input(s), which can be of one or more modalities (e.g., RGB image, natural language text, audio signal, etc.). E φ can be implemented as a function approximator with learnable parameters φ, such as a neural network.
[0046] The base model can provide a natural interface for a predetermined, fixed prompt input, e.g., in natural language text. This fixed text prompt allows for an initial specialization of the base model on a particular task (e.g., semantic segmentation) and domain (e.g., a particular set of objects, a particular environmental setting). Due to the advantages of this initial specialization, using an appropriate supported fixed text prompt is crucial for using the base model in downstream tasks. The fixed text prompt supported by the base model is directly associated with the pre - tasks that the base model was initially trained to perform; an appropriate fixed text prompt can be used to perform the corresponding tasks. However, challenges arise when there is a mismatch between the (one or more) pre - tasks and the downstream tasks that utilize the base model. Under these conditions, "custom" fixed text prompts must be formulated (or "designed") for the task at hand. In fact, for cases where the base model does not support the downstream task (e.g., semantic segmentation), the system can construct multiple fixed text prompts which include sufficient object references and task specifications for enabling the base model to generate a series of outputs that can subsequently be combined:
[0047]
[0048] where O is the set of objects that the base model can detect, and task is a single task supported by the base model. For example, in the case of using the UnifiedIO base model (Lu et al., 2022), O is the set of object classes to be detected in a downstream semantic segmentation task (e.g., road, sidewalk, person, rider, car, truck, bus, motorcycle, bicycle, caravan, trailer, building, wall, fence, guardrail, bridge, tunnel, pole, traffic sign, traffic light, vegetation, terrain, sky, ground), and task is the selected pre - task supported by the base model (e.g., object / instance segmentation).
[0049] In one embodiment, the system can inject the learned prompt representation into various models or modules of the system. For example, using only the fixed text prompt is often not sufficient for competitive performance on downstream tasks. Thus, the system can consider additionally using the learned prompt, where O is the set of objects that the base model can detect, and task is a single task supported by the base model. For example, in the case of using the UnifiedIO base model, O is the set of object classes to be detected in a downstream semantic segmentation task (e.g., road, sidewalk, person, rider, car, truck, bus, motorcycle, bicycle, caravan, trailer, building, wall, fence, guardrail, bridge, tunnel, pole, traffic sign, traffic light, vegetation, terrain, sky, ground), and task is the selected pre - task supported by the base model (e.g., object / instance segmentation).
[0050] Using fixed - text prompts alone is often insufficient for competitive performance on downstream tasks. Therefore, the system can consider additionally using the learned prompt representation Z prompt . After having forced Z prompt to be a valid sub - domain mapping (the first set of desired properties) and have a clear representation in its distribution (the second set of desired properties), as discussed in the "Learnable Image Prompt Model" subsection above, the system can choose where and how to effectively inject this learned prompt into the input interface of the multimodal base model. Since Z prompt is not an image, but a representation of an image (a continuous - valued vector representing a latent variable that acts as a concise summary), the system can choose to feed Z prompt into an intermediate layer of the image encoder of the base model, rather than using it as a direct image input. This can be done to make the training and freezing of the base model more stable and sample - efficient.
[0051] Given a fixed - text prompt T prompt and the learned image - based prompt representation Z prompt , the base model produces a multimodal base - model representation FOUND repr :
[0052] FOUND: Z prompt ×T prompt →FFOUND repr
[0053] Depending on the application, FOUND repr can be a representation from the encoder of the base model, an intermediate representation from the decoder of the base model, a representation from one of the output prediction heads of the base model, a combination of multiple base - model prediction heads (if available), or some combination of the above.
[0054] Figure 3A Illustrates the architecture of a system configured to perform transfer semantic segmentation via self - supervision and prompting of a multimodal base model. A fixed - text prompt 301 can be fed into the base model 303. The base module can be frozen. The input can also include an RGB image 303. The RGB image 303 can be fed into a learnable image - prompt network 305. As previously mentioned, the image - prompt network can be a learnable mapping function - from the input to a latent vector space of a certain fixed dimension. The input RGB image 303 can also be fed into a task encoder 307.
[0055] The task encoder 307 can output information to be sent to a fusion model 313. The fusion mechanism 313F αOne objective can be to combine a more general representation from the base model 311 with a task-specific representation Z φ from the task-specific encoder module E enc This combination results in a framework that exhibits the best of both worlds - task awareness (with strong in-domain performance) and strong cross-domain generalization. F α Take the task-specific encoder representation Z enc and the prompt multimodal base model representation FOUND repr as inputs:
[0056] F α : [Z enc ; FOUND repr → Z att ,
[0057] where the operator ";" indicates vector concatenation. Without loss of generality, the system and method include other suitable forms of vector combination or regression strategies, such as weighted Hadamard product, inner product, outer (cross) product, cross-modal attention, pairwise cross-modal attention, self-attention, distribution alignment constraint, L-norm constraint, etc. F α Output the fused multimodal context Z att , as the input to the task-specific decoder module (next subsection). F α Can be implemented as a function approximator with learnable parameters α, such as a neural network.
[0058] The task-specific decoder G ψ Can be a learnable function mapping that takes as inputs vectors from the intermediate latent representation space Z enc (produced by the task encoder 307E φ ), the (self-) attention fusion context F α 313 and the prompt multimodal base model representation FOUNDrepr from the base model 311. Depending on the application, FOUND repr Can be a representation from the encoder of the base model, an intermediate representation of the decoder of the base model, a representation from one of the output prediction heads of the base model, a combination of representations of multiple base model prediction heads (if available), or some combination of the above. G ψ Maps these inputs to the output label representation space Z P,{s,t} ; this output space can be interpreted as giving an unnormalized probability distribution over the {C s , C t} classes (see the problem definition above). If the task label space Y has more than two dimensions (as in a semantic segmentation task), then Z P,{s,t} Specifies the classes {C s , C tElement - wise set of probability distributions over
[0059] G ψ : [Z enc ; F α ; FOUND repr ⊙ mask → Z P,{st} ,
[0060] where the operator ';' indicates vector concatenation, and '⊙' indicates the Hadamard product (element - wise vector multiplication in the corresponding dimensions). Without loss of generality, the system includes other suitable forms of vector combination or regression strategies in addition to vector concatenation, such as weighted Hadamard product, inner product, outer (cross) product, cross - modal attention, pairwise cross - modal attention, self - attention, distribution alignment constraints, L - norm constraints, etc. In some applications, the decoder may require an additional bias in order to utilize the representation FOUND of the base model when it is being trained repr ; for such a case, the system can include a multiplicative vector mask on the input to decoder 315G ψ . The mask itself can be a learnable filter or just a static representation. The overall job of decoder 315G ψ can be to produce a latent vector from which the correct label for a downstream task can be sampled with high probability or otherwise generated or selected. G ψ can be implemented as a function approximator with learnable parameters ψ, such as a neural network
[0061] Figure 3B Illustrates the architecture of a system configured to perform transfer semantic segmentation via self - supervision from a multimodal base model. In such an embodiment, the system may lack a learnable image prompt network 305. Thus, the RGB image 301 can be sent directly to the base model 311. The task encoder 307 can receive the input image 303, which, in one non - limiting example, can be an RGB image. The task encoder 307 can encode the image to be sent to the task decoder 313 for decoding of the representation. Additionally, the base model 311 representation can be sent to the decoder 313 for decoding. The decoder 315 can produce a latent vector from which the correct label for a downstream task can be sampled with high probability or otherwise generated or selected. The label selection 317 processor module can utilize the latent vector
[0062] Figure 3C Illustrates the architecture of a system configured to perform transfer semantic segmentation via self - supervision from a multimodal base model. In such an embodiment, as with Figure 3B and 3CCompared with the embodiments disclosed in [reference], the system may lack a learnable image prompt network 305 and a task encoder. Therefore, both the fixed text prompt 301 and the RGB image 301 can be directly sent to the base model 311. The fixed text prompt can include both a task (e.g., "semantic segmentation") and objects (e.g., cars, roads, sky, etc.). The task decoder 313 can work with the base model 311 to provide an appropriate output 319, which, in one example, can include an image 319 with semantic segmentation.
[0063] Figure 4 FIG. illustrates an overview of a system for an application scenario according to one embodiment. In such an embodiment, the system can include a general input 401. The input 401 can include images, text, etc. The input 401 can be fed into a multimodal encoder system 403. The multimodal encoder system can be Figures 3A to 3C one of the systems disclosed in [reference], but is not limited to such systems. In one embodiment, the system can use such task-specific inputs 405a, 405b, 405c, 405d for various applications, as discussed further below. For various applications, the system can utilize downstream task encoders 407a, 407b, 407c, 407d to create outputs 407a, 407b, 407c, 407d. For example, in an autonomous driving strategy, the system can output navigation actions, trajectories, waypoints, etc. based on inputs related to autonomous driving. In another example, for a robot navigation strategy, the system can receive task-specific inputs related to the robot and output navigation actions, trajectories, waypoints, etc. In another example, it can include an embodied question-answering system that can include task-specific inputs and whose output is navigation actions, trajectories, waypoints in addition to natural language predictions, conversations, etc. In yet another example, the system can support a visual question-answering decision support system with natural language predictions, conversations, etc.
[0064] The machine learning models described herein can be used in many different applications, not just in the context of road sign image processing. Figures 6 - 11 Additional applications where vision-based prediction tasks can be used are shown in [reference]. Figure 5 The structure of the machine learning model (or base model) for training and using these applications (and other applications) is illustrated in [reference]. Figure 5A schematic diagram depicting the interaction between a computer-controlled machine 500 and a control system 502. The computer-controlled machine 500 includes an actuator 504 and a sensor 506. The actuator 504 may include one or more actuators, and the sensor 506 may include one or more sensors. The sensor 506 is configured to sense the condition of the computer-controlled machine 500. The sensor 506 may be configured to encode the sensed condition into a sensor signal 508 and transmit the sensor signal 508 to the control system 502. Non-limiting examples of the sensor 506 include video, radar, LiDAR, ultrasonic, and motion sensors. In one embodiment, the sensor 506 is an optical sensor configured to sense an optical image of the environment near the computer-controlled machine 500.
[0065] The control system 502 is configured to receive the sensor signal 508 from the computer-controlled machine 500. As described below, the control system 502 may further be configured to calculate an actuator control command 510 depending on the sensor signal and transmit the actuator control command 510 to the actuator 504 of the computer-controlled machine 500.
[0066] As Figure 5 shown, the control system 502 includes a receiving unit 512. The receiving unit 512 may be configured to receive the sensor signal 508 from the sensor 506 and transform the sensor signal 508 into an input signal x. In an alternative embodiment, the sensor signal 508 is directly received as the input signal x without the receiving unit 512. Each input signal x may be a part of each sensor signal 508. The receiving unit 512 may be configured to process each sensor signal 508 to generate each input signal x. The input signal x may include data corresponding to the image recorded by the sensor 506.
[0067] The control system 502 includes a classifier 514. The classifier 514 can be configured to classify an input signal x into one or more labels using a machine learning (ML) algorithm, such as the neural network described above. The classifier 514 is configured to be parameterized by parameters, such as those described above (e.g., parameter θ). The parameters θ can be stored in and provided by a non-volatile storage device 516. The classifier 514 is configured to determine an output signal y from the input signal x. Each output signal y includes information assigning one or more labels to each input signal x. The classifier 514 can transmit the output signal y to a conversion unit 518. The conversion unit 518 is configured to convert the output signal y into an actuator control command 510. The control system 502 is configured to transmit the actuator control command 510 to an actuator 504, which is configured to actuate a computer-controlled machine 500 in response to the actuator control command 510. In another embodiment, the actuator 504 is configured to actuate the computer-controlled machine 500 directly based on the output signal y.
[0068] When the actuator 504 receives the actuator control command 510, the actuator 504 is configured to perform an action corresponding to the associated actuator control command 510. The actuator 504 can include control logic configured to transform the actuator control command 510 into a second actuator control command for controlling the actuator 504. In one or more embodiments, instead of or in addition to an actuator, the actuator control command 510 can be used to control a display.
[0069] In another embodiment, instead of or in addition to the computer-controlled machine 500 including a sensor 506, the control system 502 includes a sensor 506. Instead of or in addition to the computer-controlled machine 500 including an actuator 504, the control system 502 can also include an actuator 504.
[0070] As Figure 5 shown, the control system 502 also includes a processor 520 and a memory 522. The processor 520 can include one or more processors. The memory 522 can include one or more memory devices. The classifier 514 of one or more embodiments (e.g., a machine learning algorithm, such as those described above with respect to the pre-trained classifier 306) can be implemented by the control system 502, which includes the non-volatile storage device 516, the processor 520, and the memory 522.
[0071] The non-volatile storage device 516 may include one or more permanent data storage devices, such as hard disk drives, optical disk drives, tape drives, non-volatile solid-state devices, cloud storage devices, or any other device capable of permanently storing information. The processor 520 may include one or more devices selected from high-performance computing (HPC) systems, including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units, field-programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other device that manipulates signals (analog or digital) based on computer-executable instructions residing in the memory 522. The memory 522 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information.
[0072] The processor 520 may be configured to read into the memory 522 and execute computer-executable instructions residing in the non-volatile storage device 516 and embody one or more ML algorithms and / or method techniques of one or more embodiments. The non-volatile storage device 516 may include one or more operating systems and applications. The non-volatile storage device 516 may store computer programs compiled and / or interpreted from those created using various programming languages and / or technologies, including but not limited to Java, C, C++, C#, Objective C, Fortran, Pascal, Java Script, Python, Perl, and PL / SQL, either alone or in combination.
[0073] When executed by the processor 520, the computer-executable instructions of the non-volatile storage device 516 may cause the control system 502 to implement one or more ML algorithms and / or method techniques as disclosed herein. The non-volatile storage device 516 may also include ML data (including data parameters) that support the functions, features, and processes of one or more embodiments described herein.
[0074] The program code embodying the algorithms and / or method techniques described herein can be distributed, alone or in combination, in various different forms as a program product. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of one or more embodiments. A computer-readable storage medium that is non-transitory in nature can include volatile and non-volatile, removable and non-removable tangible media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium can also include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid state storage technologies, portable compact disc read-only memory (CD-ROM) or other optical storage devices, magnetic tape cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be read by a computer. The computer-readable program instructions can be downloaded from a computer-readable storage medium to a computer, another type of programmable data processing apparatus, or another device, or downloaded to an external computer or external storage device via a network.
[0075] The computer-readable program instructions stored in a computer-readable medium can be used to cause a computer, other type of programmable data processing apparatus, or other device to operate in a particular manner such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions, acts, and / or operations specified in a flowchart or diagram. In certain alternative embodiments, consistent with one or more embodiments, the functions, acts, and / or operations specified in a flowchart and diagram can be reordered, processed serially, and / or processed concurrently. Additionally, any flowchart and / or diagram can include more or fewer nodes or blocks than illustrated consistent with one or more embodiments.
[0076] A process, method, or algorithm can be embodied, in whole or in part, using suitable hardware components such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices, or combinations of hardware, software, and firmware components.
[0077] Figure 6FIG. depicts a schematic diagram of a control system 502 configured to control a vehicle 600, which may be at least partially autonomous or at least partially autonomous robot. The vehicle 600 includes actuators 504 and sensors 506. The sensors 506 may include one or more video sensors, cameras, radar sensors, ultrasonic sensors, LiDAR sensors, and / or position sensors (e.g., GPS). One or more of the one or more specific sensors may be integrated into the vehicle 600. In the context of sign recognition and processing as described herein, the sensor 506 is a camera mounted to or integrated into the vehicle 600. Alternatively or in addition to the one or more specific sensors described above, the sensor 506 may include a software module configured to determine the state of the actuators 504 when executed. A non-limiting example of a software module includes an autonomous driving strategy that may provide navigation actions, trajectories, waypoints, or other items associated with the vehicle 600 or other locations.
[0078] The classifier 514 of the control system 502 of the vehicle 600 may be configured to detect an object near the vehicle 600 depending on the input signal x. In such an embodiment, the output signal y may include information characterizing the object near the vehicle 600. The actuator control command 510 may be determined based on this information. The actuator control command 510 may be used to avoid a collision with the detected object.
[0079] In an embodiment where the vehicle 600 is at least partially autonomous vehicle, the actuators 504 may be embodied in the brakes, propulsion system, engine, driveline, or steering of the vehicle 600. The actuator control command 510 may be determined to control the actuators 504 such that the vehicle 600 avoids a collision with the detected object. The detected objects may also be classified according to what the classifier 514 thinks they are most likely to be, such as pedestrians or trees. The actuator control command 510 may be determined depending on the classification. In scenarios where adversarial attacks may occur, the above system may be further trained to better detect objects or identify changes in the lighting conditions or angles of the sensors or cameras on the vehicle 600.
[0080] In other embodiments where the vehicle 600 is at least partially autonomous robot, the vehicle 600 may be a mobile robot configured to perform one or more functions such as flying, swimming, diving, and stepping. The mobile robot may be at least partially autonomous lawn mower or at least partially autonomous cleaning robot. In such an embodiment, the actuator control command 510 may be determined such that the propulsion unit, steering unit, and / or braking unit of the mobile robot may be controlled such that the mobile robot may avoid a collision with the identified object.
[0081] In another embodiment, vehicle 600 is at least partially autonomous robot in the form of a gardening robot. In such an embodiment, vehicle 600 may use an optical sensor as sensor 506 to determine the state of plants in the environment near vehicle 600. Actuator 504 may be a nozzle configured to spray a chemical substance. Depending on the type and / or status of the identified plant, an actuator control command 510 may be determined to cause actuator 504 to spray an appropriate amount of an appropriate chemical substance onto the plant.
[0082] Vehicle 600 may be at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include washing machines, stoves, ovens, microwave ovens, or dishwashers. In such a vehicle 600, sensor 506 may be an optical sensor configured to detect the state of an object to be processed by the household appliance. For example, in the case where the household appliance is a washing machine, sensor 506 may detect the state of the laundry inside the washing machine. An actuator control command 510 may be determined based on the detected state of the laundry.
[0083] Figure 7 A schematic diagram of control system 502 is depicted. Control system 502 is configured to control system 700 (e.g., a manufacturing machine), such as a stamping machine, a cutting machine, or a gun drill, which is part of manufacturing system 702 (e.g., a production line). Control system 502 may be configured to control actuator 504, which is configured to control system 700 (e.g., a manufacturing machine).
[0084] Sensor 506 of system 700 (e.g., a manufacturing machine) may be an optical sensor configured to capture one or more attributes of manufactured product 704. Classifier 514 may be configured to determine the state of manufactured product 704 based on the one or more captured attributes. Actuator 504 may be configured to control system 700 (e.g., a manufacturing machine) depending on the determined state of manufactured product 704 for subsequent manufacturing steps of manufactured product 704. Actuator 504 may be configured to control the function of system 700 (e.g., a manufacturing machine) on subsequent manufactured product 106 of system 700 (e.g., a manufacturing machine) depending on the determined state of manufactured product 704.
[0085] Figure 8 A schematic diagram of control system 502 is depicted. Control system 502 is configured to control power tool 800 having at least a partially autonomous mode, such as a drill or a driver. Control system 502 may be configured to control actuator 504, which is configured to control power tool 800.
[0086] The sensor 506 of the power tool 800 can be an optical sensor configured to capture one or more attributes of the work surface 802 and / or the fastener 804 driven into the work surface 802. The classifier 514 can be configured to determine the state of the work surface 802 and / or the fastener 804 relative to the work surface 802 based on one or more captured attributes. The state can be that the fastener 804 is flush with the work surface 802. Alternatively, the state can be the hardness of the work surface 802. The actuator 504 can be configured to control the power tool 800 such that the driving function of the power tool 800 is adjusted depending on the determined state of the fastener 804 relative to the work surface 802 or one or more captured attributes of the work surface 802. For example, if the state of the fastener 804 is flush with the work surface 802, the actuator 504 can interrupt the driving function. As another non-limiting example, the actuator 504 can apply additional or less torque depending on the hardness of the work surface 802.
[0087] Figure 9 A schematic diagram depicting a control system 502 configured to control an automated personal assistant 900 is shown. The control system 502 can be configured to control an actuator 504, which is configured to control the automated personal assistant 900. The automated personal assistant 900 can be configured to control household appliances such as a washing machine, a stove, an oven, a microwave oven, or a dishwasher.
[0088] The sensor 506 can be an optical sensor and / or an audio sensor. The optical sensor can be configured to receive a video image of the pose 904 of the user 902. The audio sensor can be configured to receive a voice command of the user 902.
[0089] The control system 502 of the automated personal assistant 900 can be configured to determine an actuator control command 510 configured to control the control system 502. The control system 502 can be configured to determine the actuator control command 510 based on the sensor signal 508 of the sensor 506. The automated personal assistant 900 is configured to transmit the sensor signal 508 to the control system 502. The classifier 514 of the control system 502 can be configured to execute a pose recognition algorithm to identify the pose 904 made by the user 902, determine the actuator control command 510, and transmit the actuator control command 510 to the actuator 504. The classifier 514 can be configured to retrieve information from the non-volatile storage device in response to the pose 904 and output the retrieved information in a form suitable for the user 902 to receive.
[0090] Figure 10A schematic diagram depicting a control system 502 configured to control a surveillance system 1000 is shown. The surveillance system 1000 can be configured to physically control access through a door 1002. A sensor 506 can be configured to detect a scenario relevant to determining whether to grant access. The sensor 506 can be an optical sensor configured to generate and transmit image and / or video data. The control system 502 can use such data to detect a human face.
[0091] A classifier 514 of the control system 502 of the surveillance system 1000 can be configured to interpret the image and / or video data by matching the identities of known persons stored in a non-volatile storage device 516 to determine the identity of a person. The classifier 514 can be configured to generate an actuator control command 510 in response to the interpretation of the image and / or video data. The control system 502 is configured to transmit the actuator control command 510 to an actuator 504. In this embodiment, the actuator 504 can be configured to lock or unlock the door 1002 in response to the actuator control command 510. In other embodiments, non-physical logical access control is also possible.
[0092] The surveillance system 1000 can also be a monitoring system. In such an embodiment, the sensor 506 can be an optical sensor configured to detect a scenario under monitoring, and the control system 502 is configured to control a display 1004. The classifier 514 is configured to determine the classification of the scenario, such as whether the scenario detected by the sensor 506 is suspicious. The control system 502 is configured to transmit the actuator control command 510 to the display 1004 in response to the classification. The display 1004 can be configured to adjust the displayed content in response to the actuator control command 510. For example, the display 1004 can highlight an object considered suspicious by the classifier 514. With an embodiment of the disclosed system, a monitoring system can use semantic segmentation to highlight certain objects or suspicious activities.
[0093] Figure 11 A schematic diagram of a control system 502 is depicted, which is configured to control an imaging system 1100, such as an MRI device, an x-ray imaging device, or an ultrasound device. The sensor 506 can be an imaging sensor, for example. The classifier 514 can be configured to determine the classification of all or part of a sensed image. The classifier 514 can be configured to determine or select an actuator control command 510 in response to the classification obtained by a trained neural network. For example, the classifier 514 can interpret a region of the sensed image as a potential anomaly. In such a case, an actuator control command 510 can be determined or selected to cause the display 1102 to display the imaging and highlight the potential anomaly region.
[0094] While the exemplary embodiments have been described above, it is not intended that these embodiments describe all possible forms covered by the claims. The words used in the specification are descriptive rather than restrictive, and it is understood that various changes can be made without departing from the spirit and scope of the disclosure. As previously mentioned, the features of the various embodiments can be combined to form additional embodiments of the invention that may not be explicitly described or illustrated. Although the various embodiments may have been described as providing advantages over other embodiments or prior art implementations in one or more desired characteristics or being preferred to other embodiments or prior art implementations, one of ordinary skill in the art recognizes that one or more features or characteristics may be compromised to achieve the desired overall system attributes, depending on the particular application and implementation. These attributes can include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, suitability, weight, manufacturability, ease of assembly, etc. Accordingly, to the extent that any embodiment is described as less desirable than other embodiments or prior art implementations in one or more characteristics, these embodiments are not outside the scope of the disclosure and may be desirable for a particular application.
Claims
1. A computer-implemented method comprising: receiving one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images; outputting an intermediate representation from generating a sequence of objects and tasks in response to utilizing the fixed text prompt and the one or more images at a base model associated with a machine learning network; decoding the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with a fixed textual prompt and an image; as well as In response to identifying a highest probability associated with the matrix using the label selection, a final label associated with the vision-based prediction task is output.
2. The computer-implemented method of claim 1, wherein the base model is a multimodal model.
3. The computer-implemented method of claim 1 , wherein the method comprises utilizing the one or more images at an image hinting network to generate a continuous latent vector of fixed dimension.
4. The computer-implemented method of claim 1, wherein the decoder utilizes a task-specific decoder.
5. The computer-implemented method of claim 4, wherein the method comprises utilizing a task-specific encoder.
6. The computer-implemented method of claim 1, wherein the decoder is a task-specific decoder comprising a learnable function mapping that utilizes input from an intermediate latent representation space of an encoder as a vector.
7. The computer-implemented method of claim 1, wherein the vision-based prediction task is not included in pre-training of the base model.
8. The computer-implemented method of claim 1, wherein the vision-based prediction task is semantic segmentation and the final label comprises a semantically segmented image.
9. A method comprising: receiving one or more fixed text prompts at a base model; receiving, at a learnable image prompting network, one or more images, wherein the fixed text prompt is associated with the one or more images; generating a continuous latent vector of fixed dimension at a learnable image cueing network using the one or more images; outputting intermediate representations from generating a sequence of objects and tasks in response to utilizing a fixed textual cue and a continuous latent vector of fixed dimensionality at a base model associated with a machine learning network; decoding the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with a fixed textual prompt and an image; as well as In response to identifying a highest probability associated with the matrix using the label selection, a final label associated with the vision-based prediction task is output.
10. The method of claim 9, wherein the method further comprises combining the representation from the base model with the task-specific representation from the encoder using a fusion model.
11. The method of claim 9, wherein the vision-based prediction task is semantic segmentation and the final label comprises a semantically segmented image.
12. The method of claim 9, wherein the vision-based prediction task is not included in the pre-training of the base model. The method of claim 9 , wherein the task is a single task.
14. The system of claim 9, wherein the base model is configured to output an intermediate representation from one of a base model encoder, a base model decoder, an output prediction head associated with the base model, or a representation of a combination of multiple base model prediction heads.
15. A system comprising: The controller is configured as: receiving one or more fixed text prompts and one or more images, wherein the fixed text prompts are associated with the one or more images; outputting an intermediate representation from generating a sequence of objects and tasks in response to utilizing the fixed text prompt and the one or more images at a base model associated with a machine learning network; decoding the intermediate representation using a decoder associated with the base model to generate a matrix associated with a task associated with a fixed textual prompt and an image; as well as In response to identifying a highest probability associated with the matrix using the label selection, a final label associated with the vision-based prediction task is output.
16. The system of claim 15, wherein the vision-based prediction task is semantic segmentation and the final label comprises a semantically segmented image.
17. The system of claim 15, wherein the vision-based prediction task is not included in pre-training of the base model.
18. The system of claim 15, wherein the base model comprises a language interface and a visual interface.
19. The system of claim 15, wherein the system comprises a task encoder configured to send one or more visual representations to a task decoder.
20. The system of claim 15, wherein the one or more images include a red-green-blue (RGB) image, an audio image, a video image, or a radar image.