Automated media content identification for understanding multimedia

By automatically generating custom computational code for visual language models from large language models, the problem of low efficiency in manually writing code in existing technologies is solved, and fast, low-cost multi-task processing is achieved.

CN121008780APending Publication Date: 2025-11-25NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510666758.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-22
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies require expert coders to manually write custom computational code for visual language models, resulting in low efficiency, high cost, and difficulty in handling various tasks.

Method used

By using large language models (LM) to generate custom computational codes, media content can be automatically identified and corresponding computational codes can be generated. By combining visual language models and content detection models, automated content recognition without expert coding can be achieved.

Benefits of technology

It improves the application speed and versatility of visual language systems, making them easier for non-professional users to use, reducing the need for manual coding, and lowering development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008780A_ABST
    Figure CN121008780A_ABST
Patent Text Reader

Abstract

The invention relates to automated media content identification for understanding multimedia. Apparatuses, systems, and techniques are disclosed for automated content recognition that generates custom compute code using a language model. The techniques include obtaining a first prompt for describing a media item; the media item and the representation of the first cue are processed using a content detection model to obtain a representation of the media item. The techniques also include generating a second cue using the first cue and the characterization of the media item, the second cue including instructions for a language model (LM). The techniques further include causing the LM to process a second hint to generate a compute code associated with the characterization of the media item; and causing the compute code to be executed to generate a responsiveness description for the media item.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] At least one embodiment relates to content generation using artificial intelligence (AI) systems. For example, at least one embodiment relates to AI systems and techniques for recognizing media content and representing an understanding of the recognized content in natural language form. BACKGROUND

[0002] Well-trained language models (e.g., large language models (LLMs)) are capable of supporting conversations in natural language, understanding a speaker’s intent and emotion, explaining complex topics, generating new text upon receiving suitable prompts, providing recommendations about topics of interest to a user, processing images, audio, and / or other data types, and / or performing other functions. LLMs are typically self-supervised trained on large amounts of text data and / or other data types according to an embodiment, and learn to predict the next token in a phrase / sentence and / or missing tokens (which can correspond to subwords, symbols, words, etc.), detect a human speaker’s intent and / or emotion, determine whether two sentences are related or unrelated, and / or perform other basic language tasks. After initial training, LLMs are typically subjected to guided (prompt-based) supervised fine-tuning, which enables the LLMs to gain deeper language proficiency and / or master more specialized tasks. Supervised fine-tuning includes using learning prompts (questions, hints, etc.) that are accompanied by example text (e.g., answers, example articles, etc.) as training ground truth. In reinforced fine-tuning, human evaluators assign a rating that indicates how similar the generated text is to human-produced text. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 is a block diagram of an example computer architecture that is capable of automated content recognition and custom computational code generation using language models, according to at least one embodiment;

[0004] Figure 2 An example computing device that supports a system that deploys custom computational code generation using language models to facilitate automated content recognition is shown, according to at least one embodiment;

[0005] Figure 3 An example data flow of automated content recognition using custom computational code generation using language models is shown, according to at least one embodiment;

[0006] Figure 4 An example architecture of an open-vocabulary content detection model that can be used for automated content recognition through custom computational code generation is shown, according to at least one embodiment;

[0007] Figures 5A-5BContent detection that can be used with custom compute code generation, according to at least one embodiment, is illustratively shown;

[0008] Figure 6 is a flow diagram of an example method of automated content recognition with language model generation of custom compute code, according to at least one embodiment;

[0009] Figure 7A shows inference and / or training logic, according to at least one embodiment;

[0010] Figure 7B shows inference and / or training logic, according to at least one embodiment;

[0011] Figure 8 shows training and deployment of a neural network, according to at least one embodiment;

[0012] Figure 9 is an example dataflow graph of a high-level compute pipeline, according to at least one embodiment; and

[0013] Figure 10 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in a high-level compute pipeline, according to at least one embodiment. DETAILED DESCRIPTION

[0014] Computer vision AI provides computers with the ability to detect various objects of interest (e.g., people, animals, cars, etc.), actions and events (e.g., sports actions, game actions), certain expected or unexpected behaviors and / or situations (e.g., traffic jams, unsafe or poor manufacturing conditions), etc., in images and videos. Computer vision (CV) automates tasks that would traditionally be performed by a human observer. The output of a CV model can include localization of objects (e.g., using bounding boxes or other segmentation techniques), classification of objects (e.g., by type or class in a number of classes learned in training), confidence in the localization / classification obtained, etc. Such output can be used by downstream systems, such as an on-board planner for an autonomous vehicle.

[0015] A visual language model combines the functionality of a CV with that of a language model in order to perform natural language (NL) understanding on data and / or tasks to be performed on data. For example, a prompt to a visual language model (e.g., “Is the animal in the image a cat or a dog?”) can call for classification of an object, and the output can contain a set of classes and probabilities (e.g., “cat, 0.8; dog, 0.1”). In some cases (e.g., as in the above example), the output can be readily understandable to humans, including laypersons. In more complex cases (e.g., “How many pedestrians are crossing the street in this video?”), the user can need to perform some additional analysis on the output, e.g., filtering various bounding boxes and classifications generated by the model and manually identifying relevant objects. Alternatively, a developer can write custom post-processing code that parses the output and extracts actionable data to generate a response to the user’s query (e.g., “Five pedestrians are crossing the street”). Writing such code requires at least some coding experience, which most non-experts lack, or can be impractical or cost-prohibitive in cases where multiple different tasks need to be handled. For example, code for extracting information about “How many bags did the passenger wearing a blue shirt carry?” can be quite different from code created to determine whether the positioning and location of cars in an image indicate that a traffic accident occurred, and thus needs to be written separately.

[0016] Aspects and embodiments of the present disclosure address these and other challenges of visual language technology by providing systems and techniques that automate content recognition and generate custom computing code using language models, without the need for expert coders. In some embodiments, keyword extraction can be performed on a user prompt (query, question, etc.) to identify the type of target content in the media data (e.g., objects present in an image / video or audio file). Based on the identified keywords, a content detection model (e.g., a CV model) can be selected from a library of available models. In those cases where one or more models are available to detect a particular target content of interest (e.g., a trained pedestrian detection model), such specialized models can be used. In those cases where there are no trained models available to detect the target content, one or more open-vocabulary content detection models can be selected. Open-vocabulary models can contain a media processing portion (e.g., an object detection portion) and a pre-trained language understanding portion that together are capable of detecting content that was not previously encountered by the media processing portion (e.g., in training). Specifically, open-vocabulary models leverage their language understanding capabilities to identify features of previously unseen objects. For example, a media processing portion can have never encountered an image of a lion, but the language understanding portion can have read a large amount of text describing lions, including information that lions are large felines, have large heads, rounded ears, a tawny body color, adult males typically have a thick mane, and / or other information. The linkage between these two portions of the model enables the language descriptions of the features of the target object to propagate to the visual neurons of the model and facilitate recognition of the unfamiliar target object.

[0017] The content detection model can generate information related to the target content, such as bounding boxes and object types of the target objects referenced in the model input (e.g., keywords derived from the user prompt). The content detection output and code writing instructions (e.g., “write a script to count the number of times Y appears in [content detection output]”) can then be used to augment the user prompt. The augmented prompt can also include instructions (explanations) on how to understand the format of the detection output. The augmented prompt can then be used as input to a teaching language model (LM) trained to generate computation code. The teaching LM processes the received prompt and generates code (e.g., Python code, C++ code, JavaScript code, etc.) that is capable of extracting the target information requested in the user prompt and contained in the content detection model output. For example, the LM can assemble a list of operators that includes instructions on how to obtain data from the model output and computation instructions on how to process the obtained data. The source code generated by the teaching LM can be compiled, e.g., using a suitable compiler to convert the source code to machine code, and then executed to generate a response. The output of the code can include natural language phrases or sentences, e.g., “X pedestrians are crossing the street,” where the value of X is computed and replaced during the execution of the code.

[0018] Advantages of the disclosed embodiments include, but are not limited to, eliminating the task of manually writing code by shifting this task to a teaching coding LM. The deployed technology allows developers to provide a single set of descriptions for each content detection model to inform the teaching LM how to read the output of the model. The prompt to the teaching LM can then be fully automated according to the received user prompt without the need for manually generating separate task-related code. The disclosed technology improves the speed and versatility of applications of visual language systems and facilitates the use of such systems by non-expert users.

[0019] Figure 1 is a block diagram of an example computer architecture 100 capable of automated content recognition and custom computation code generation using a language model in accordance with at least one embodiment. As Figure 1 shown, the computer architecture 100 can include a user device 102, a media content recognition (MCR) server 110, a LM service 130, a model library 150, a training server 160, any one, some, or all of which can be connected through a network 140. The network 140 can be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wireless network, a personal area network (PAN), combinations thereof, and / or other network types.

[0020] The user device 102 can include a desktop computer, a laptop computer, a smartphone, a tablet computer, a server, a wearable device, a virtual / augmented / mixed reality head-mounted device or heads-up display, a digital avatar or chatbot kiosk, an in-vehicle infotainment computing device, and / or any other suitable computing device capable of performing the techniques described herein. The user device 102 can be configured to communicate with the user 101 through a user interface (UI) 104. The user 101 can be an individual user (e.g., an owner of a computer, a vehicle, an entertainment device), a collective user (e.g., a business organization, an institution, a government agency, etc.), an agent of a repair facility, etc. In some embodiments, the prompt generated by the user 101 can include text (e.g., a sequence of one or more typed words), speech (e.g., a sequence of one or more spoken words), an image, a gesture, and / or a combination thereof. The prompt can be generated as part of the user 101 interacting with the MCR server 110 using the LM service 130.

[0021] The UI 104 can include one or more devices of various modalities, such as a keyboard, a touchscreen, a touchpad, a handwriting pad, a graphical interface, a mouse, a stylus, and / or any other pointer device capable of selecting a word / phrase displayed on a screen, and / or other suitable devices. In some embodiments, the UI 104 can include an audio device (e.g., a combination of a microphone and a speaker), a video device (such as a digital camera for capturing an image or a sequence of two or more images (video frames)), or both. In some embodiments, the text, speech, and / or video input devices can be integrated together (e.g., integrated into a smartphone, a tablet computer, a desktop computer, etc. device).

[0022] In some embodiments, the MCR server 110 can be located on one or more computing devices / servers, such as cloud-based servers. The user device 102 can download the MCR application programming interface (API) 106 from the MCR server 110 and deploy the MCR API 106 to facilitate communication with the MCR server 110. The MCR server 110 can perform processing of prompts generated by the user 101. These prompts can be natural language prompts that involve instructing the MCR server 110 to detect and analyze content of one or more media items 108 (e.g., provided by the user 101). The media items 108 can include images, videos (e.g., sequences of images / frames that are related in time, visually, and / or contextually), audio, and / or any other data items produced by suitable sensors, including but not limited to lidar sensors, radar sensors, infrared camera sensors, temperature sensors, pressure sensors, and / or any other physical or chemical sensors. The MCR server 110 can process the media items 108, e.g., as directed by the user prompts. In some embodiments, the LM 132 provided by the LM service 130 can facilitate the processing of the media items 108 by the MCR server 110.

[0023] In some embodiments, the MCR server 110 can include a memory 112 (e.g., one or more memory devices or units) that is communicably coupled to one or more processing devices, such as one or more central processing units (CPUs) 114, one or more graphics processing units (GPUs) 116, one or more data processing units (DPUs), one or more parallel processing units (PPUs), and / or other processing devices (e.g., field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.). The memory 112 can include read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), static memory such as static random access memory (SRAM), and / or some other memory capable of storing digital data. The memory 112 can store one or more content detection models 120 trained to detect and / or classify content of the media items 108, a LM prompt generation module 122 for generating prompt requests to the LM 132 to author computational code capable of processing outputs of the content detection models, a LM API 124 for facilitating communication with the LM 132, and an MCR code execution module 126 for executing (and, where applicable, compiling) code generated by the LM 132. The MCR server 110 can also support Figure 1any number of additional components and modules not explicitly shown, such as any application capable of generating, displaying processing, editing, and / or otherwise using textual data, audio data, image data, video data, and the like. In some embodiments, the MCR server 110 can also be operated by the LM service 130. Although depicted as separate from the LM service 130 in Figure 1 In some embodiments, the MCR server 110 can host the LM 132, although it is depicted as separate from the LM service 130.

[0024] In some embodiments, the LM 132 can be a large language model, e.g., a model having at least 100,000 learnable parameters, provided by the LM service 130. The LM 132 can be trained by the LM training engine 134. In some embodiments, the LM 132 can be a model pre-trained and deployed by a separate entity. In some embodiments, the LM 132 can be trained in multiple stages. First, the LM training engine 134 can train the LM 132 to capture the syntax and semantics of human language, e.g., by training to predict the next word, the previous word, and / or the missing word in a sequence of words (e.g., one or more sentences of human speech or text). The LM 132 can be trained using training data containing a large amount of text, such as human conversations, newspaper text, magazine text, book text, web-based text, and / or any other text. Since the ground truth (e.g., the next word) for such training is embedded in the text itself, the LM training engine 164 can use these texts to perform self-supervised training of the LM 132. This teaches the LM 132 to converse in natural language with a user (a human user or another computer) in a manner very similar to a conversation with a human speaker, including understanding the user’s intent and responding in a manner that the user would expect a conversational partner to do.

[0025] After the initial self-supervised training, the LM training engine 134 can perform supervised fine-tuning or instruction fine-tuning of the LM 132 to teach the LM 132 more specialized skills, including expertise in writing computer code for various computational tasks that can be formulated using natural language. In some embodiments, the LM training engine 134 can facilitate any, some, or all of the training stages of the LM 132. For example, the LM training engine 134 can oversee the self-supervised training stage, focusing on developing general language capabilities, and then pass the pre-trained LM 132 to another entity for additional fine-tuning of the LM 132. In some cases, the LM 132 can receive a pre-trained LM 132 from another entity and fine-tune the LM 132. In some cases, the LM training engine 134 can perform both the pre-training of the LM 132 and the domain-specific fine-tuning of the LM 132.

[0026] The content detection models 120 can be trained to recognize particular target content (possibly mentioned in the prompt) in any relevant input data (e.g., media items 108), e.g., detecting target objects in image / video frames, target words or descriptions in audio files, occurrences of particular conditions in sensor data, etc. The content detection models 120 can be stored in a model library 150 and downloaded and deployed on the MCR server 110. The models available in the model library 150 can include a target content detection model 122-A, which can include a model trained to detect a particular type / category of content of interest, e.g., cars, trucks, buses, pedestrians, bicycles, and / or other objects. Such a model can have a fixed number of output channels associated with the target category. In addition, the models in the model library 150 can store an open-vocabulary content detection model 122-B, which can be trained to detect a number of types of content but also able to detect objects not encountered in training, e.g., by leveraging language understanding capabilities learned from a variety of text containing descriptions of a large number of content items, including many items for which the model has not seen images (or other representations) of.

[0027] In some implementations, the training of the content detection models 120 can be performed by a training server 160. In at least one embodiment, any, some, or all of the content detection models 120 can be implemented as a deep learning neural network with multiple layers of linear or non-linear operations. For example, any, some, or all of the content detection models 120 can include a convolutional neural network, a recurrent neural network, a fully connected neural network, a long short-term memory (LSTM) neural network, a neural network with attention (e.g., a transformer neural network), etc. In at least one embodiment, any, some, or all of the content detection models 120 can include multiple neurons, a single neuron receiving inputs from other neurons and / or from external sources and producing an output by applying an activation function to a sum of the inputs modified by (trainable) weight and bias values. In at least one embodiment, any, some, or all of the content detection models 120 can include multiple neurons arranged in layers, including an input layer, one or more hidden layers, and / or an output layer. Neurons from adjacent layers can be connected by weighted edges. In some embodiments, different content detection models can have different architectures, numbers of layers of neurons, numbers of neurons in each layer, etc.

[0028] Any, some, or all of the content detection models 120 can be trained by a training engine 162 hosted by a training server 160, which can be (or include) a desktop computer, a laptop computer, a smartphone, a tablet computer, a server, and / or any suitable computing device capable of performing the techniques described herein. Training of the target content detection model 122-A can be performed using training data that includes content (e.g., depicted or otherwise represented in images, videos, audio, and / or other related data), which can be labeled with ground truth, which can include correct identifications of target content and / or non-target content. Training of the open-vocabulary detection model 122-B can also include zero-shot training, in which the model is given training cues to identify content (e.g., depictions of objects) that were not encountered in previous training rounds.

[0029] During training, predictions of the suitable model 165 can be compared to ground truth annotations. More specifically, the training engine 162 can cause the model to process training inputs 164 (which can include media items and training cues) and generate training outputs 166 (which represent identifications of content in the respective training inputs 164). During training, the training engine 162 can also generate mapping data 167 (e.g., metadata) that associates the training inputs 164 with correct target outputs 168. The target outputs 168 can include ground truth content identifications for the respective training inputs 164. Training causes the model 165 to identify patterns in the training inputs 164 based on the desired target outputs 168 and learn to accurately classify input data.

[0030] Initially, the edge parameters (e.g., weights and biases) of the model being trained can be assigned some starting values (e.g., random values). For each training input 164, the training engine 162 can compare the training output 166 to the target output 168. The resulting error or mismatch, e.g., the difference between the expected target output 168 and the model’s generated training output 166, can be backpropagated through the model, and at least some of the model’s parameters can be changed in a way that brings the training output 166 closer to the target output 168. This adjustment can be repeated until the output error for a given training input 164 meets a predetermined condition (e.g., is below a predetermined error). Subsequently, a different training input 164 can be selected, a new training output 166 generated, and a new series of adjustments implemented until the model is trained to a target accuracy or until the model converges to its limit of accuracy (determined by its architecture).

[0031] The training server 160 can train any number of content detection models in this (or a similar) manner, using different sets of training inputs 164 and target outputs 168. The trained content detection models can be deployed on any suitable machine, such as the MCR server 110. The trained content detection models can be stored in the model library 150 and downloaded to the MCR server 110. After download at the MCR server 110, these models can be deployed for inference, such as automated identification of content in media items 108, as disclosed in greater detail below.

[0032] Figure 2 An example computing device 200 is shown that supports a system that facilitates deployment of automated content identification using language models to generate custom computing code, in accordance with at least one embodiment. In at least one embodiment, the computing device 200 can be part of the MCR server 110 and / or part of the user device 102 (refer to Figure 1 ). In at least one embodiment, the computing device 200 can deploy an MCR API 206 (which can be a server counterpart of the MCR API 106 running on the user device 102, as shown in Figure 1 ). As shown in Figure 2 , the automated content identification pipeline can include receiving a prompt 202 and a media item 204 associated with the prompt 202, processing the prompt 202 and the media item 204 using a content detection stage 210 to obtain a description of the media item 204 (which can be responsive to the scope of the prompt), and performing prompt augmentation 220 using the obtained description. The prompt 202 augmented with an interpretation of the format of the description of the media item 204 can be provided to an LM (e.g., the LM 132 in Figure 1 ) trained to generate computing code through the LM API 124. Code execution 230 can then execute the code produced by the LM to generate a description of the media item 204.

[0033] The operations of MCR API 206, content detection stage 210, hint augmentation 220, LM API 124, code execution 230, and various modules that operate in conjunction with the automated content recognition pipeline and / or other software / firmware instantiated on computing device 200 can be performed using one or more CPUs 114, one or more GPUs 116, one or more parallel processing units (PPUs) or accelerators such as deep learning accelerators, data processing units (DPUs), etc. In at least one embodiment, GPU 116 includes a plurality of cores 211. A single core 211 can be capable of executing a plurality of threads 212. A single core 211 can run multiple threads 212 concurrently (e.g., in parallel). In at least one embodiment, threads 212 can have access to registers 213. Registers 213 can be thread- private registers, access to which is limited to respective threads. Additionally, shared registers 214 can be accessed by one or more (e.g., all) threads of a core 211. In at least one embodiment, individual cores 211 can include schedulers 215 to allocate computing tasks and processes among different threads 212 of a core. Dispatch units 216 can implement scheduled tasks on appropriate threads using correct private registers 213 and shared registers 214. Computing device 200 can include input / output components 217 to facilitate exchange of information with one or more users or developers.

[0034] In at least one embodiment, GPU 116 can have a (high-speed) cache 218, access to which can be shared by multiple cores 211. Additionally, computing device 200 can include GPU memory 219 in which GPU 116 can store intermediate and / or final results (outputs) of various computations performed by GPU 116. After completion of a particular task, GPU 116 (or CPU 114) can move outputs to (main) memory 112. In at least one embodiment, CPU 114 can execute processes that involve serial computing tasks, while GPU 116 can perform tasks that are amenable to parallel processing, such as multiplying inputs of a neural node by weights and adding a bias.

[0035] The systems and methods described herein can be used for various purposes, such as but not limited to machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twin, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, generative AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0036] The disclosed embodiments can be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines, in-vehicle infotainment systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart regional monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems for performing medical operations, systems for performing factory operations, systems for performing analytics operations, systems implemented using edge devices, systems for generating or presenting at least one of augmented reality content, virtual reality content, mixed reality content, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in data centers, systems for performing conversational AI operations, systems for performing optical transport simulation, systems for performing collaborative content creation for 3D assets, systems implementing one or more language models (such as a large language model (LLM) or a visual language model (VLM) that can process text, speech, images, and / or other data types to generate output in one or more formats), systems implemented at least partially using cloud computing resources, systems for performing generative AI operations, and / or other types of systems.

[0037] Figure 3 An example data flow 300 for automated content recognition using language models to generate custom computing code is shown in accordance with at least one embodiment. Figure 3 The illustrated operations can be performed by the MCR server 110 (refer to FIG. 1) in accordance with at least one embodiment. Figure 1) are performed. In some embodiments, these operations include receiving a prompt 302. The prompt 302 can be received from a user, e.g., as part of a real-time conversation, or can be previously generated (and stored) and subsequently retrieved from a memory device (e.g., memory 112 of MCR server 110 or memory of user device 102). The prompt 302 can be associated with a media item 304, which can include an image, a video (e.g., a sequence of images / frames that are temporally, visually, and contextually related), audio, and / or any other data item generated by a suitable sensor, which can include a camera, a video camera, an infrared camera, a microphone, a sonar, a lidar, a radar, and / or any other physical or chemical sensor, e.g., a temperature sensor, a pressure sensor, a humidity sensor, a smoke detection sensor, a chemical composition sensor, a motion detection sensor, an accelerometer, an altitude sensor, a global positioning sensor, etc. The media item 304 can be (or contain) any time-sequential data, e.g., a sequence of video frames. The media item 304 can be associated with the prompt 302. For example, the media item 304 can be explicitly referenced in the prompt 302 (e.g., by specifying a storage location of the media item 304), directly attached (e.g., as a data file) to the prompt 302, implicitly associated with the prompt 302, and / or in any other manner that unambiguously identifies the media item 304.

[0038] The prompt 302 can be a natural language prompt, e.g., any suitable description for the media item 304, which can be (or contain) a quantitative description (e.g., a request for a number of objects in the media item 304 or a particular type), a qualitative or conceptual description of the content of the media item 304 (e.g., whether a red team scored or missed a goal). The prompt 302 can be expressed as (or contain) a question (e.g., “How many players in white-blue jerseys were on the ice before the game stopped?”), an instruction (e.g., “Count the number of players in white-blue jerseys on the ice before the game stopped”), a task (e.g., determine whether the white-blue team had too many players on the ice before the game stopped”), and / or any other suitable form of inquiry. In some embodiments, the prompt 302 can be a text representation of audio data or visual data from a user, or can be retrieved from memory.

[0039] In some embodiments, the prompt 302 can undergo keyword extraction 310 to identify a type of target content in the media item 304. In some embodiments, the keyword extraction 310 can use an LM, e.g., a base model trained to understand human language (but not necessarily trained on certain specialized tasks). In some embodiments, the LM used for keyword extraction 310 can be the same as the model used to generate the computational code (e.g., the encoding LM 330). In some embodiments, the LM used for keyword extraction 310 can be different from the model used to generate the computational code. In some embodiments, the keyword extraction 310 can use any combination of morphological, syntactic, statistical, graph-based methods, and / or any combination thereof to extract keywords from the prompt 302. In some embodiments, the keyword extraction 310 can use a trained machine learning classifier, e.g., a discriminative classifier, a decoder-only classifier, etc.

[0040] The keyword extraction 310 can generate a list of keywords (also referred to as a “base noun group,” etc.) for the prompt 302. For example, the list of keywords can include “player,” “white and blue uniform,” “on ice,” etc. In some cases, the list of keywords can contain one or more words that are not part of the prompt 302 but are semantically close to one or more words in the prompt 302. For example, a list of keywords for a prompt that determines how many puppies a doggy sitter walked in a picture can contain the word “dog,” even though this word is not explicitly contained in the original prompt. In some implementations, the keyword extraction 310 can generate a representation of the prompt 302 that is different from the list of keywords. For example, the representation of the prompt 302 can be (or contain) a list of word embeddings or tokens that can be understood by a machine learning model, generated by a suitable tokenizer (not explicitly shown in FIG. 3). Figure 3

[0041] The representation of the prompt 302 can be used by the model selection stage 320, which selects one or more content detection models 120 to process the media item 304. In some embodiments, the representation of the prompt 302 can be (or contain) the prompt 302 itself. In some embodiments, the representation of the prompt 302 can contain the list of keywords of the prompt 302. In some embodiments, the representation of the prompt 302 can contain a set of tokens or a set of tokens of the keywords of the prompt 302.

[0042] ​The tokens can encode speech units (e.g., words, syllables, etc.) as numbers. In one example of GPT-4 tokens, the word “the” can be represented by the token “280,” the word “import” can be represented by the token “476,” the word “description” can be represented by the token “4097,” and so on. In some embodiments, individual words can be represented by any number of tokens or word transformations. For example, a long word or a word that contains multiple words can be represented by multiple tokens, e.g., one token for representing the beginning portion of the word and another token for representing the middle or end portion of the word. In some cases, even long / compound words can be represented by a single token. Thus, tokenization can be performed in any manner suitable for input to a language-based content detection model.

[0043] The content detection model 120 can be selected from a model library 150, which can include a target content detection model 122-A trained to detect a particular target content of interest, an open-vocabulary content detection model 122-B trained to recognize unfamiliar content, and / or other suitable models. In some embodiments, the model selection stage 320 can compare the list of keywords to a list of reference keywords associated with individual target content detection models 122-A and compute similarity scores (e.g., cosine similarity values) between the keywords and the reference keywords. If at least some of the similarity scores are above a certain empirically set threshold, the model selection stage 320 can select the corresponding target content detection model 122-A. If the similarity scores are below the threshold, the model selection stage 320 can select one of the available open-vocabulary content detection models 122-B.

[0044] The selected content detection model 120 can process the representation of the prompt 302 and the media item 304 and output an appropriate characterization of the media item 304. The characterization can include any relevant information contained in the media item 304 about the entities identified in the prompt 302, e.g., by the keywords of the prompt 302. In one example embodiment, e.g., when the media item 304 includes an image or a video, the characterization of the media item 304 can include a bounding box, an object type / category, and / or other identifying information of the object in the media item 304. In another non-limiting example, the output of the content detection model 120 can include a segmentation map for the media item 304, e.g., a classification of individual pixels (or groups of pixels) into two or more categories, e.g., “target object pixels,” “background pixels,” and so on.

[0045] Figure 4 An example architecture of an open-vocabulary content detection model 400 that can be used for automated content recognition through custom code generation is shown, in accordance with at least one embodiment. In some embodiments, the open-vocabulary content detection model 400 can beFigure 1 One of the open vocabulary content detection models 122-B in FIG. 4, and can be deployed as one of the content detection models 120 on the MCR server 110. The open vocabulary content detection model 400 can include a language understanding portion (e.g., a text trunk 410 that processes text input 402, such as the prompt 302 of FIG. 3, reference) and a media processing portion (e.g., a media trunk 420 that processes media input 404, such as the media item 304). In one example, the media trunk 420 can be trained to recognize visual patterns in images of various objects, and the text trunk 410 can be trained to recognize contextual and semantic links between individual units (e.g., words, phrases, etc.) in text. The text trunk 410 and / or the media trunk 420 can include one or more self-attention blocks to identify associations between different units of the individual inputs. The output of the processing by the text trunk 410 and the media trunk 420 can be processed by a multi-modal transformer that uses one or more cross-attention blocks (but can also contain any number of self-attention blocks) to identify associations between units of the text input 402 and content units of the media input 404. The intermediate output of the multi-modal transformer 430 can be processed by a suitable classifier, e.g., a media decoder 440 that generates a content classification 450 including any suitable characterization of the media input 404, such as pixel-level classifications, object-level detections (e.g., bounding boxes, convex hulls, etc.) and classifications, audio feature detections, detections of features appearing in sensor data, etc. The content classification 450 can be used for prompt augmentation 220 and code generation, e.g., as disclosed in more detail below. Figure 3

[0046] Referring again to FIG. 4, Figure 3 In one non-limiting example embodiment, the characterization of the media item 304 obtained by the content detection model can have the following illustrative format:

[0047] “filename”: “basketball court.jpg”

[0048] “height”: 693

[0049] “width”: 1024

[0050] “title”: “Basketball player on court”

[0051] “detections”:

[0052] “instances”:

[0053] “id”: 1

[0054] “bbox”: [262, 210, 323, 338],

[0055] ​"confidence": 0.93,

[0056] "class_name": "basketball_player"

[0057] "id": 2

[0058] "bbox": [354, 330, 426, 514],

[0059] "confidence": 0.91,

[0060] "class_name": "basketball_player"

[0061] "id": 3

[0062] "bbox": [126, 202, 276, 282],

[0063] "confidence": 0.94,

[0064] "class_name": "basketball_player"

[0065] specifies the filename of the media item 304 ("basketball court.jpg"), the dimensions of the media item (693 x 1024 pixels), a list of keywords / captions that have been performed content detection ("basketball players on the court"), and various detection results identified with bounding boxes ("bbox") of detected objects, classes of detected objects ("basketball_player"), and confidence of detection.

[0066] Figures 5A-5B Content detection that can be used with custom computational code generation, in accordance with at least one embodiment, is schematically illustrated. Figure 5A An image 500 of a shipping warehouse including a robot 502 and a plurality of packages 504 is depicted. The image 500 can be used as a media item 304, in conjunction with a suitable prompt 302, such as "How many packages is the robot transporting?" Figure 5B An image 510 of a shipping warehouse annotated with detection results performed by a content detection model is depicted. As shown, the annotations (shown as bounding boxes) include a robot detection 512 and package detections 514.

[0067] Referring again to Figure 3The representation of the content of the media item 304 obtained by the one or more content detection models 120 can be used to prompt augmentations 220. The prompt augmentations 220 can be used to augment the prompt 302 using the media item representation. The augmented prompt 322 can also contain instructions that require the coding LM 330 to write code that, given the information contained in the representation of the media item 304 (generated by the content detection models 120), performs the computational task described in the prompt 302 related to the media item 304. In one non-limiting example, the augmented prompt 322 can be: “Write Python code to identify how many basketball players are on the court according to the provided representation of the image.” The augmented prompt 322 can also include an explanation of the representation format, e.g., a description of the individual fields in the representation (“bounding box,” “object,” “confidence,” etc.). In some embodiments, the explanation can be provided in natural language form.

[0068] The coding LM 330 can process the augmented prompt 322 and generate code 332 (e.g., source code or object code) that is capable of extracting the information contained in the representation of the media item 304 in response to the prompt 302. The code 332 produced by the LM 330 can include a list of commands (operators) that instruct a processing device how to (i) extract and analyze the data contained in the representation of the media item 304 and (ii) generate a response to the prompt 302.

[0069] The code 332 written (generated) by the coding LM 330 can be processed by the code execution module 230. The code execution 230 can include a compiler (if needed) to convert source code to machine code and processing logic to execute the machine code. The output of the code execution 230 can be a response 340, e.g., a natural language phrase that can be understood by a human user (e.g., the creator of the prompt 302). In some embodiments, the response 340 can be pre-formatted by the coding LM 330 with fillable blanks, e.g., “[…] basketball players on the court,” where the blanks are filled with numerical or other suitable values when the code 332 is executed.

[0070] In certain embodiments, the code 332 generated by the coding LM 330 can not always be able to successfully compile and / or execute during code execution 230. In the event of an error, the code 332 can be returned to the coding LM 330 along with a log of the compilation and / or execution error, and the coding LM 330 can debug the code 332. Such debugging can be performed iteratively using multiple attempts. In certain cases, the coding LM 330 can write new code 332 in a different language. For example, if C++ code still does not run after a set number of iterations, the coding LM 330 can be requested (e.g., by the hint augmentation 220 generating a new augmented hint 322) to write Python code or JavaScript code or code in some other language. In certain embodiments, after a predetermined maximum number of attempts still does not generate viable code (e.g., code that can be successfully compiled and / or executed), the coding LM 330 can output a final error, and a human developer can manually correct and / or debug the code 332. The corrected code 332 can then be used for the teachable on-the-fly training of the LM 330, so that the coding LM 330 does not make the same mistake in the future.

[0071] Figure 6 is a flowchart of an example method 600 of automated content recognition using a language model to generate custom computing code, in accordance with at least one embodiment. In at least one embodiment, method 600 can be performed using a processing unit of a computing device 200 in Figure 2 may be (or include) devices associated with MCR server 110, LM service 130, user device 102, and / or other devices. In at least one embodiment, a processing unit performing method 600 can execute instructions stored on a non-transitory computer-readable storage medium. In at least one embodiment, method 600 can be performed using multiple processing threads (e.g., CPU threads and / or GPU threads), where individual threads perform one or more separate functions, routines, subroutines, or operations of the method. In at least one embodiment, processing threads implementing method 600 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, processing threads implementing method 600 can be executed asynchronously relative to one another. Various operations of method 600 can be performed in a different order than shown in Figure 6 . In at least one embodiment, one or more operations shown in Figure 6 are not always performed.

[0072] At block 610, the method 600 can include obtaining a first prompt to obtain a responsive description of a media item for the first prompt. In some embodiments, the media item can be or include an image item, a video item, an audio item, a sensor data item, etc. For example, a user can input any suitable query (e.g., the prompt 302 in Figure 3 FIG. 3) about an image, a video, a dataset (e.g., the media item 304), etc. In some embodiments, the first prompt can be or include a natural language prompt. In some embodiments, the first prompt can be a textual representation of audio data or visual data from a user, or can be retrieved from memory.

[0073] At block 620, the method 600 can include processing the media item and a representation of the first prompt using at least one content detection model to obtain a representation of the media item. In some embodiments, the representation of the first prompt can include one or more keywords associated with the first prompt. In some embodiments, the at least one content detection model can be selected from a plurality of trained models using the representation of the first prompt. In some embodiments, the at least one content detection model can include an object detection model trained to detect one or more objects associated with the representation of the first prompt. In some embodiments, the at least one content detection model can include an open-vocabulary model (e.g., the open-vocabulary content detection model 400 in Figure 4 FIG. 4) that includes a computer vision portion (e.g., the media backbone 420) to process at least the computer vision portion of the media item, a language understanding portion (e.g., the text backbone 410) to process at least the language understanding portion of the media item, and a classifier portion (e.g., the multi-modal transformer 430 and / or the media decoder 440) to process outputs of the computer vision portion and the language understanding portion to obtain the representation of the media item. In some embodiments, the representation of the media item can include one or more bounding boxes (e.g., as shown in Figure 5B FIG. 5) for respective one or more objects in the media item. In some embodiments, the representations of the media item from all content detection models can be aggregated (combined) before downstream use.

[0074] At block 630, the method 600 can continue to generate a second prompt (e.g., the augmented prompt 322 in Figure 3 FIG. 3) that includes instructions to a language model (LM) using the first prompt and the representation of the media item. In some embodiments, the instructions to the LM can include a natural language explanation of a format of the representation of the media item. The LM can be trained to generate computational code to perform a computational task in response to natural language instructions that include a description of the computational task.

[0075] At block 640, the method 600 can continue with the LM processing the second prompt to generate a computational code (e.g., code 332) associated with the representation of the media item. At block 650, the method 600 can include causing the computational code to be executed to generate a responsive description (e.g., response 340) of the media item.

[0076] The systems and methods described herein can be used for a variety of purposes, such as, but not limited to, for performing one or more operations related to machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twin, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation and / or digital twin, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0077] The disclosed embodiments can be included in a variety of different systems, such as an automotive system (e.g., an in-vehicle infotainment system for an autonomous or semi-autonomous machine), a system implemented using robotics, an aviation system, a medical system, a boating system, a smart area monitoring system, a system for performing deep learning operations, a system for performing simulation operations, a system for performing digital twin operations, a system for performing medical operations, a system for performing factory operations, a system for performing analytics operations, a system implemented using edge devices, a system containing one or more virtual machines (VMs), a system for performing synthetic data generation operations, a system implemented at least partially in a data center, a system for performing conversational AI operations, a system for performing light transport simulation, a system for performing collaborative content creation of 3D assets, a system for performing generative AI operations, a system implemented at least partially using cloud computing resources, and / or other types of systems.

[0078] Inference and training logic

[0079] Figure 7A Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. In one implementation, the inference and / or training logic 715 are used in a system to perform one or more operations described herein.

[0080] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters of neurons or layers of a neural network configured in aspects of one or more embodiments that are trained and / or used for inferencing. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs) or simply circuits) of a processor. In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which the code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0081] In at least one embodiment, any portion of code and / or data storage 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 701 can be cache memory, dynamic random addressable memory (“DRAM”), static random addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or code and / or data storage 701 is internal or external to a processor, e.g., or comprised of DRAM, SRAM, flash or some other storage type, can depend on available storage space on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.

[0082] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 705 to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 705 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 705 to store graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).

[0083] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 705 is internal or external to a processor, e.g., whether it is made up of DRAM, SRAM, Flash memory, or some other storage type, depends on whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data being used in inferencing and / or training of a neural network, or some combination of these factors.

[0084] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be a combined storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0085] In at least one embodiment, inference and / or training logic 715 can include, without limitation, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. Results of such operations can result in activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 720, which are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations stored in activation storage 720 are generated by ALUs 710 in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALUs 710, where weight values stored in code and / or data storage 705 and / or data storage 701 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 705 or code and / or data storage 701 or another on-chip or off-chip storage.

[0086] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 710, while in another embodiment, one or more ALUs 710 may be located outside the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may share a processor or other hardware logic device or circuitry, while in another embodiment, they may be located in different processors or other hardware logic devices or circuitry, or in some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation storage 720 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0087] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or some other memory type.

[0088] In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7AThe illustrated inference and / or training logic 715 can be used in combination with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (e.g., field programmable gate arrays (“FPGAs”)).

[0089] Figure 7B Inference and / or training logic 715 are illustrated as a portion of the system 700 in the example of FIG. 7. In other examples, the inference and / or training logic 715 can be used in combination with other systems or devices, including but not limited to other systems or devices that include one or more other processors 710. Figure 7B The inference and / or training logic 715 illustrated in FIG. 7 can be used in combination with application-specific integrated circuit (ASIC) hardware, such as Google’s Tensor Processing Unit (TPU), Graphcore’s AI processing units from Graphcore TM Inference Processing Units (IPUs) from Intel Corp, or “Lake Crest”) processors. In at least one embodiment, the inference and / or training logic 715 illustrated in FIG. 7 can be used in combination with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate arrays (FPGAs)). Figure 7B Inference and / or training logic 715 are illustrated as a portion of the system 700 in the example of FIG. 7. In other examples, the inference and / or training logic 715 can be used in combination with other systems or devices, including but not limited to other systems or devices that include one or more other processors 710. Figure 7B In at least one embodiment, each of the code and / or data storage 701 and the code and / or data storage 705 are respectively associated with a dedicated computing resource (e.g., computing hardware 702 and computing hardware 706). In at least one embodiment, each of the computing hardware 702 and the computing hardware 706 includes one or more ALUs that only perform mathematical functions (e.g., linear algebraic functions) on information stored in the code and / or data storage 701 and the code and / or data storage 705, respectively, and the results of the functions performed are stored in the activation storage 720.

[0090] In at least one embodiment, each of code and / or data stores 701 and 705 and corresponding compute hardware 702 and 706, respectively, correspond to different layers of a neural network, such that activations resulting from one storage / compute pair 701 / 702 of code and / or data store 701 and compute hardware 702 are provided as input to next storage / compute pair 705 / 706 of code and / or data store 705 and compute hardware 706 in order to reflect a conceptual organization of a neural network. In at least one embodiment, each storage / compute pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in inference and / or training logic 715 after or in parallel with storage / compute pairs 701 / 702 and 705 / 706.

[0091] Neural network training and deployment

[0092] Figure 8 Training and deployment of a deep neural network is shown, in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training dataset 802. In at least one embodiment, training framework 804 is a PyTorch framework, while in other embodiments, training framework 804 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, training framework 804 trains untrained neural network 806 and enables it to be trained using processing resources described herein to generate a trained neural network 808. In at least one embodiment, weights can be chosen randomly or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.

[0093] In at least one embodiment, an untrained neural network 806 is trained using supervised learning, where a training dataset 802 includes inputs paired with desired outputs for inputs, or where a training dataset 802 includes inputs with known outputs and the output of the neural network 806 is manually graded. In at least one embodiment, an untrained neural network 806 is trained in a supervised manner and processes an input from a training dataset 802 and compares a resulting output to a set of expected or desired outputs. In at least one embodiment, an error is then propagated back through the untrained neural network 806. In at least one embodiment, a training framework 804 adjusts weights that control the untrained neural network 806. In at least one embodiment, a training framework 804 includes tools to monitor how well an untrained neural network 806 is converging towards a model (e.g., a trained neural network 808) that is suitable for generating correct answers (e.g., results 814) based on input data (e.g., new datasets 812). In at least one embodiment, a training framework 804 trains an untrained neural network 806 repeatedly while adjusting weights to improve outputs of the untrained neural network 806 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, a training framework 804 trains an untrained neural network 806 until the untrained neural network 806 reaches a desired accuracy. In at least one embodiment, a trained neural network 808 can then be deployed to implement any number of machine learning operations.

[0094] In at least one embodiment, an untrained neural network 806 is trained using unsupervised learning, where an untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 802 will include input data without any associated output data or “ground truth” data. In at least one embodiment, an untrained neural network 806 can learn groupings within a training dataset 802 and can determine how individual inputs relate to the untrained dataset 802. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 808 that is capable of performing operations useful for reducing a dimensionality of new datasets 812. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in new datasets 812 that deviate from a normal pattern of new datasets 812.

[0095] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mix of labeled and unlabeled data is included in training dataset 802. In at least one embodiment, training framework 804 can be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning enables trained neural network 808 to adapt to new dataset 812 without forgetting knowledge that was imprinted into trained neural network 808 during initial training.

[0096] Referring to Figure 9 , Figure 9 is an example dataflow graph for a process 900 for generating and deploying processing and inference pipelines in accordance with at least one embodiment. In at least one embodiment, process 900 can be deployed for performing game name identification analysis and inference on user feedback data at one or more facilities 902, such as data centers.

[0097] In at least one embodiment, process 900 can be executed within a training system 904 and / or a deployment system 906. In at least one embodiment, training system 904 can be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for deployment system 906. In at least one embodiment, deployment system 906 can be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements of facilities 902. In at least one embodiment, deployment system 906 can provide a pipeline platform for selecting, customizing, and implementing virtual instruments for use with computing devices at facilities 902. In at least one embodiment, virtual instruments can include software-defined applications for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in a pipeline can use or call services (e.g., inference, visualization, computation, AI, etc.) of deployment system 906 during application execution.

[0098] In at least one embodiment, some applications used in advanced processing and inference pipelines can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models can be trained using feedback data 908 (e.g., imaging data) stored at facilities 902 or feedback data 908 from another facility or facilities, or a combination thereof, at facilities 902. In at least one embodiment, training system 904 can be used to provide applications, services, and / or other resources to generate working, deployable machine learning models for deployment system 906.

[0099] In at least one embodiment, the model registry 924 may be supported by an object storage system that supports version control and object metadata. In at least one embodiment, it may be available from within a cloud platform via, for example, cloud storage (e.g., Figure 10 The system uses a Cloud 1026-compatible Application Programming Interface (API) to access object storage. In at least one embodiment, machine learning models within the model registry 924 can be uploaded, listed, modified, or deleted by the developer or partner of the system interacting with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate models with applications, enabling the models to be executed as part of the containerized instantiation of the application.

[0100] In at least one embodiment, training pipeline 1004 ( Figure 10 The scenario may include where facility 902 is training its own machine learning model or has an existing machine learning model that needs optimization or updating. In at least one embodiment, feedback data 908 may be received from various channels, such as forums, web forms, etc. In at least one embodiment, once feedback data 908 is received, AI-assisted annotation 910 may be used to help generate annotations corresponding to the feedback data 908 for use as ground truth data for the machine learning model. In at least one embodiment, AI-assisted annotation 910 may include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations corresponding to certain types of feedback data 908 (e.g., from certain devices) and / or certain types of anomalies in the feedback data 908. In at least one embodiment, AI-assisted annotation 910 may then be used directly or may be adjusted or fine-tuned using annotation tools to generate ground truth data. In at least one embodiment, in some examples, labeled data 912 may be used as ground truth data for training the machine learning model. In at least one embodiment, AI-assisted annotation 910, labeled data 912, or a combination thereof may be used to train the machine learning model (e.g., via...). Figures 9-10 The model is trained using ground-based real-world data (914). In at least one embodiment, the trained machine learning model may be referred to as output model 916 and may be used by deployment system 906, as described herein.

[0101] In at least one embodiment, training pipeline 1004 ( Figure 10) can include situations in which the facility 902 needs a machine learning model for performing one or more processing tasks for one or more applications in the deployment system 906, but the facility 902 can not currently have such a machine learning model (or can not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from the model registry 924. In at least one embodiment, the model registry 924 can include machine learning models that are trained to perform a variety of different inferencing tasks on imaging data. In at least one embodiment, the machine learning models in the model registry 924 can have been trained on imaging data from different facilities (e.g., facilities located remotely from the facility 902). In at least one embodiment, the machine learning models can have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a particular location (which can be in the form of feedback data 908), the training can occur at that location, or at least in a manner that protects the confidentiality of the imaging data or limits the transfer of the imaging data from outside the facility (e.g., in compliance with HIPAA regulations, privacy regulations, etc.). In at least one embodiment, once a model is trained, or partially trained, at a location, the machine learning model can be added to the model registry 924. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in the model registry 924. In at least one embodiment, a machine learning model can then be selected from the model registry 924 (and referred to as an output model 916), and can be used in the deployment system 906 to perform one or more processing tasks for one or more applications of the deployment system.

[0102] In at least one embodiment, the training pipeline 1004( Figure 10) can be used in a scenario that includes a facility 902 that needs a machine learning model for performing one or more processing tasks for deploying one or more applications in a deployment system 906, but the facility 902 can not currently have such a machine learning model (or can not have an optimized, efficient, or effective model for that purpose). In at least one embodiment, due to population differences, genetic variations, robustness of training data used to train the machine learning model, diversity of training data anomalies, and / or other issues with the training data, the machine learning model selected from the model registry 924 can not be fine-tuned or optimized for feedback data 908 generated at the facility 902. In at least one embodiment, AI assisted annotation 910 can be used to help generate annotations corresponding to the feedback data 908 for use as ground truth data to retrain or update the machine learning model. In at least one embodiment, labeled data 912 can be used as ground truth data to train the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 914. In at least one embodiment, the model training 914 (e.g., AI assisted annotation 910, labeled data 912, or a combination thereof) can be used as ground truth data to retrain or update the machine learning model.

[0103] In at least one embodiment, the deployment system 906 can include software 918, services 920, hardware 922, and / or other components, features, and functionality. In at least one embodiment, the deployment system 906 can include a software “stack” such that the software 918 can be built on top of the services 920 and can use the services 920 to perform some or all processing tasks, and the services 920 and software 918 can be built on top of the hardware 922 and use the hardware 922 to perform processing, storage, and / or other computing tasks for the deployment system 906.

[0104] In at least one embodiment, software 918 can include any number of different containers, where each container can execute an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks in a high-level processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, there can be any number of containers for each type of computing device that can perform data processing tasks on feedback data 908 (or other data types, such as data types described herein). In at least one embodiment, in addition to containers that receive and configure imaging data for use by each container and / or by facility 902 after processing through a pipeline, a high-level processing and inference pipeline can be defined based on selection of different containers that are desired or required for processing feedback data 908 (e.g., to convert output back to a usable data type for storage and display at facility 902). In at least one embodiment, a combination of containers within software 918 (e.g., that make up a pipeline) can be referred to as a virtual instrument (as described in greater detail herein), and a virtual instrument can utilize services 920 and hardware 922 to perform some or all processing tasks of applications instantiated in containers.

[0105] In at least one embodiment, data can be pre-processed as part of a data processing pipeline to prepare data for processing by one or more applications. In at least one embodiment, post-processing can be performed on output of one or more inference tasks or other processing tasks of a pipeline to prepare output data for a next application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks can be performed by one or more machine learning models, such as trained or deployed neural networks, which can include output models 916 of training system 904.

[0106] In at least one embodiment, tasks of a data processing pipeline can be encapsulated in one or more containers, each of which represents a discrete, fully-functional instantiation of an application and virtualized computing environment that can reference a machine learning model. In at least one embodiment, containers or applications can be published into a private (e.g., limited access) area of a container registry (described in greater detail herein), and trained or deployed models can be stored in model registry 924 and associated with one or more applications. In at least one embodiment, images of applications (e.g., container images) can be available in a container registry, and once a user selects an image from a container registry for deployment in a pipeline, the image can be used to generate a container for instantiation of an application for use by a user’s system.

[0107] In at least one embodiment, the developer can develop, publish, and store an application (e.g., as a container) for performing processing and / or inference on the provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publication, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, on data from a first facility), the SDK serving as a system (e.g.,...). Figure 10 Architecture 1000 may support at least some services 920. In at least one embodiment, once verified by architecture 1000 (e.g., for accuracy, etc.), the application becomes available in the container registry for users (e.g., hospitals, clinics, laboratories, healthcare providers, etc.) to select and / or implement one or more processing tasks on data at the user's facility (e.g., a second facility).

[0108] In at least one embodiment, the developer can then share the application or container over a network for the system (e.g., Figure 10 The architecture 900 allows for user access and use. In at least one embodiment, a completed and validated application or container may be stored in a container registry, and associated machine learning models may be stored in a model registry 924. In at least one embodiment, a requesting entity (which provides an inference or image processing request) may browse the container registry and / or model registry 924 for applications, containers, datasets, machine learning models, etc., select the desired combination of elements to include in the data processing pipeline, and submit a processing request. In at least one embodiment, the request may include input data necessary to execute the request, and / or may include a selection of the application and / or machine learning model to be executed when the request is processed. In at least one embodiment, the request may then be passed to one or more components of deployment system 906 (e.g., the cloud) to perform processing in the data processing pipeline. In at least one embodiment, the processing performed by deployment system 906 may include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 924. In at least one embodiment, once results are generated through the pipeline, the results may be returned to the user for reference (e.g., for viewing in a suite of viewing applications executed on a local machine, local workstation, or terminal).

[0109] In at least one embodiment, to help process or execute applications or containers in a pipeline, services 920 can be utilized. In at least one embodiment, services 920 can include compute services, collaborative content creation services, simulation services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, services 920 can provide functionality that is common to one or more applications in software 918, and thus can abstract functionality that can be called or utilized by applications. In at least one embodiment, functionality provided by services 920 can run dynamically and more efficiently, while also scaling well by allowing applications to process data in parallel (e.g., using parallel computing platform 1030 in FIG. 10B). Figure 10 In at least one embodiment, rather than requiring each application that requires the same functionality provided by a shared service 920 to have a respective instance of service 920, services 920 can be shared among and between various applications. In at least one embodiment, as a non-limiting example, services can include inference servers or engines that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included, which can provide machine learning model training and / or retraining capabilities.

[0110] In at least one embodiment, where services 920 include AI services (e.g., inference services), as part of application execution, one or more machine learning models associated with an application for anomaly detection (e.g., tumors, growth anomalies, scarring, etc.) can be executed by calling (e.g., as an API call) an inference service (e.g., inference server) to execute one or more machine learning models or processing thereof. In at least one embodiment, where another application includes one or more machine learning models for segmentation tasks, the application can call an inference service to execute the machine learning models for performing one or more processing operations associated with segmentation tasks. In at least one embodiment, software 918 implementing an advanced processing and inference pipeline can be pipelined, as each application can call the same inference service to perform one or more inference tasks.

[0111] In at least one embodiment, hardware 922 can include GPUs, CPUs, graphics cards, AI / deep learning systems (e.g., AI supercomputers such as NVIDIA’s DGX TM(Supercomputer system), cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 922 may be used to provide efficient, specially built support for software 918 and services 920 in deployment system 906. In at least one embodiment, GPU processing may be used to perform local processing (e.g., at facility 902) within the AI / deep learning system, in the cloud system, and / or other processing components of deployment system 906 to improve the efficiency, accuracy, and performance of game name recognition.

[0112] In at least one embodiment, as a non-limiting example, regarding deep learning, machine learning and / or high-performance computing, simulation and visual computing, software 918 and / or service 920 may be optimized for GPU processing. In at least one embodiment, at least some of the computing environment in which the deployment system 906 and / or training system 904 are located may have GPU-optimized software (e.g., NVIDIA DGX). TM The system's hardware and software combination is executed in a data center or one or more supercomputers or high-performance computing systems. In at least one embodiment, as described herein, hardware 922 may include any number of GPUs that can be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU-optimized execution for deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, AI / deep learning supercomputers and / or GPU-optimized software (e.g., such as NVIDIA's DGX) may be used. TM The system provides a hardware abstraction and scaling platform to execute cloud platforms (e.g., NVIDIA's NGC). TM In at least one embodiment, the cloud platform can integrate application container cluster systems or coordination systems (e.g., KUBERNETES) across multiple GPUs to achieve seamless scaling and load balancing.

[0113] Figure 10 This is a system diagram of an example architecture 1000 for generating and deploying a deployment pipeline according to at least one embodiment. In at least one embodiment, architecture 1000 can be used to implement Figure 9 The process 900 and / or other processes include high-level processing and inference pipelines. In at least one embodiment, architecture 1000 may include a training system 904 and a deployment system 906. In at least one embodiment, the training system 904 and the deployment system 906 may be implemented using software 918, services 920, and / or hardware 922, as described herein.

[0114] In at least one embodiment, architecture 1000 (e.g., training system 904 and / or deployment system 906) can be implemented in a cloud computing environment (e.g., using cloud 1026). In at least one embodiment, architecture 1000 can be implemented locally (with respect to a facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1026 can be limited to authorized users by instituting security measures or protocols. In at least one embodiment, security protocols can include network tokens that can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, APIs of a virtual instrument (described herein) or other instances of architecture 1000 can be limited to a set of Internet Service Providers (ISPs) that have been vetted or authorized for interaction.

[0115] In at least one embodiment, various components of architecture 1000 can communicate with and within each other using any of a plurality of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between facilities and components of architecture 1000 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be communicated through one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (such as Ethernet), etc.

[0116] In at least one embodiment, similar to training pipeline 1004 described herein with respect to Figure 9 In at least one embodiment, training system 904 can execute training pipeline 1004. In at least one embodiment, where deployment system 906 is to use one or more machine learning models in deployment pipeline 1010, training pipeline 1004 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1006 (e.g., without retraining or updating). In at least one embodiment, as a result of training pipeline 1004, an output model 916 can be generated. In at least one embodiment, training pipeline 1004 can include any number of processing steps, AI-assisted annotation 910, labeling or annotation of feedback data 908 for use in generating labeled data 912, selection of a model from a model registry, model training 914, training, retraining, or updating of a model, and / or other processing steps. In at least one embodiment, different training pipelines 1004 can be used for different machine learning models used by deployment system 906. In at least one embodiment, training pipeline 1004 of a first example described with respect to Figure 9 In at least one embodiment, training pipeline 1004 of a first example described with respect to Figure 9The training pipeline 1004 of the second example described can be used for a second machine learning model, similar to the description of the first example Figure 9 The training pipeline 1004 of the third example described can be used for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 904 can be used according to the requirements of each respective machine learning model. In at least one embodiment, one or more machine learning models can already be trained and ready for deployment, so the training system 904 can not perform any processing on the machine learning model and the machine learning model can be implemented by the deployment system 906.

[0117] In at least one embodiment, the one or more output models 916 and / or pre-trained models 1006 can include any type of machine learning model, according to embodiment. In at least one embodiment and without limitation thereto, machine learning models used by the architecture 1000 can include using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machine, etc.), and / or other types of machine learning models.

[0118] In at least one embodiment, training pipeline 1004 can include AI assisted annotation. In at least one embodiment, labeled data 912 (e.g., traditional annotation) can be generated by any number of techniques. In at least one embodiment, labels or other annotations can be generated in a drawing program (e.g., annotation program), a computer aided design (CAD) program, a labeling program, another type of application suitable for generating ground truth annotations or labels, and / or can be hand drawn, in at least one embodiment, ground truth data can be synthetically generated (e.g., generated from computer models or renderings), realistically generated (e.g., designed and generated from real world data), automatically generated by a machine (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a labeler or annotation specialist defining locations of labels), and / or combinations thereof. In at least one embodiment, for each instance of feedback data 908 (or other data type used by a machine learning model), there can be corresponding ground truth data generated by training system 904. In at least one embodiment, AI assisted annotation can be performed as part of deployment pipeline 1010; in addition to or instead of AI assisted annotation included in training pipeline 1004. In at least one embodiment, architecture 1000 can include a multi-tiered platform that can include a software tier (e.g., software 918) of diagnostic applications (or other application types) that can perform one or more medical imaging and diagnostic functions.

[0119] In at least one embodiment, software tier can be implemented as a secure, encrypted, and / or authenticated API through which applications or containers can be invoked (e.g., called) from an external environment (e.g., facility 902). In at least one embodiment, applications can then call or execute one or more services 920 to perform computing, AI, or visualization tasks associated with respective applications, and software 918 and / or services 920 can leverage hardware 922 to perform processing tasks in an efficient and effective manner.

[0120] In at least one embodiment, deployment system 906 can execute deployment pipeline 1010. In at least one embodiment, deployment pipeline 1010 can include any number of applications that can be sequential, non-sequential, or otherwise applied to feedback data (and / or other data types) - including AI assisted annotation, as described above. In at least one embodiment, a deployment pipeline 1010 for an individual device can be referred to as a virtual instrument for a device, as described herein. In at least one embodiment, for a single device, there can be more than one deployment pipeline 1010, depending on information desired from data generated by a device.

[0121] In at least one embodiment, applications available for deployment pipeline 1010 can include any applications that can be used to perform processing tasks on feedback data or other data from devices. In at least one embodiment, because various applications can share common image operations, in some embodiments, a data augmentation library (e.g., as one of services 920) can be used to accelerate these operations. In at least one embodiment, to avoid bottlenecks of traditional processing methods that rely on CPU processing, parallel computing platform 1030 can be used for GPU acceleration of these processing tasks.

[0122] In at least one embodiment, deployment system 906 can include a user interface 1014 (e.g., graphical user interface, web interface, etc.) that can be used to select applications to include in deployment pipeline 1010, arrange applications, modify or change applications or parameters or constructs thereof, use and interact with deployment pipeline 1010 during setup and / or deployment, and / or otherwise interact with deployment system 906. In at least one embodiment, although not shown with respect to training system 904, UI 1014 (or a different user interface) can be used to select models for use in deployment system 906, to select models for training or retraining in training system 904, and / or to otherwise interact with training system 904. In at least one embodiment, training system 904 and deployment system 906 can include DICOM adapters 1002A and 1002B.

[0123] In at least one embodiment, in addition to application coordination system 1028, pipeline manager 1012 can be used to manage interactions between applications or containers of deployment pipeline 1010 and services 920 and / or hardware 922. In at least one embodiment, pipeline manager 1012 can be configured to facilitate interactions from application to application, from application to service 920, and / or from application or service to hardware 922. In at least one embodiment, although shown as included in software 918, this is not intended to be limiting, and in some examples, pipeline manager 1012 can be included in services 920. In at least one embodiment, application coordination system 1028 (e.g., Kubernetes, DOCKER, etc.) can include a container coordination system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from deployment pipeline 1010 (e.g., reconstruction applications, segmentation applications, etc.) with individual containers, each application can execute in a self-contained environment (e.g., at kernel level) to improve speed and efficiency.

[0124] In at least one embodiment, each application and / or container (or image thereof) can be separately developed, modified, and deployed (e.g., a first user or developer can develop, modify, and deploy a first application, a second user or developer can develop, modify, and deploy a second application separate from first user or developer), which can allow for focus and attention to be directed to tasks of a single application and / or container without being impeded by tasks of other applications or containers. In at least one embodiment, pipeline manager 1012 and application coordination system 1028 can facilitate communication and cooperation between different containers or applications. In at least one embodiment, as long as intended inputs and / or outputs of each container or application are known to system (e.g., based on construction of application or container), application coordination system 1028 and / or pipeline manager 1012 can facilitate communication between and among each application or container and sharing of resources. In at least one embodiment, as one or more applications or containers in deployment pipeline 1010 can share same services and resources, application coordination system 1028 can coordinate, load balance, and determine sharing of services or resources between and among various applications or containers. In at least one embodiment, a scheduler can be used to track resource needs of applications or containers, current or planned use of these resources, and resource availability. Accordingly, in at least one embodiment, a scheduler can allocate resources to different applications and among and between applications, taking into account needs and availability of system. In some examples, a scheduler (and / or other components of application coordination system 1028) can determine resource availability and distribution based on constraints imposed on system (e.g., user constraints), such as quality of service (QoS), urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0125] In at least one embodiment, services 920 utilized by and shared by applications or containers in deployment system 906 can include compute services 1016, collaboration content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, and / or other service types. In at least one embodiment, an application can invoke (e.g., execute) one or more services 920 to perform processing operations for the application. In at least one embodiment, an application can utilize compute services 1016 to perform supercomputing or other high performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1016 can be utilized to perform parallel processing (e.g., using parallel computing platform 1030) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1030 (e.g., NVIDIA’s DGX A100) can be used to perform parallel processing of data by one or more applications and / or one or more tasks of a single application. General-purpose computing on GPUs (GPGPU) (e.g., GPU 1022) can be implemented. In at least one embodiment, software layers of parallel computing platform 1030 can provide access to virtual instruction sets and parallel computing elements of a GPU to execute a compute kernel. In at least one embodiment, parallel computing platform 1030 can include memory, and in some embodiments, can share memory between and among multiple containers, and / or between and among different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls can be generated for multiple containers and / or multiple processes within a container to use the same data from a shared memory segment of parallel computing platform 1030 (e.g., where multiple different stages of an application or applications are processing the same information). In at least one embodiment, rather than copying data and moving data to different locations in memory (e.g., read / write operations), the same data in the same location in memory can be used for any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, as data is used to generate new data as a result of processing, this information of the new location of the data can be stored and shared between various applications. In at least one embodiment, location of data, as well as location of updated or modified data, can be part of a definition of how to understand a payload in a container.

[0126] In at least one embodiment, AI services 1018 can be utilized to perform inferencing services for executing machine learning models associated with applications (e.g., task is to perform one or more processing tasks for an application). In at least one embodiment, AI services 1018 can utilize AI system 1024 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inferencing tasks. In at least one embodiment, an application of deployment pipeline 1010 can use one or more output models 916 from training system 904 and / or other models of an application to perform inferencing on imaging data (e.g., DICOM data, RIS data, CIS data, REST-compliant data, RPC data, raw data, etc.). In at least one embodiment, two or more examples of inferencing can be available using application coordination system 1028 (e.g., a scheduler). In at least one embodiment, a first category can include a high priority / low latency path, which can implement a higher service level agreement, such as for performing inferencing on urgent requests in emergency situations, or for radiologists during a diagnosis process. In at least one embodiment, a second category can include a standard priority path, which can be used for requests that can not be urgent or can have analysis performed at a later time. In at least one embodiment, application coordination system 1028 can allocate resources (e.g., services 920 and / or hardware 922) based on a priority path for different inferencing tasks of AI services 1018.

[0127] In at least one embodiment, shared storage can be installed to AI services 1018 in architecture 1000. In at least one embodiment, shared storage can operate as a cache (or other storage device type) and can be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of deployment system 906 can receive the request and can select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request can be input into a database, a machine learning model can be located from model registry 924 if not already in cache, a validation step can ensure that appropriate machine learning model is loaded into cache (e.g., shared storage), and / or a copy of the model can be saved to cache. In at least one embodiment, if an application is not already running or there are not enough instances of an application, a scheduler (e.g., of pipeline manager 1012) can be used to start the application referenced in the request. In at least one embodiment, if an inference server is not already started to execute the model, an inference server can be started. In at least one embodiment, each model can start any number of inference servers. In at least one embodiment, in a pull model of clustering inference servers, a model can be cached whenever load balancing is favorable. In at least one embodiment, inference servers can be statically loaded into respective distributed servers.

[0128] In at least one embodiment, inference can be performed using inference servers running in containers. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally multiple versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, when an inference server is started, a model can be passed to the inference server so that the same container can be used to service different models as long as the inference server is run as a different instance.

[0129] In at least one embodiment, during application execution, an inference request for a given application can be received and a container (e.g., an instance hosting an inference server) can be loaded (if not already loaded) and a launcher can be invoked. In at least one embodiment, pre-processing logic in a container can load, decode, and / or perform any additional pre-processing on incoming data (e.g., using CPU and / or GPU). In at least one embodiment, once data is ready for inference, a container can infer on data as needed. In at least one embodiment, this can include a single inference call on one image (e.g., a hand X-ray), or can require inference on hundreds of images (e.g., a chest CT). In at least one embodiment, an application can summarize results before completion, which can include, without limitation, a single confidence score, a pixel-level segmentation, a voxel-level segmentation, generating a visualization, or generating text to summarize results. In at least one embodiment, different priorities can be assigned for different models or applications. For example, some models can have a real-time (turnaround time less than 1 minute) priority, while other models can have a lower priority (e.g., turnaround time less than 10 minutes). In at least one embodiment, model execution time can be measured from a requesting authority or entity, and can include a cooperative network traversal time as well as an inference service’s execution time.

[0130] In at least one embodiment, transfer of requests between service 920 and inference applications can be hidden behind a software development kit (SDK), and robust transfer can be provided through a queue. In at least one embodiment, requests will be placed in a queue through an API for individual application / tenant ID combinations, and the SDK pulls requests from the queue and provides them to the application. In at least one embodiment, a name of a queue can be provided in an environment from which the SDK pulls requests. In at least one embodiment, asynchronous communication through a queue can be useful because it can allow any instance of an application to pick up work when it is available. In at least one embodiment, results can be transferred back through a queue to ensure no data loss. In at least one embodiment, a queue can also provide an ability to split work, as highest priority work can go into a queue that connects to most instances of an application, while lowest priority work can go into a queue that connects to a single instance that processes tasks in order of receipt. In at least one embodiment, an application can run on GPU-accelerated instances that are spawned in cloud 1026, and an inference service can perform inference on a GPU.

[0131] In at least one embodiment, visualization service 1020 can be utilized to generate visualizations for viewing application and / or deployment pipeline 1010 output. In at least one embodiment, visualization service 1020 can utilize GPU(s) 1022 to generate visualizations. In at least one embodiment, visualization service 1020 can implement rendering effects such as ray tracing or other light transport simulation techniques to generate higher quality visualizations. In at least one embodiment, visualizations can include, without limitation, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, virtualized environments can be used to generate virtual interactive displays or environments (e.g., virtual environments) for interaction by system users (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, visualization service 1020 can include internal visualizers, movie and / or other rendering or image processing capabilities or functionality (e.g., ray tracing, rasterization, internal optics, etc.).

[0132] In at least one embodiment, hardware 922 can include GPU(s) 1022, AI system(s) 1024, cloud 1026, and / or any other hardware used to execute training system 904 and / or deployment system 906. In at least one embodiment, GPU(s) 1022 (e.g., NVIDIA’s and / or QUADRO GPUs) can include any number of GPUs that can be used to perform processing tasks for any features or functionality of computing services 1016, collaboration content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, other services, and / or software 918. For example, for AI services 1018, GPU(s) 1022 can be used to perform pre-processing on imaging data (or other data types used by machine learning models), perform post-processing on outputs of machine learning models, and / or perform inferencing (e.g., to execute machine learning models). In at least one embodiment, cloud 1026, AI system(s) 1024, and / or other components of architecture 1000 can use GPU(s) 1022. In at least one embodiment, cloud 1026 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system(s) 1024 can use GPUs, and one or more AI system(s) 1024 can be used to perform cloud 1026 (or at least portions of tasks that are deep learning or inferencing). Likewise, although hardware 922 is illustrated as discrete components, this is not intended to be limiting, and any component of hardware 922 can be combined with, or utilized by, any other component of hardware 922.

[0133] In at least one embodiment, AI system 1024 can include a purpose-built computing system (e.g., a supercomputer or HPC) configured for inferencing, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPUs, RAM, storage, and / or other components, features, or functionality, AI system 1024 (e.g., NVIDIA’s DGX TM ) can also include software (e.g., a software stack) that can perform GPU-optimized using multiple GPUs 1022. In at least one embodiment, one or more AI systems 1024 can be implemented in cloud 1026 (e.g., in a data center) to perform some or all of AI-based processing tasks of architecture 1000.

[0134] In at least one embodiment, cloud 1026 can include a GPU-accelerated infrastructure (e.g., NVIDIA’s NGC TM ) that can provide a GPU-optimized platform for performing processing tasks of architecture 1000. In at least one embodiment, cloud 1026 can include AI system 1024 for performing one or more AI-based tasks of architecture 1000 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1026 can integrate with application coordination system 1028 that utilizes multiple GPUs to enable seamless scaling and load balancing between and among applications and services 920. In at least one embodiment, cloud 1026 can be responsible for performing at least some services 920 of architecture 1000, including compute services 1016, AI services 1018, and / or visualization services 1020, as described herein. In at least one embodiment, cloud 1026 can perform batched inference (e.g., perform NVIDIA’s TensorRT TM ), provide accelerated parallel computing APIs and platforms 1030 (e.g., NVIDIA’s ), execute application coordination system 1028 (e.g., KUBERNETES), provide graphics rendering APIs and platforms (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher quality cinematic effects), and / or can provide other functionality for architecture 1000.

[0135] In at least one embodiment, to protect patient confidentiality (e.g., in instances where patient data or records are used off-site), cloud 1026 can include a registry - e.g., a deep learning container registry. In at least one embodiment, the registry can store containers for instantiating applications that can perform pre-processing, post-processing, or other processing tasks on patient data. In at least one embodiment, cloud 1026 can receive data including patient data and sensor data in containers, perform only the requested processing on those sensor data in containers, and then forward results and / or visualizations to appropriate parties and / or devices (e.g., local medical devices for visualization or diagnosis) without extracting, storing, or otherwise accessing patient data. In at least one embodiment, patient data is kept confidential in accordance with HIPAA and / or other data regulations.

[0136] Other variations are within the spirit of the disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. It should be understood, however, that there is no intention to limit the disclosure to the specific form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure, as defined in the appended claims.

[0137] The use of the terms “a” and “one” and “the” and similar referents in the context of describing the disclosed embodiments (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be interpreted to mean one or more items from the list, but not limited to a single item from the list or the list as a whole, unless other wise indicated herein or clearly contradicted by context. The use of the term “one or more of’ followed by a list of one or more items (for example, “one or more of A and B”) is to be interpreted to mean one item from the list, or one or more items from the list, but not limited to a single item from the list or the list as a whole, unless other wise indicated herein or clearly contradicted by context. The use of the terms “includes,” “including,” “has,” “having” and the like are not intended to be exclusive or exhaustive in describing a process, method, article or apparatus, or any combination thereof. The terms “comprises,” “comprising,” “has,” “having,” “includes,” “including,” or “containing,” “containing” are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. The use of the terms “based on” and “based upon” are intended to be supporting and evidentiary in nature, rather than limiting. The use of any of the words “including,” “containing,” “comprising,” “consisting of,” “by” or the like, are not intended to exclude or omit other items from the claimed disclosure. The use of the terms “set,” “subset,” or the like, are not intended to be limiting in nature, unless otherwise indicated herein or clearly contradicted by context. The use of the terms “plurality” and “plural” are not intended to be limiting in nature, unless otherwise indicated herein or clearly contradicted by context.

[0138] Unless explicitly stated otherwise or apparent from context, a phrase such as "at least one of A, B, and C" or "at least one of A, B, or C" shall indicate that the group is inclusive of any of the items, elements, etc. individually or any combination of the items, elements, etc. For example, the phrases "at least one of A, B and C" and "at least one of A, B, or C" shall cover: (a) A alone, (b) B alone, (c) C alone, (d) at least one of A and B together, (e) at least one of A and C together, (f) at least one of B and C together, and (g) all of A, B, and C together. In other words, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" shall mean that the group is an inclusive- or group. In addition, unless otherwise stated or clear from context, the term "plurality" shall indicate a state of more than one (e.g., a plurality of items shall indicate more than one item). In at least one embodiment, a plurality of items shall indicate at least two items, but if explicitly indicated or otherwise clear from context, a plurality of items can indicate more than two items. Further, unless otherwise stated or clear from context, the phrase "based on" is intended to refer to the established fact that something is based on a combination of more than one thing. For example, "based on" is intended to mean "based, at least in part, on" unless explicitly stated otherwise or clear from context.

[0139] The operations of a process described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. In at least one embodiment, a process, such as those described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions to perform the operations of the process, and the processes can be implemented as code (e.g., executable instructions, one or more computer programs or one or more applications) executing collectively on one or more processing units, by hardware or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, such as a computer program product, which is readable by a computer system in the computing device. The code is executed by the computer system to perform the operations of the process. In at least one embodiment, the code is stored on one or more computer-readable storage medium, which is non-transitory computer-readable storage medium, excluding transitory, propagating signals.

[0140] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software to enable implementation of the operations. Moreover, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices operating in different manners such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.

[0141] The use of any and all examples, or exemplary language (e.g., "such as") provided herein, is intended merely to better illuminate embodiments of the disclosure and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0142] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0143] In the description and claims, the terms "coupled" and "connected," along with derivatives thereof, can be used. It should be understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, "connected" or "coupled" can be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" can also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other.

[0144] Unless specifically stated otherwise, it can be appreciated that throughout the specification terms such as "processing," "computing," "calculating," "determining," or the like, refer to the action and / or processes of a computer or computing system, or similar electronic

[0145] In a similar manner, the term "processor" can refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities such as tasks, threads, and intelligent agents that perform work over time. Likewise, each process can refer to multiple processes to sequentially or concurrently execute instructions, either continuously or intermittently. In at least one embodiment, the terms "system" and "method" can be used interchangeably herein so long as a system can embody one or more methods and a method can be considered a system.

[0146] In this document, obtaining, acquiring, receiving, or importing analog or digital data into a subsystem, computer system, or computer-implemented machine can be referenced. In at least one embodiment, the process of obtaining, acquiring, receiving, or importing analog and digital data can be accomplished in a variety of ways such as, for example, by receiving data as a parameter to a function call or call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or importing analog or digital data can be accomplished by transferring data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or importing analog or digital data can be accomplished by transferring data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, providing, outputting, transferring, sending, or presenting analog or digital data can also be referenced. In various examples, the process of providing, outputting, transferring, sending, or presenting analog or digital data can be accomplished by transferring data as an input or output parameter to a function call, a parameter to an application programming interface, or an interprocess communication mechanism.

[0147] Although the description herein sets forth example embodiments of the described technology, other architectures can be used to implement the described functionality, and are intended to be within the scope of this disclosure. Moreover, although specific distributions of responsibilities are defined above for the components, for purposes of description only, there is no intention to be bound by such distributions, as such components can define alternative structures that are within the scope of the disclosure.

[0148] Moreover, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A method comprising: obtaining a first prompt to obtain a responsive description of a media item for the first prompt; processing a representation of the media item and the first prompt using a content detection model to obtain a characterization of the media item; generating a second prompt using the first prompt and the characterization of the media item, the second prompt comprising instructions to a language model (LM); causing the LM to process the second prompt to generate computational code associated with the characterization of the media item; and causing the computational code to be executed to generate the responsive description of the media item for the first prompt.

2. The method of claim 1, wherein the first prompt comprises a natural language prompt.

3. The method of claim 1, wherein the representation of the first prompt comprises one or more keywords associated with the first prompt. selecting the content detection model from a plurality of trained models based on the representation of the first prompt.

4. The method of claim 1, further comprising:

5. The method of claim 4, wherein the content detection model comprises an object detection model trained to detect one or more objects associated with the representation of the first prompt.

6. The method of claim 4, wherein the content detection model comprises an open vocabulary model comprising: a computer vision portion to process at least the media item, a language understanding portion to process at least the characterization of the media item, and a classifier portion to process outputs of the computer vision portion and the language understanding portion to obtain the characterization of the media item.

7. The method of claim 1, wherein the media item comprises at least one of: an image item, a video item, an audio item, or a sensor data item.

8. The method of claim 1, wherein the characterization of the media item comprises: one or more bounding boxes for respective one or more objects in the media item.

9. The method of claim 1, wherein the instructions to the LM comprise: a natural language explanation of a format of the characterization of the media item.

10. The method of claim 1, wherein the LM is trained to generate computational code to perform a computational task in response to natural language instructions comprising a description of the computational task.

11. A system comprising: one or more processing units to: process a media item and a representation of a first prompt corresponding to the media item using a content detection model to obtain a characterization of the media item; generate a second prompt using the first prompt and the characterization of the media item, the second prompt comprising instructions to a language model (LM); and cause the LM to process the second prompt to generate computational code associated with the characterization of the media item, wherein the computational code, when executed, generates a responsive description of the media item for the first prompt.

12. The system of claim 11, wherein the representation of the first prompt comprises one or more keywords associated with the first prompt. ​ ​ 13. The system of claim 11, wherein the one or more processing units are further to use the representation of the first hint to select the content detection model from a plurality of trained models.

14. The system of claim 13, wherein the content detection model comprises at least one of: an object detection model trained to detect one or more objects associated with the representation of the first hint; or an open-vocabulary model comprising: a computer vision portion to process at least the media item, a language understanding portion to process at least the representation of the media item, and a classifier portion to process outputs of the computer vision portion and the language understanding portion to obtain the representation of the media item.

15. The system of claim 11, wherein the media item comprises at least one of: an image item, a video item, an audio item, or a sensor data item.

16. The system of claim 11, wherein the representation of the media item comprises: one or more bounding boxes for respective one or more objects in the media item.

17. The system of claim 11, wherein the instructions for the LM comprise: a natural language explanation of a format of the representation of the media item.

18. The system of claim 11, wherein the LM is trained to generate computational code to perform a computational task in response to natural language instructions comprising a description of the computational task.

19. The system of claim 11, wherein the system is comprised in at least one of: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing one or more medical operations; a system for performing one or more factory operations; a system for performing one or more analytics operations; a system for performing optical transport simulations; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using robots; a system for performing one or more conversational AI operations; a system implementing one or more large language models (LLMs); a system implementing one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

20. A non-transitory computer-readable memory having instructions stored thereon that, when executed by a processing device, cause the processing device to: generating a second prompt using the first prompt and a representation of a media item, the representation of the media item being obtained by applying the media item and the first prompt to at least one content detection model; and causing the LM to process the second prompt to generate a computational code associated with the representation of the media item, wherein, the invocation of the computing code generates a responsive description of the media item for the first prompt.