DISPLAYING IMAGES IN CHATBOT REPLIES
By associating text data with image identifiers and using a search function to include images in chatbot responses, the system addresses the resource-intensive challenge of multimodal models, enhancing chatbot interactions with reduced computational demands.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-03-26
AI Technical Summary
Traditional language models require significant computational resources for multimodal capabilities, making it impractical to implement multimodal models for chatbots that generate images or other media data.
A system that identifies and retrieves image identifiers during response generation, associating relevant text data with images, and uses a search function to include images in chatbot outputs, reducing the need for expensive multimodal models.
Enables the augmentation of chatbot responses with images without requiring additional computational resources, providing more context and enhancing user interaction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Traditional language models are trained / updated with large text databases to understand and generate natural language data. Certain language models, known as multimodal language models, are trained / updated to process information across different media modalities, including images, videos, and audio. However, due to the additional parameters used to implement multimodal capabilities, multimodal language models require significantly more computational resources for training / updating and execution. SUMMARY
[0002] Conversational agents or chatbots can be implemented using machine learning models, including large language models (LLMs), vision language models (VLMs), multimodal language models, and so on. Chatbots or conversational agents function by receiving a natural language text prompt as input and autoregressively generating a natural language text response to that prompt. Previously submitted prompts and generated responses can be used as context for further responses, allowing the large language model to function as a conversational chatbot, virtual agent, non-player character (NPC), digital avatar, and so forth.Although image or other media data can be helpful in supplementing text-based responses for language models, implementing traditional multimodal models for chatbots / conversational agents requires significantly more computing resources compared to typical LLMs.
[0003] To address the limitations of conventional approaches, the systems and methods described here enable the augmentation of responses from language models with additional media data without requiring the use of more expensive multimodal machine learning models. Identifiers of images or other media are identified and retrieved during the generation of a response by a large language model and used to provide images as part of the chatbot output. Media used to augment chatbot responses can be extracted from a repository of documents containing text data, image data, and / or other media data.
[0004] Text data located near images / media within documents is identified as relevant to those images / media, extracted, and stored with a unique identifier corresponding to the image / media data. When a prompt is received for the chatbot / conversational agent / NPC / digital avatar / etc., a search function can be used to identify relevant stored text data. Once relevant text data has been identified, the associated identifier of the image / media data can be used to retrieve the image / media data for inclusion in the chatbot output. The image / media data, along with text data generated using the chatbot / conversational agent's language model, can then be presented as output in response to the initial prompt to provide more context and assist the user during the conversation.
[0005] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of protection of the claims are described herein.
[0006] Systems and methods are disclosed that relate to displaying images in responses from chatbots / NPCs / virtual agents / digital avatars / etc. A system can identify text corresponding to an image in an electronic document and store a representation of the text associated with an identifier of the image. The system can receive an input prompt for a machine learning model. Using the machine learning model, the system can generate a response to the input prompt. The response can include the image resulting from the identification of the text representation using a search function and output from the machine learning model.
[0007] At least one aspect involves one or more processors. The one or more processors may comprise one or more circuits. The one or more circuits can identify text that corresponds to an image in an electronic document. The one or more circuits can store a representation of the text associated with an identifier of the image. The one or more circuits can receive an input prompt for a machine learning model. The one or more circuits can generate a response to the input prompt using the machine learning model, the response containing the image as a result of identifying the representation of the text using a search function and an output from the machine learning model.
[0008] In some implementations, one or more circuits can identify the text corresponding to the image by extracting the text near the image in the electronic document. In some implementations, one or more circuits can identify the text corresponding to the image by extracting a predetermined portion of the text near the image in the electronic document. In some implementations, one or more circuits can generate the representation of the text by providing the text as input for an embedding model.
[0009] In some implementations, one or more circuits can store the text representation in a vector database. In some implementations, one or more circuits can store the image in an image database, where the image is identified by its identifier. In some implementations, one or more circuits can identify a multitude of images using the search function and the output of the machine learning model. In some implementations, one or more circuits can select at least one of the multitude of images to include in the response, based on at least one image selection parameter.
[0010] In some implementations, one or more circuits can receive the image selection parameter along with the input prompt for the machine learning model. In some implementations, one or more circuits can present the output of the machine learning model, including the image, via a graphical user interface in response to the input prompt. In some implementations, the search function includes a vector similarity search function.
[0011] At least one aspect pertains to a system. The system can contain one or more processors. The system can receive an input prompt for a machine learning model. The system can generate a response message using the input prompt and the machine learning model. The system can identify coded text data using a search function and the response message, where the coded text data is stored in association with an image identifier. The system can provide the response message and the image for display as a response to the input prompt.
[0012] In some implementations, the coded text data includes embedding data. In some implementations, the search function is a vector search function. In some implementations, the system can identify a set of search results that contains the coded text data. In some implementations, the system can select the coded text data based on at least a similarity between the coded text data and the response message. In some implementations, the system can extract text data from an electronic document. The text data may be located near the image. In some implementations, the system can encode the text data to generate the coded text data. In some implementations, the system can store the image identifier in a database in association with the coded text data.In some implementations, the system can encode the text using an embedding model that corresponds to the machine learning model.
[0013] At least one aspect pertains to a procedure. The procedure may include identifying, using one or more processors, text that corresponds to media in an electronic document. The procedure may include storing, using one or more processors, a representation of the text associated with an identifier of the media. The procedure may include receiving, using one or more processors, an input prompt for a machine learning model. The procedure may include generating, using one or more processors, a response to the input prompt using the machine learning model, wherein the response includes the media as a reaction to the identification of the text representation using a search function and an output from the machine learning model.
[0014] In some implementations, the method may include identifying, using one or more processors, the text corresponding to the media by extracting the text near the media in the electronic document. The method may also include identifying, using one or more processors, the text corresponding to the media by extracting a predetermined portion of the text near the media in the electronic document. Finally, the method may include generating, using one or more processors, the representation of the text by providing the text as input to an embedding model.
[0015] The processors, systems, and / or methods described herein can be implemented by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing generative AI operations using a large language model; a system for performing generative AI operations using a video language model; a system implemented using an edge device; a system implemented using a robot; or a system for performing conversational AI operations.a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources.
[0016] The revelation extends to all novel aspects or features described and / or illustrated herein.
[0017] Further features of the disclosure are characterized by the independent and dependent claims.
[0018] Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0019] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0020] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.
[0021] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0022] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0023] The disclosure also includes a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out the procedures described herein and / or for embodying the device or system features described herein.
[0024] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0025] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0026] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0027] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The systems and methods presented here for displaying images in chatbot responses are described in detail below with reference to the accompanying diagrams. These diagrams show: Fig. 1 a block diagram of an exemplary system for displaying images in chatbot responses, according to some embodiments of the present disclosure; Fig. 2 an exemplary graphical user interface showing how images can be displayed in chatbot responses according to some embodiments of the present disclosure; Fig. 3 a flowchart of an exemplary method for displaying images in chatbot responses, according to some embodiments of the present disclosure; Fig. 4A is a block diagram of an exemplary generative language model system suitable for use in the implementation of at least some embodiments of the present disclosure; Fig. 4B a block diagram of an exemplary generative language model containing a transformer-encoder-decoder suitable for use in the implementation of at least some embodiments of the present disclosure; Fig. 4C a block diagram of a generative language model containing a decoder-only transformer architecture suitable for use in the implementation of at least some embodiments of the present disclosure; Fig. 5 a block diagram of an exemplary computing device suitable for use in the implementation of at least some embodiments of the present disclosure; and Fig. 6 a block diagram of an exemplary data center suitable for use in the implementation of at least some embodiments of the present disclosure. DETAILED DESCRIPTION
[0029] This disclosure concerns systems and methods for providing images, videos, and / or other media in responses from chatbots, conversational agents, NPCs, digital avatars, virtual assistants, etc. Conversational agents or chatbots can be implemented using machine learning models, such as large language models (LLMs). Such machine learning models are trained / updated using large repositories or volumes of text data to generate responses based on learned language patterns. Such conventional chatbots operate by receiving a natural language text prompt as input and autoregressively generating a natural language text response to the input text prompt. Previously submitted prompts and previously generated responses can be used as context for further responses, enabling the large language model to function as a conversational chatbot.
[0030] However, since the machine learning models used to implement conversational agents are trained / updated to operate using natural language text data, such models cannot generate images as output. This is a disadvantage for machine learning models that generate highly technical or nuanced responses, as visual aids may be more appropriate than, or significantly complement, a natural language response. While some machine learning models, such as vision-based transformer models or multimodal large language models, can produce synthetically generated images as output, such models typically require significant computing resources to train and run in computational environments.
[0031] The systems and methods of this disclosure address these problems by encoding identifiers for images that can be identified and retrieved during the generation of a response from a large language model. To provide images as part of the chatbot output, the systems and methods described herein can process a large repository of documents containing text data, video data, image data, audio data, and other media types. Text data located near the images in the documents can be considered relevant to those images. The relevant text can be extracted and assigned a unique identifier to the corresponding image. The images and text data extracted from the database can be stored for later retrieval, for example, using a search algorithm.
[0032] When generating the response, a search algorithm (such as a vector search operation) can be used to search the database for text entries related to the response generated by the chatbot's machine learning model. If a related text entry is identified, the identifier of one or more images stored in association with that text entry (e.g., extracted from documents) can be accessed and used to retrieve the one or more corresponding images provided with the response. The retrieved images can then be displayed alongside the machine learning-generated response via a graphical interface.
[0033] Referring to Fig. 1 is Fig. 1 An exemplary computing environment that includes a system for displaying images in chatbot responses, according to some embodiments of the present disclosure. It is understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, arrays, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in conjunction with other components, in any suitable combination and location. Various functions performed by entities described herein may be executed by hardware, firmware, and / or software.Various functions can be performed, for example, by a processor executing instructions stored in main memory. In some embodiments, the system and method described here can be implemented using one or more generative language models (e.g., as in ). Fig. 4A-4C), one or more computing devices or components thereof (e.g., as described in Fig. 5 described) and / or one or more data centers or components thereof (e.g. as described in Fig. 6) will be implemented.
[0034] System 100 can be used to process electronic documents 108 containing text and image data (e.g., multimedia data 110), so that the images / multimedia data can be included in relevant chatbot responses. The system is shown to include a data processing system 102, a memory 106, one or more databases 112, and one or more client devices 122. The data processing system 102 is shown to implement a document processor 118, a multimedia retriever 120, and a language model 121. The memory 106 is shown to store one or more electronic documents 108 and multimedia data 110. In some implementations, the multimedia data 110 can be stored in a storage device, system, or database separate from the memory 106 (which stores the electronic documents).The database 112 can be separate from or included as part of the memory 106 and is shown storing encoded text data 114 and multimedia identifiers 116. The data processing system 102 can receive input prompts 124 from the client device 122 and generate one or more output responses 126 according to the techniques described herein. The output responses 126 are shown containing a text response 128 and a multimedia response 130.
[0035] The data processing system 102 can include one or more processors, circuits, memory, and / or computing devices / systems capable of performing the various techniques described herein. For example, the data processing system 102 can be implemented in a cloud computing environment capable of maintaining, updating, and / or executing one or more language models 121. The data processing system 102 can implement the various techniques described herein to extract text data and multimedia data 110 from electronic documents 108 and to automatically select and provide relevant multimedia data 110 for inclusion in chatbot responses (e.g., the output responses 126).
[0036] As shown, in this example, the data processing system 102 is connected to the storage system 106. The storage system 106 can be an external server, a distributed storage / computing environment (e.g., a cloud storage system), or any other type of storage device or system connected to the data processing system 102. Although shown as external to the data processing system 102, it is understood that the storage system 106 can form part of the data processing system 102 or otherwise be located within it. The storage system 106 can be accessed in any of the several ways shown.
[0037] As shown, in this example, the data processing system 102 is associated with the database 112. The database 112 can be similar to the storage 106 and can be an external server, a distributed storage / computing environment (e.g., a cloud storage system), or any other type of storage device or system associated with the data processing system 102. Although shown as external to the data processing system 102, it is understood that the database 112 can form part of the data processing system 102 or otherwise reside within it. The database 112 can be any type of vector database capable of storing encoded text data 114 in association with one or more corresponding multimedia identifiers 116.
[0038] Memory 106 can store electronic documents 108, for example, in one or more data structures. Although shown as being stored in memory 106, it is understood that in some implementations, memory can also store identifiers of electronic documents 108, such as hyperlinks or network location identifiers that identify a network location of the electronic document 108 in one or more networks (e.g., a local area network, a wide area network, the internet, etc.). The electronic documents 108 can be any type of electronic document, which may contain both text and additional media data (e.g., images, audio, video, combinations thereof, etc.).
[0039] In some implementations, electronic documents can contain word processing files (e.g., DOCX, ODT), Portable Document Format (PDF) files, presentation files (e.g., PPTX, ODP), or spreadsheets (e.g., XLSX, ODS). In some implementations, electronic documents can also contain web pages (e.g., HTML, HTM), e-books (e.g., EPUB, MOBI), or Rich Text Format (RTF) files. Electronic documents can contain metadata or other formatting data that specifies the location of text data within electronic documents, as well as the one or more locations of multimedia data within the electronic documents, so that the data processing system (or its components) can determine the proximity of text information to multimedia data within the electronic documents.
[0040] The multimedia data 110 can contain any type / format of images, including, but not limited to, JPEG, PNG, GIF, BMP, or TIFF images, as well as vector images, such as those in SVG or EPS format. In some implementations, the multimedia data 110 of the electronic documents 108 can also contain other types of media in addition to images, including, but not limited to, audio data in formats such as MP3, WAV, AAC, or FLAC, and video data in formats such as MP4, AVI, MKV, or MOV, among others. In some implementations, the multimedia data 110 embedded in the electronic documents 108 can contain multimedia elements, such as three-dimensional (3D) and animated images, such as images in Graphics Interchange Format (GIF), and Web Media (WEBM) data.
[0041] The data processing system 102 can execute the document processor 118 to access electronic documents 108 and populate the database 112. The document processor 118 can include hardware, software, or a combination of hardware and software. The document processor 118 can access one or more of the electronic documents 108. In some implementations, the document processor 118 can process the electronic documents 108 in response to a request received from an external computing device (e.g., a client device 122, another external computing system) or in response to input received from an operator of the data processing system 102. The request can refer to one or more electronic documents 108 or to one or more locations (e.g., uniform resource identifiers (URIs), network locations, etc.).) from which one or more electronic documents 108 can be accessed. In some implementations, the data processing system 102 can perform web scraping of one or more servers, websites, or network locations to access and retrieve one or more electronic documents for processing.
[0042] To process an electronic document 108, the document processor 118 can parse the electronic document 108 to identify one or more elements of multimedia data 110. For example, the document processor 118 can parse the electronic document 108 to identify each element of multimedia data 110 (e.g., each image) in the file, as well as any text data located near the multimedia data element 110 (in this example, near the image). In some implementations, the document processor 118 can iterate through the electronic document 18 to extract embedded images, diagrams, and other media. The document processor 118 can access layout information (e.g., tags, metadata, the structure of the electronic document 108, etc.) to identify the relative distance between portions of text data and multimedia data 110 embedded in the electronic document 108.For example, the layout information can specify one or more relative or absolute locations within the document where text data, image data, or other multimedia data should appear.
[0043] After identifying text data located near the multimedia data 110, the document processor 118 can retrieve a portion of the text data to associate it with the multimedia data 110. As described herein, text data in electronic documents 108 located near multimedia data 110 can be assumed to be relevant to the corresponding multimedia data 110. In some implementations, if an element of multimedia data 110 (e.g., an image) is identified in an electronic document and no corresponding text data is present near the element of multimedia data 110, the document processor 118 may refrain from further processing with respect to the element of multimedia data 110. Otherwise, if text data is identified near the element of multimedia data 110, the document processor 118 may extract the multimedia data 110 and store it in memory 106, as shown.
[0044] The document processor 118 can also extract at least a portion of the nearby text data. The amount of text data extracted from the electronic document can be predetermined. For example, the document processor 118 can extract a predetermined number of characters, partial words, words, phrases, sentences, or paragraphs that are identified as being near the corresponding multimedia data element 110. In some implementations, the amount of extracted text data can correspond to the type of electronic document 108 in which the multimedia data element 110 was identified. For example, if the multimedia data element 110 is identified in a PDF, DOCX, DOC, RTF, or HTM / HTML document, the document processor 118 can extract a predetermined number of characters (e.g., 500 characters) near the multimedia data element 110.In another example, if the multimedia data element 110 is identified in a presentation document (e.g., PPT, PPTX, etc.), the document processor 118 can extract all text information on the same slide / page as the detected multimedia data 110. The rules / conditions for extracting text data from different types of multimedia data 110 can be stored in the configuration settings of the data processing system 102 and can be modified in response to requests from one or more external computing systems and / or input from one or more operators of the data processing system 102.
[0045] The document processor 118 can generate a unique identifier (e.g., a universally unique identifier (UUID), etc.) for each element of multimedia data 110 extracted from each electronic document 108. The unique identifier of the multimedia data 110 can be stored as part of the multimedia identifiers 116 in the database 112, associated with a corresponding set of coded text data 114. The document processor 118 can generate the multimedia identifier 116, for example, using a hash function, a timestamp, an identifier of the electronic document 108, combinations thereof, etc. The multimedia identifiers 116 can be stored in the database in such a way that they can be identified by searching the associated coded text data 114.
[0046] The document processor 118 can encode the text data extracted from the electronic document 108, which is associated with the multimedia data 110. Encoding the text data can include converting the text data into embeddings. In some implementations, encoding the text data can involve using one or more embedding models, which may include pre-trained transformer-based models that generate contextual embeddings capturing the semantic meaning of the text, and storing these embeddings as encoded text data 114. In some implementations, the text data can be encoded using word2vec or GloVe embedding to convert the words of the extracted text data into fixed-length vectors, which are then stored as vectors in the database 112. The encoded text data 114 are stored in the database 112 (e.g.,a vector database) stored, shown to retrieve one or more corresponding multimedia identifiers 116 from multimedia data 110 that were close to the corresponding text data. The coded text data 114 can be searched using a suitable similarity search function, as described in more detail herein, to identify coded text data 114 that closely match input prompts 124 provided by one or more client devices 122.
[0047] The document processor 118 can repeat this extraction process for each element of multimedia data 110 in a set of electronic documents 108 to populate the database 112 with coded text data 114 and associated unique multimedia identifiers 116. In some implementations, the document processor 118 can iterate over a large repository of electronic documents 108 covering a wide range of fields, topics, or other types of information. In some implementations, the document processor 118 can periodically, according to a schedule, or in response to one or more requests from external computing systems or one or more operators of the data processing system 102, perform web scraping of one or more sources of electronic documents 108 to update the database 112.In some implementations, the document processor 118 can manage (e.g., delete, modify, etc.) one or more entries in the database 112 in response to one or more requests from external computing systems or one or more operators of the data processing system 102. For example, automatic or manual verification processes can modify coded text data or exchange multimedia identifiers 116 to ensure that appropriate multimedia data 110 is returned as part of the output responses 126.
[0048] The data processing system 102 can receive and process input prompts 124 received from one or more client devices 122 using one or more speech models 121. The client device 122 can include any type of device capable of communicating with the data processing system 102 (e.g., via a network), including, but not limited to, smartphones, laptops or mobile computers, augmented and / or virtual reality devices, digital assistive devices, accessibility devices (e.g., hearing aids or devices, etc.), personal computers, servers, cloud computing systems, or other types of computing systems that can provide input prompts 124 to the data processing system 102.In some implementations, the client device 122 may contain one or more communication interfaces that enable the transmission of input prompts 124 to one or more external computing systems, which may contain the data processing system 102.
[0049] The input prompts 124 can contain any type of data that can be provided as input to one or more language models 121, including, but not limited to, text data, audio data, or video data, among others. In some implementations, the input prompts 124 may appear in response to one or more interactions with a graphical user interface (e.g., the graphical user interface of Fig. 2 etc.). The interactions can be provided via one or more input devices of the client device 122, such as a touchscreen, keyboard, mouse, or other input device. In some implementations, the input prompts 124 can be stored in one or more data structures on the client device 122. In some implementations, the client device 122 can run one or more applications that allow a user to provide data in one or more input formats, including but not limited to text, audio, or video, to serve as input for the language model 121. In some implementations, the application can include a front end for a conversational agent.
[0050] Input prompts 124, generated or retrieved by the client device 122, can be transmitted to the data processing system 102 for processing using the language model 121. In some implementations, the input prompts 124 can be provided via an input supplied by an operator of the data processing system 102. Upon receiving the input prompts 124, the data processing system 102 can convert them into a format compatible with one or more language models 121. For example, the data processing system 102 can execute one or more tokenizers to generate a sequence of tokens representing the one or more input prompts 124 in a numeric format compatible with one or more input layers of the language model 121.In some implementations, the sequence of tokens can be stored in association with input prompt 124.
[0051] The data processing system 102 can provide the sequence of input tokens representing the one or more input prompts 124 as input for the language model 121. The language model 121 can be any type of machine learning model that can be implemented as part of a chatbot or conversational agent. For example, the language model 121 can be or include a transformer-based model (e.g., a generative pre-trained transformer, GPT, model). In some implementations, the language model 121 can be or include an LLM or a vision language model (VLM). In some implementations, the language model 121 can include or be associated with one or more tokenizers that, as described herein, convert the input prompts 124 into a coded format (e.g.,a sequence of one or more tokens or a “tokenized” format) that is compatible with the layers of language model 121.
[0052] The data processing system 102 can execute the language model 121 by providing the tokenized input prompt 124 to one or more input layers of the language model 121. The data processing system 102 can perform the mathematical operations of each layer of the language model 121, passing the results of each layer to the next layer for processing until one or more output distributions of token probabilities are generated (e.g., from an output softmax layer, etc.). The data processing system 102 can use one or more configuration settings to select one or more tokens from the one or more output distributions to include in the output response. The data processing system 102 can autoregressively execute the large language model 121 to accurately model sequences of output tokens that correspond to natural language.For example, the data processing system 102 can execute the language model 121 to predict one or more next tokens in an output sequence, which can then be included in the input context for the next iteration, as described herein.
[0053] The data processing system 102 can iteratively execute the language model 121, using previously generated tokens as context for generating subsequent tokens until a termination condition is met. A termination condition can be a constraint on the context length or a configurable limit on the number of tokens that can be generated and / or processed by the language model 121. In some implementations, the termination condition can be met when the language model 121 generates a token that represents the end of a response. The language model 121 can be trained / updated to a conversational agent in some implementations.
[0054] The sequence of tokens generated by language model 121 can be stored as model output 123 once the termination condition is met. Model output 123 can be a sequence of output tokens generated by language model 121 that can be converted into a text format using a detokenization model. The detokenization model can be or include any software, hardware, or combination thereof that performs the reverse operations of the tokenizer associated with language model 121. When converted into a text format, model output 123 can be provided as a text response 128, which in this example is shown as part of the output response 126 provided in response to input prompt 124.In some implementations, the text response 128 may contain formatting instructions for displaying the text data generated by the language model 121, which may include Markdown formatting instructions, HTML instructions, or other formatting instructions for presenting the text data in the text response 128.
[0055] As shown, the output response 126 contains a multimedia response 130 in addition to the text response 128. The multimedia response 130 can contain relevant multimedia data 110, which is identified by the data processing system 102 based on the model output 123, the input prompt 124, or combinations thereof. The multimedia data 110 to be included as part of the multimedia response 130 can be identified by searching the database 112 using a search query to identify a corresponding multimedia identifier 116 that is relevant to the search query. To identify relevant multimedia data, the data processing system 102 can execute the multimedia retriever 120.
[0056] The multimedia retriever 120 can include any software, hardware, or combination thereof capable of performing one or more search queries against the database 112 to identify relevant coded text data 114. The multimedia retriever 120 can generate a search query against the database 112. In some implementations, the multimedia retriever 120 can use the text response 128, the input prompt 124, or both as the search query against the database 112. As described herein, the database 112 can be a vector database. Queries generated against the vector database can be determined using the same encoding process used to generate the coded text data 114, so that similarity numerical values between the query and the coded text data 114 can be calculated using a suitable search function.
[0057] As described herein, coded text data 114 can be generated from portions of text that are close to / relevant to multimedia data 110, which can be identified by corresponding multimedia identifiers 116. An entry in the database 112 contains coded text data 114 stored in association with one or more multimedia identifiers 116 for the relevant / closely related multimedia data 110. To identify relevant multimedia data 110 for inclusion in the output response 126, the multimedia retriever 120 can execute a search query against the database 112 to identify a set of similar / relevant coded text data 114.
[0058] To generate a search query for database 112, the multimedia retriever 120 can convert the text response 128 and / or the input prompt 124 into an encoded format (e.g., an encoded representation) using the same encoding process used to generate the encoded text data 114. As described herein, this may involve executing one or more embedding models, word2vec processes, or other functions to numerically encode the text response 128 and / or the input prompt 124 in a format that preserves their semantic meaning. In some implementations, the multimedia retriever 120 can generate a vector representation of the text response 128 and / or the input prompt 124 using a pre-trained language model (e.g., language model 121, another pre-trained language model, etc.).The vector representation can be used as a search query to retrieve relevant coded text data 114 from the database 112, based on the similarity between the vector representation of the search query and the coded text data 114.
[0059] The multimedia retriever 120 can use the generated vector representation (e.g., the search query) to search database 112 to identify relevant parts of coded text data 114. Executing the search against database 112 can involve comparing the vector representation of the search query with each entry of the coded text data 114 in database 112 to calculate a similarity score for each. The comparison can be performed using brute-force or approximate nearest neighbor (ANN) techniques. In some implementations, locality-sensitive hashing (LSH) can be implemented to improve the efficiency of searching the vector database 112.In some implementations, the vector space of database 112 can be partitioned into KD trees or sphere trees to improve the performance of nearest neighbor search techniques.
[0060] In some implementations, the comparison can be a distance calculation, where a shorter distance indicates a greater similarity between the search query and the entry in the coded text data 114, and a longer distance indicates a lesser similarity between the search query and the entry in the coded text data 114. In some implementations, the multimedia retriever 120 can identify the entry in the coded text data 114 in the database 112 with the greatest similarity (e.g., the shortest distance) to the search query. In some implementations, the multimedia retriever 120 can identify a set of top entries in the coded text data 114 in the database 112 with the greatest similarity (e.g., the shortest distance) to the search query.The number of top entries of coded text data 114 identified by the multimedia retriever 120 can be stored as a "Top-k" configuration setting, where k corresponds to the number of entries of coded text data 114 returned by the search. For example, the "Top-k" setting can specify the number of multimedia data elements 110 to be provided with the output response 126, and can be equal to one, two, or three multimedia data elements 110.
[0061] Once one or more entries of coded text data 114 have been identified, the multimedia retriever 120 can access the multimedia identifiers 116 associated with the identified entries of coded text data 114. In some implementations, multimedia identifiers 116 can be accessed / used for each of the one or more top entries of coded text data 114, regardless of the similarity numerical values between the top coded text data 114 and the search query. In some implementations, a multimedia identifier 116 can be accessed for top entries of the coded text data 114 that meet a threshold. The threshold can be stored as a configuration setting that can be modified via operator input into the data processing system 102 or specified in one or more requests from one or more external computing systems.In some implementations, a single multimedia identifier 116 can be assigned to an entry of encoded text data 114, and the multimedia retriever 120 can access it. In some implementations, multiple multimedia identifiers 116 can be assigned to an entry of encoded text data 114, and the multimedia retriever 120 can access it.
[0062] The multimedia retriever 120 can use the one or more multimedia identifiers 116 that have been accessed to retrieve corresponding elements of multimedia data 110. As described herein, the multimedia data 110 extracted from the electronic documents 108 can be stored in memory 106 and / or in one or more multimedia databases (e.g., one or more image databases). The memory and / or multimedia databases can store the multimedia data 110 in such a way that it can be retrieved using its corresponding unique multimedia identifiers 116. In some implementations, the multimedia identifiers 116 can be keys for the one or more databases that store the multimedia data 110.In some implementations, the multimedia retriever 120 can use the one or more accessed multimedia identifiers 116 as key values to access the multimedia data 110 from memory 106 (or, in some implementations, from the image database). In some implementations, the multimedia retriever 120 can use the one or more accessed multimedia identifiers 116 as part of one or more search queries to retrieve the corresponding multimedia data 110.
[0063] The multimedia data 110 accessed by the multimedia retriever 120 can be provided as the multimedia response 130 in the output response 126, as shown. The multimedia response 130 can contain one or all of the elements of multimedia data 110 retrieved by the multimedia retriever 120. In some implementations, the number of elements of multimedia data 110 provided can be specified via a configuration setting, which may be the same as or different from the top-k for searching the database 112, as described herein. In some implementations, the multimedia retriever 120 can generate the multimedia response 130 as one or more hyperlinks accessing the multimedia data 110 or display instructions to present the multimedia data 110. The hyperlinks or display instructions can cause the computing system (e.g.,The client device 122, which receives the multimedia response 130, retrieves and presents the multimedia data 110 identified by the hyperlinks or display instructions. In an implementation where one or more hyperlinks are provided, the computing system receiving the multimedia response 130 can display the hyperlink with one or more details about the name, type, or other attributes of the multimedia data 110 identified by the hyperlink.
[0064] Once generated, the output response 126 (including the text response 128 and the multimedia response 130) can be provided to one or more computing systems as a response to the input prompt 124. For example, the output response 126 can be provided to the client device 122 that transmitted the input prompt 124 and, in some implementations, stored in memory 106 in association with a recording of the input prompt 124. The computing system 102 can store / maintain a recording of one or more conversations in a log, which may contain sequences of one or more input prompts 124 and corresponding output responses 126 to implement a conversational agent / chatbot.In some implementations, the one or more input prompts 124 and the one or more corresponding output responses 126 can be stored in one or more historical conversation repositories for use in future training / update processes for the language model 121 or other machine learning models. An example of a graphical user interface displaying a sample output response 126 is provided in conjunction with... Fig. 2 described.
[0065] With reference to Fig. 2 in connection with the in connection with Fig. The components described in Section 1 illustrate an exemplary graphical user interface 200, which shows how image / multimedia data can be displayed in chatbot responses 204, according to some embodiments of the present disclosure. The graphical user interface 200 can be a web-based interface or an interface provided within a client application running on a client device 122. In some implementations, the graphical user interface 200 can be provided by the data processing system 102 and presented on a client device 122. For example, the graphical user interface can be provided directly via a web interface by the data processing system 102 and / or indirectly via a client application that is at least partially supported by the data processing system 102.
[0066] As shown, the graphical user interface 200 includes an area in which responses generated by the data processing system 102 are presented. The graphical user interface 200 includes a prompt input field 208 that can receive natural language prompts via user input. In some implementations, the prompt input field 208 can be populated using speech-to-text or other alternative input techniques. The graphical user interface 200 may include a submit button 210 which, when interacted with, causes the data processing system 102 to transmit the text / information in the prompt input field 208 to the data processing system 102.
[0067] Fig. Figure 2 shows the graphical user interface 200 after the input prompt 202 (e.g., an input prompt 124) has been transmitted, which in this example is: “What is a graphics card?”. In response to the input prompt 202, an output response 204 (e.g., an output response 126) is generated (e.g., by the data processing system 102 according to the techniques described herein) and displayed below the input prompt 202. The output response 204 is shown containing a text response (e.g., the text response 128) that, in this example, explains details about graphics cards in natural language. The text response can be generated by the language model 121, as described herein. In addition, the output response 204 contains one or more response images 206 (e.g., the multimedia response 130).
[0068] In some implementations, the multimedia data may contain one or more videos, audio files, and / or other types of multimedia data, instead of the multimedia data of the output response 204 being a single image. Although the response image 206 is shown as positioned below the text data of the output response 204, in some implementations the response image 206 may also be positioned anywhere within the output response 204 or the graphical user interface 200. In some implementations, instructions for displaying the one or more response images 206 (or other multimedia data), which may specify the size, position, or other attributes of the multimedia response 130, may be provided by the data processing system 102 with the output response 204.In some implementations, these instructions can be generated by the language model 121 using a corresponding input prompt that requests the generation of display instructions for the one or more response images 206.
[0069] In some examples, one or more of the machine learning models described herein (e.g., language models) can be packaged as a microservice, such as an inference microservice (e.g., NVIDIA's NIMs), which can contain a container (e.g., an operating system virtualization package) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or a model engine. For example, the inference microservice can contain the container itself and the model (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., when it has a sufficiently small number of parameters), the model can be contained within the container itself. In other examples, such as when the model is large, the model can be hosted in the cloud (e.g.,The model can be hosted / stored in a data center and / or hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside the container). In such embodiments, the model can be accessible via one or more APIs, such as REST APIs. Therefore, in some embodiments, the machine learning models described herein can be used as an inference microservice to accelerate the deployment of models in any cloud, data center, or edge computing system while ensuring data security. For example, the inference microservice can include one or more APIs, a pre-configured container for simplified deployment, and an optimized inference engine (e.g., built using standardized software for deploying and running AI models, such as...).NVIDIA's Triton inference server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that provide low latency and high throughput for production applications (such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The one or more machine learning models described herein may be included as part of the microservice along with accelerated infrastructure capable of being deployed with a single command and / or orchestrated and automatically scaled using a container orchestration system on accelerated infrastructure (e.g., on a single device up to the size of a data center). Therefore, the inference microservice may include the one or more machine learning models (e.g.,The inference microservice may include (optimized for high-performance inference), inference runtime software to execute one or more machine learning models and provide outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software to provide health checks, identity verification, and / or other monitoring. In some embodiments, the inference microservice may include software to perform an on-premises replacement and / or update of one or more machine learning models. During the replacement or update, the software performing the replacement / update may retain the user configurations of the inference runtime software and the enterprise management software.
[0070] In some embodiments, the system and methods described herein can be used in a talking or intelligent kiosk application. For example, a kiosk, tablet, smart display, or other device can include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing the model, image database, etc.). In some embodiments, the kiosk / tablet / display can communicate (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) with one or more locally hosted servers / computing devices and / or with one or more remote servers / computing devices (e.g., in one or more data centers).In such examples, the kiosk can communicate with the language model and / or image database hosted on the local and / or remote servers using one or more APIs - such as, without limitation, REST APIs.
[0071] In one or more embodiments, the system and methods described herein can be used in a gaming application. For example, a game console, PC, tablet, or other gaming device may include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing the game model, game assets, player data, etc.). These devices may use one or more machine learning models (e.g., language models) to enhance the gaming experience, generate dynamic content in real time, and personalize user experiences based on in-game behavior or pre-stored player profiles. In some embodiments, the system may be used in a cloud gaming environment. In such cases, a client device (e.g., a server) may be used to access the game's content.A smart display, tablet, or gaming controller can be used to interact with the game, while one or more speech models and / or visual rendering can take place on one or more remote servers / computing devices (e.g., in one or more data centers). The speech model and / or AI processing described herein can be run in the cloud, processing player input and generating corresponding in-game responses.
[0072] In some embodiments, the system and methods described herein can be used in a videoconferencing application. For example, a videoconferencing device, such as a dedicated conference unit, computer, tablet, or smartphone, may include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing video, audio, or other communication-related data). The system may use one or more language models to enhance videoconferencing functionality, including real-time transcription, speech translation, and background noise suppression. In one or more embodiments, the system may allow users to interact with the videoconferencing platform using natural language input.For example, users can give voice commands to schedule, join or leave meetings, or to manage participants and screen sharing.
[0073] In some embodiments, the system and methods described herein can be used in a robotics application. For example, a robot or robotic system may include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and main memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models) that enable it to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects or navigating environments using sensors, such as camera, LiDAR, radar, ultrasonic sensors, and more. The system may use sensor fusion techniques to combine data from multiple sensors (e.g.,The system combines various sensors (cameras, infrared, LiDAR, radar, accelerometers) to create a comprehensive model of the robot's environment. This data can be processed locally on the robot or sent to remote servers for more computationally intensive tasks, such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where centralized AI models can analyze it and distribute optimized commands to an entire fleet.
[0074] In some embodiments, the system and methods described herein can be used in an in-vehicle infotainment (IVI) system. For example, the infotainment system within a vehicle (e.g., cars, trucks, or autonomous vehicles) can include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and working memory and / or storage (e.g., for storing entertainment content, navigation data, and user preferences). The system can use these processors to execute one or more machine learning models (e.g., language models) to enable features such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services via network connectivity.The vehicle's infotainment system can also use natural language processing (NLP) models to enable voice-based interaction. The one or more machine learning models can be stored locally or accessed via one or more APIs connected to cloud services, allowing the system to process requests in real time.
[0075] Now, with reference to Fig. Each block of the Method 300 described herein contains a computational process that can be performed using any combination of hardware, firmware, and / or software. Various functions can be performed, for example, by one or more processors executing instructions stored in memory. The Method can also be embodied as computer-usable instructions stored on computer storage media. The Method can be provided by a standalone application, a service or hosted service (alone or in combination with another hosted service), or a plug-in for another product, to name just a few. Furthermore, the Method 300 is described, for example, in relation to the system of Fig. 1 described. However, this procedure can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.
[0076] Fig. Figure 3 is a flowchart illustrating a method 300 for displaying images in chatbot responses, according to some embodiments of the present disclosure. Method 300 includes, in block B302, the identification of text corresponding to an image in an electronic document (e.g., an electronic document 108). The electronic document 108 can be any type of document, including word processing documents (e.g., DOCX, DOC, PDF, etc.), presentation documents (e.g., PPT, PPTX, etc.), spreadsheet documents, web pages, or other types of documents. The electronic document can be parsed to identify each element of multimedia data (e.g., multimedia data 110) within the electronic document. In some implementations, the electronic document can contain both text data and multimedia data (e.g., image data).
[0077] Text corresponding to an image can be identified as text within the electronic document that is located near the image. Proximity can be determined in some implementations based on layout instructions or metadata specified in the electronic document. For example, the nearby text might include a caption or a relevant description of the image / multimedia data. Text identified as being near the image can be extracted from the electronic document for further processing. In some implementations, a predetermined number of characters, partial words, words, phrases, sentences, or paragraphs of the nearby text data can be extracted. For example, five hundred characters of the nearby text data can be extracted from the electronic document.
[0078] Furthermore, the image / multimedia data can be extracted from the electronic document and stored in a repository (e.g., storage 106, an image database, etc.). For each element of extracted image / multimedia data, an identifier (e.g., a multimedia identifier 116) can be generated to uniquely identify the image / multimedia data within the repository. The generated identifiers can be or contain UUIDs for each image or each element of multimedia data extracted from the electronic document.
[0079] Procedure 300, in block B304, involves storing a representation of the text (e.g., the coded text data 114) in association with an identifier of the image. The text data representation can be generated by encoding the text data extracted from the electronic document into a vector format and storing the coded text data in a vector database (e.g., database 112) in association with the identifier of the corresponding image / multimedia data. In some implementations, the coded text data can be generated by providing the text as input for an embedding model and / or a pre-trained language model. In some implementations, the coded text data can be generated using a word2vec model or a similar vectorization function for text information that preserves the semantic meaning of the text data.The encoded text data can be used as a key in the vector data for the unique identifier for the corresponding image / multimedia data.
[0080] Procedure 300, in block B306, includes receiving an input prompt (e.g., input prompt 124) for a machine learning model (e.g., language model 121). The input prompt can be received by a client device (e.g., client device 122) or by an operator of the computing system executing Procedure 300. The input prompts can contain any type of data that can be provided as input to a language model, including, but not limited to, text, audio, or video data. In some implementations, the input prompts may be received in response to one or more interactions with a graphical user interface (e.g., the graphical user interface of [missing information]). Fig. 2 etc.). The input prompts can be provided via a frontend for a chatbot or a conversational agent.
[0081] Procedure 300, in block B308, involves generating a response (e.g., output response 126) to the input prompt using the machine learning model. The response can be generated to include the image (e.g., multimedia response 130) as a reaction to identifying the text representation using a search function, and output from the machine learning model. For example, once an input prompt is received, it can be provided as input to the machine learning model to generate a text-based output (e.g., text response 128). The text response can then be used to identify relevant media data to be provided as part of the response.
[0082] To identify relevant media data, the text response and / or the input prompt can be encoded using the same encoding procedure used to generate the coded representation of text data extracted from electronic documents, as described herein. The encoded data is then used to perform a search query against the database (e.g., Database 112), which stores the encoded text data and the corresponding multimedia identifiers. The search function can be a vector search function that returns a predetermined number (e.g., according to the "Top-k" configuration setting) of results that identify corresponding multimedia identifiers. The one or more multimedia identifiers that are most similar to / relevant to the search query can be used to retrieve the corresponding multimedia data (e.g.,to retrieve the image from the repository where the extracted multimedia data is stored (e.g., memory 106, an image database, etc.). The retrieved images, along with the text response, can be provided as an output response message in response to the input prompt, as described in conjunction with [method / command]. Fig. 1 described. An example of the output response, which includes text and an image response, is shown in Fig. 2 shown.
[0083] The systems and methods described herein can be used for a variety of purposes, including but not limited to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational artificial intelligence (AI), light transport simulations (e.g., ray tracing, path tracing, etc.), collaborative content creation for three-dimensional (3D) assets, cloud computing, generative AI, and / or other suitable applications.
[0084] The disclosed embodiments can include a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing operations to generate synthetic data, systems implemented at least partially in a data center, systems for performing conversational AI operations, and systems implementing one or more language models, such as...one or more large language models (LLMs), systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems. EXEMPLARY LANGUAGE MODELS
[0085] In at least some embodiments, language models such as large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, video, computer-aided design (CAD) assets, omniverse and / or metaverse file information (e.g., in USD format, such as OpenUSD), and / or the like, based on the context provided in input prompts or queries. These language models may be considered "large" in certain embodiments because the models are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases), such as millions or billions of parameters.The LLMs / VLMs / MMLMs / etc. can be implemented for summarizing text data, analyzing and extracting insights from data (e.g., text, image, video, etc.), and generating new text / image / video / etc. in user-defined styles, tones, and / or formats. The LLMs / VLMs / MMLMs / etc. of this disclosure can, in some embodiments, be used solely for text processing, while in other embodiments, multimodal LLMs can be implemented to accept, understand, and / or generate text and / or other types of content such as images, audio, 2D and / or 3D data (e.g., in USD formats), and / or video. For example, vision language models (VLMs) or, more generally, multimodal language models (MMLMs) can be implemented to process image, video, audio, text, 3D design (e.g.,to accept CAD) and / or other input data types and / or to generate or output image, video, audio, text, 3D design and / or other output data types.
[0086] In various embodiments, different types of LLMs / VLMs / MMLMs / etc. architectures can be implemented. For example, different architectures can be implemented that use different techniques for understanding and generating output—such as text, audio, video, image, 2D and / or 3D design or asset data, etc. In some embodiments, LLMs / VLMs / MMLMs / etc. architectures such as recurrent neural networks (RNNs) or long-short-term memory (LSTM) networks can be used, while in other embodiments, transformer architectures such as... B. those based on self-awareness and / or cross-awareness mechanisms (e.g., between contextual data and textual data) are used to understand and detect relationships between words or tokens and / or contextual data (e.g., other textual, video, image, design data, USD, etc.).One or more generative processing pipelines containing LLMs / VLMs / MMLMs / etc. may also contain one or more diffusion blocks (e.g., denoisers). The LLMs / VLMs / MMLMs / etc. of this disclosure may contain one or more encoder and / or decoder blocks. For example, discriminative or encoder-only models such as BERT (Bidirectional Encoder Representations from Transformers) may be implemented for tasks requiring language understanding, such as classification, sentiment analysis, question answering, and named entity recognition. As another example, generative or decoder-only models such as GPT (Generative Pretrained Transformer) may be implemented for tasks involving the generation of speech and content, such as text completion, story generation, and dialogue generation.LLMs / VLMs / MMLMs / etc., which contain both encoder and decoder components such as T5 (Text-to-Text Transformer, Text-to-Text Transformer), can be implemented to understand and generate content, such as for translations and summaries. These examples are not intended as limitations, and any architecture type, including but not limited to those described herein, can be implemented depending on the specific implementation and the one or more tasks performed using the LLMs / VLMs / MMLMs / etc.
[0087] In various embodiments, LLMs / VLMs / MMLMs / etc. can be trained using unsupervised learning, where they learn patterns from large amounts of unlabeled text / audio / video / image / design / USD / etc. data. Due to this comprehensive training, the models in some embodiments may not require task- or domain-specific training. LLMs / VLMs / MMLMs / etc. that have undergone extensive pre-training with massive amounts of unlabeled data can be considered baseline models and may be suitable for a variety of tasks such as answering questions, summarizing, filling in missing information, translating, and generating images / videos / designs / USD / data. Some LLMs / VLMs / MMLMs / etc. can be further refined using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), and the addition of adapters (e.g.,custom neural networks and / or neural network layers that tune or adapt input prompts or tokens to align the language model with a specific task or domain) and / or are adapted to a specific use case using other fine-tuning or adaptation techniques that optimize the models for use in particular tasks and / or within particular domains.
[0088] In some embodiments, the LLMSs / VLMSs / MMLMs / etc. of this disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify inappropriate or unwanted inputs (e.g., prompts) and / or outputs of the models. The system can use the guardrails and / or other model alignment techniques to either prevent a specific unwanted input from being processed using the LLMs / VLMs / MMLMs / etc., and / or to prevent the output or presentation (e.g., display, audio output, etc.) of information generated using the LLMs / VLMs / MMLMs / etc. In some embodiments, one or more additional models—or layers thereof—can be implemented to identify problems with the inputs and / or outputs of the models.For example, these "protection models" can be trained to identify inputs and / or outputs that are "safe" or otherwise acceptable or desirable, and / or that are "unsafe" or otherwise undesirable for the particular application / implementation. As a result, the LLMs / VLMs / MMLMs / etc. of this disclosure are less likely to output language / text / audio / video / design / USD data / etc. that may be offensive, vulgar, inappropriate, unsafe, out of domain, and / or otherwise undesirable for the particular application / implementation.
[0089] In some implementations, the LLMs / VLMs / etc. can be configured to access or use one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations for which the model is not ideally suited, it may have instructions (e.g., as a result of training and / or based on instructions in a given prompt) to access one or more plug-ins (e.g., third-party plug-ins) to assist in processing the current input. In such an example, where at least part of a prompt relates to restaurants or the weather, the model may access one or more restaurant or weather plug-ins (e.g., via one or more APIs) to retrieve the relevant information.As another example, if at least part of an answer requires a mathematical calculation, the model can access one or more math plugins or APIs to assist in solving the one or more problems and can then use the plugin and / or API's response in the model's output. This process can be repeated, for example recursively, for any number of iterations and using any number of plugins and / or APIs until a prompt response can be generated that addresses every question / issue / requirement / process / operation, etc. Therefore, the one or more models can rely not only on their own knowledge gained from training on one or more large datasets but also on the expertise or optimization of one or more external resources, such as APIs, plugins, and / or the like.
[0090] In some embodiments, multiple language models (e.g., LLMs / VLMs / MMLMs / etc.), multiple instances of the same language model, and / or multiple prompts provided to the same language model or the same instance of the same language model can be implemented, executed, or accessed (e.g., using one or more plug-ins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, multiple language models, e.g., language models with different architectures, language models trained on different (e.g., updated) datasets, can be provided with the same input query and the same prompt (e.g., a set of constraints, conditioners, etc.).In one or more embodiments, the language models can be different versions of the same basic model. In one or more embodiments, at least one language model can be instantiated as multiple agents; for example, more than one prompt can be provided to restrict, direct, or otherwise influence the style, content, or character of the output provided. In one or more exemplary non-restrictive embodiments, the same language model can be prompted to provide output corresponding to a different role, perspective, character, or knowledge base, as defined, for example, by a provided prompt.
[0091] In each of these embodiments, the output of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated agents of at least one language model, and / or two further prompts provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output of a language model—or a version, instance, or agent—can be provided as input to another language model for further processing and / or validation. In one or more embodiments, a language model can be instructed to generate or otherwise obtain output with respect to input source material, the output being associated with the input source material.Such an assignment might involve, for example, generating a label or a portion of text that is embedded (e.g., as metadata) in input source text or image. In one or more embodiments, an output from a language model can be used to determine the validity of input source material for further processing or insertion into a dataset. For example, a language model can be used to evaluate the presence (or absence) of a target word in a portion of text or an object in an image, annotating the text or image to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a maintained dataset, for example, without restriction.
[0092] Fig. Figure 4A is a block diagram of an exemplary generative language model system 400, which is suitable for use in the implementation of some embodiments of the present disclosure. In the Fig. 4A illustrated example contains the generative language model system 400 a Retrieval-Augmented Generation (RAG) component 492, an input processor 405, a tokenizer 410, an embedding component 420, plug-ins / APIs 495 and a generative language model (LM) 430 (which may contain an LLM, a VLM, a multimodal LM, etc.).
[0093] At a high level, the input processor 405 can receive an input 401 comprising text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Descriptor (USD) data, such as OpenUSD, etc.), depending on the architecture of the generative LM 430 (e.g., LLM / VLM / MMLM / etc.). In some embodiments, the input 401 contains plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 401 can contain numerical sequences, pre-calculated embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML).In some implementations where the generative LM 430 is capable of handling multimodal inputs, the 401 input can combine text (or omit text) with image data, audio data, video data, design data, USD data, and / or other types of input data, such as, but not limited to, those described herein. Using raw input text as an example, the 405 input processor can prepare the raw input text in various ways. For instance, the 405 input processor can perform various types of text filtering to remove noise (such as special characters, punctuation, HTML markup, stop words, portions of one or more images, portions of audio, etc.) from relevant text content. In an example involving stop words (common words that tend to have little semantic meaning), the 405 input processor can remove stop words to reduce noise and allow the generative LM 430 to focus on more meaningful content.The 405 input processor can apply text normalization, for example by converting all characters to lowercase, removing accents, and / or handling special cases such as contractions or abbreviations to ensure consistency. These are just a few examples, and other types of input processing can also be applied.
[0094] In some embodiments, a RAG component 492 (which may contain one or more RAG models and / or be performed using the generative LM 430 itself) can be used to retrieve additional information to be used as part of the input 401 or the prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM / etc. with external knowledge, making the answers to specific questions, queries, or requirements more relevant, such as in a case where specific knowledge is needed. The RAG component 492 can retrieve this additional information (e.g., grounding information such as grounding text / image / video / audio / USD / CAD / etc.) from one or more external sources, which can then be fed into the LLM / VLM / MMLM / etc. along with the prompt to improve the accuracy of the model's responses or outputs.
[0095] For example, in some embodiments, the input 401 can be generated using the query or input for the model (e.g., a question, a requirement, etc.) in addition to data retrieved using the RAG component 492. In some embodiments, the input processor 405 can analyze the input 401 and communicate with the RAG component 492 (or the RAG component 492 can be part of the input processor 405 in some embodiments) to identify relevant text and / or other data that can be provided to the generative LM 430 as additional context or information from which the answer, response, or output 490 can generally be identified.For example, if the input indicates that the user is interested in a desired tire pressure for a specific make and model of vehicle, the RAG component 492—using a RAG model that performs a vector search in an embedding space—can retrieve the tire pressure information or the relevant text from a digital (embedded) version of the owner's manual for that specific vehicle make and model. Similarly, if a user revisits a chatbot about a particular product offering or service, the RAG component 492 can retrieve a previously saved conversation history—or at least a summary of it—and include the previous conversation history, along with the current question / request, as part of the 401 input in the generative LM 430.
[0096] The RAG component 492 can employ various RAG techniques. For example, naive RAG can be used when documents are indexed, split into pieces, and applied to an embedding model to generate embeddings that correspond to the pieces. A user request can also be applied to the embedding model and / or another embedding model of the RAG component 492, and the piece embeddings can be compared with the request's embeddings to identify the most similar embeddings that can be fed to the generative LM 430 to generate output.
[0097] In some embodiments, more advanced RAG techniques can be used. For example, the pieces can undergo pre-fetching processes (e.g., forwarding, rewriting, metadata analysis, extension, etc.) before being passed to the embedding model. Furthermore, post-fetching processes (e.g., re-ranking, prompt compression, etc.) can be performed on the outputs of the embedding model before the final embeddings are generated and used as a comparison to an input query.
[0098] As another example, modular RAG techniques can be used, such as those that are similar to naive and / or extended RAG, but may also include features such as hybrid search, recursive retrieval and query engines, step-back approaches, subqueries, and hypothetical document embedding.
[0099] As another example, Graph-RAG can use knowledge graphs as a source of contextual or factual information. Graph-RAG can be implemented using a graph database as a source of contextual information, which is then sent to the LLM / VLM / MMLM / etc. Instead of (or additionally) providing the model with snippets of data extracted from larger documents—which can result in a lack of context, factual accuracy, linguistic precision, etc.—Graph-RAG can also provide structured entity information to the LLM / VLM / MMLM / etc. by combining the structured text description of the entity with its many properties and relationships, allowing the model deeper insights. When implementing Graph-RAG, the systems and procedures described herein use a graph as a content store, extract relevant snippets from documents, and request them from the LLM / VLM / MMLM / etc.to respond using this. In such embodiments, the knowledge graph can contain relevant text content and metadata about the knowledge graph and be integrated into a vector database. In some embodiments, the graph RAG can use a graph as a domain expert, extracting descriptions of concepts and entities relevant to a query / prompt and passing them to the model as semantic context. These descriptions can include relationships between the concepts. In other examples, the graph can be used as a database, where part of a query / prompt can be allocated to a graph query, the graph query can be executed, and the LLM / VLM / MMLM / etc. can summarize the results.In such an example, the graph can store relevant factual information, and a query (natural language query) to a graph query tool (NL-to-Graph-query tool; NL for natural language) and an entity join can be used. In some implementations, graph RAG (e.g., using a graph database) can be combined with standard RAG (e.g., vector database) and / or other RAG types to benefit from multiple approaches.
[0100] In all embodiments, the RAG component 492 can implement a plugin, API, user interface, and / or other functionality to perform RAG. For example, a graph RAG plugin can be used by the LLM / VLM / MMLM / etc. to query the knowledge graph to extract relevant information for feeding into the model, and a standard or vector RAG plugin can be used to query a vector database. For example, the graph database can interact with a plugin's REST interface, thus decoupling the graph database from the vector database and / or the embedding models.
[0101] The Tokenizer 410 can segment (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, the tokens can represent individual words, partial words, characters, parts of audio / video / images, etc. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Partial word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, word stems), enabling the generative LM 430 to understand morphological variations and handle words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate token, allowing the generative LM 430 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as...The tokenizer 410 can convert the (e.g., processed) text into a structured format according to the tokenization scheme implemented in the respective embodiment.
[0102] The embedding component 420 can use any known embedding technique to convert discrete tokens into (e.g., dense, continuous vector) representations of semantic meaning. For example, the embedding component 420 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term frequency-inverse document frequency (TF-IDF) coding, one or more embedding layers of a neural network, and / or other techniques.
[0103] In some implementations where the input 401 contains image data / video data / etc., the input processor 401 can resize the data to a standard size compatible with the format of a corresponding input channel and / or normalize pixel values to a common range (e.g., 0 to 1) to ensure uniform representation. The embedding component 420 can encode the image data using any known technique (e.g., using one or more convolutional neural networks, CNNs, to extract visual features). In some implementations where the input 401 contains audio data, the input processor 401 can resample an audio file to a consistent sampling rate for uniform processing, and the embedding component 420 can use any known technique to extract and encode audio features—such as in the form of a spectrogram (e.g., a spectrogram).(of a Mel spectrogram). In some implementations where the input contains video data, the input processor can extract frames or apply resizing to extracted frames, and the embedding component can extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where the input contains multimodal data, the embedding component can fuse representations of the different data types (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.
[0104] The generative LM 430 and / or other components of the generative LM System 400 can use different types of neural network architectures, depending on the implementation. For example, transformer-based architectures, such as those used in models like GPT, can be implemented and may include self-attention mechanisms that weight the importance of different words or tokens in the input sequence and / or forward networks that process the output of the self-attention layers by applying nonlinear transformations to the input representations and extracting higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, crossmodal embedding models that learn shared embedding spaces, graph neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative adversarial networks (GANs) or adversarial autoencoders (AAEs) for joint distributional learning, and others. Depending on the implementation and architecture, the embedding component 420 can therefore apply a coded representation of the input 401 to the generative LM 430, and the generative LM 430 can process the coded representation of the input 401 to generate an output 490 that may contain responsive text and / or other data types.
[0105] As described herein, the generative LM 430, in some embodiments, may be configured to access or use plug-ins / APIs 495 (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations for which the generative LM 430 is not ideally suited, the model may have instructions (e.g., as a result of training and / or based on instructions in a given prompt, such as those retrieved using the RAG component 492) to access one or more plug-ins / APIs 495 (e.g., third-party plug-ins) to assist in processing the current input.In such an example, where at least part of a prompt relates to restaurants or the weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least part of the prompt relating to the respective plugin / API 495 to the plugin / API 495, the plugin / API 495 can process the information and return a response to the generative LM 430, and the generative LM 430 can use the response to generate the output 490. This process can be repeated, e.g., recursively, for any number of iterations and using any number of plugins / APIs 495 until an output 490 can be generated that addresses every question / request / request / process / operation / etc. from the input 401.Therefore, the models can rely not only on their own knowledge from training with one or more large datasets and / or on data retrieved using the RAG component 492, but also on the expertise or optimization of one or more external resources, such as the plug-ins / APIs 495.
[0106] Fig. Figure 4B is a block diagram of an example implementation where the generative LM 430 contains a transformer-encoder-decoder. Suppose, for example, that input text such as "Who discovered gravity?" (e.g., through the Tokenizer 410 in Fig. 4A) is tokenized into tokens, such as words, and each token (e.g., through the embedding component 420 in Fig. 94A) is encoded into a corresponding embedding (e.g., of size 512). Since these token embeddings typically do not represent the token's position in the input sequence, any known technique can be used to add positional encoding to each token embedding to encode the sequential relationships and context of the tokens in the input sequence. Therefore, the (e.g., resulting) embeddings can be applied to one or more encoders 435 of the generative LM 430.
[0107] In an exemplary implementation, the one or more Encoders 435 form an encoder stack, with each encoder containing a self-attention layer and a feedforward network. In an exemplary Transformer architecture, each token (e.g., a word) flows through a separate path. Therefore, each encoder can accept a vector sequence, with each vector passing through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used.For example, to calculate a self-attention score for each token (word), a query vector, a key vector, and a value vector can be generated for each token. A self-attention score can be calculated for token pairs by taking the dot product of the query vector with the corresponding key vectors, normalizing the resulting scores, multiplying them by the corresponding value vectors, and summing the weighted value vectors. The encoder can apply multi-headed attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. To generate a context vector that encodes the input, any number of encoders can be cascaded. An attention projection layer 440 can convert the context vector into attention vectors (keys and values) for one or more decoders 445.
[0108] In an exemplary implementation, the one or more decoders 445 form a decoder stack, with each decoder containing a self-attention layer, an encoder-decoder self-attention layer that uses the attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network. As with the one or more encoders 435, in an exemplary transformer architecture, each token (e.g., word) flows through a separate path in the one or more decoders 445. During a first pass, the one or more decoders 445, a classifier 450, and a generation mechanism 455 can generate an initial token, and the generation mechanism 455 can apply the generated token as input during a second pass. The process can be repeated in a loop, with successive tokens (e.g.,words) are generated and added to the output of the previous pass, and the token embeddings of the composite sequence with positional codes are applied as input to one or more Decoder 445s during a subsequent pass, generating one token at a time sequentially (known as autoregression) until a symbol or token representing the end of the response is predicted. Within each decoder, the self-attention layer is typically restricted to paying attention only to the preceding positions in the output sequence by applying a masking technique before the softmax operation (e.g., by setting future positions to negative infinity). In an example implementation, the encoder-decoder attention layer works similarly to the (e.g.,Multi-headed self-attention in the one or more 435 encoders, except that it generates its queries from the layer below and takes the keys and values (e.g., matrix) from the output of the one or more 435 encoders.
[0109] Therefore, the one or more decoders 445 can output a decoded (e.g., vector) representation of the input applied during a given pass. The classifier 450 can include a multi-class classifier, comprising one or more neural network layers that project the decoded (e.g., vector) representation into an appropriate dimensionality (e.g., one dimension for each supported word or token in the output vocabulary) and a softmax operation that converts logits into probabilities. The generation mechanism 455 can therefore select or sample a word or token based on an appropriate predicted probability (e.g., selecting the word with the highest predicted probability) and append it to the output of a previous pass, with each word or token being generated sequentially.The 455 generation mechanism can repeat the process, triggering successive decoder inputs and corresponding predictions, until a symbol or token is selected or sampled that represents the end of the response, at which point the 455 generation mechanism can output the generated response.
[0110] Fig. 4C is a block diagram of an exemplary implementation where the generative LM 430 incorporates a decoder-only transformer architecture. For example, one or more 460 decoders can be made from Fig. 4C similar to one or more 445 decoders. Fig. 4B work, with the exception that each of the one or more decoders 460 is out. Fig. In 4C, the encoder-decoder self-attention layer is omitted (since there is no encoder in this implementation). Thus, the one or more Decoder 460 can form a decoder stack, with each decoder containing a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or token representing the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., corresponding embeddings with positional encodings) can be applied to the one or more Decoder 460. As with the one or more Decoder 445 of Fig. 4B allows each token (e.g., word) to flow through a separate path in the one or more decoders 460, and the one or more decoders 460, a classifier 465, and a generation mechanism 470 can use autoregression to generate one token at a time until a symbol or token representing the end of the response is predicted. The classifier 465 and the generation mechanism 470 can be used similarly to the classifier 450 and the generation mechanism 455 of Fig. 4B operates wherein the generation mechanism 470 selects or samples each subsequent output token based on a corresponding predicted probability and appends it to the output from a previous pass, each token being generated sequentially until a symbol or token is selected or sampled that represents the end of the response. This and other architectures described herein are to be understood as examples only, and other suitable architectures may also be implemented within the scope of this disclosure. EXAMPLE CALCULATION DEVICE
[0111] Fig. Figure 5 is a block diagram of an exemplary computing device 500 suitable for use in implementing some embodiments of the present disclosure. The computing device 500 may include a connection system 502 that directly or indirectly couples the following devices: main memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., display(s)), and one or more logic units 520. In at least one embodiment, the one or more computing devices 500 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 508 can comprise one or more vGPUs, one or more of the CPUs 506 can comprise one or more vCPUs, and / or one or more of the logic units 520 can comprise one or more virtual logic units. Thus, a computing device 500 can contain discrete components (e.g., a complete GPU allocated to computing device 500), virtual components (e.g., a portion of a GPU allocated to computing device 500), or a combination thereof.
[0112] Although the various blocks of Fig. Where components 5 are shown as connected via the connection system 502, this is not intended as a limitation and serves only for clarity. In some embodiments, for example, a presentation component 518, such as a display device, can be considered an I / O component 514 (e.g., if the display is a touchscreen). As another example, the CPUs 506 and / or GPUs 508 can contain main memory (e.g., the main memory 504 can represent a storage device in addition to the main memory of the GPUs 508, the CPUs 506, and / or other components). Therefore, the computing device of Fig. 5 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are within the scope of protection of the computing device of Fig. 5 are being considered.
[0113] The 502 interconnect system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 502 interconnect system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the CPU 506 can be directly connected to the memory 504. Furthermore, the CPU 506 can be directly connected to the GPU 508.In a direct or point-to-point connection between components, the 502 connection system can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the 500 computing device.
[0114] The 504 main memory can contain a variety of computer-readable media. Computer-readable media can be any available media that the 500 computer can access. Computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.
[0115] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies; CD-ROM, Digital Versatile Discs (DVDs), or other optical disk storage; magnetic cartridges, magnetic tapes, magnetic disk storage, or other magnetic storage devices; or any other medium that can be used to store the desired information and that the computer can access.As used here, computer storage media do not inherently contain signals.
[0116] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media can include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included in the scope of protection of the computer-readable media.
[0117] The one or more CPUs 506 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to perform one or more of the procedures and / or processes described herein. The one or more CPUs 506 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a plurality of software threads simultaneously. The one or more CPUs 506 can contain any type of processor and can contain different types of processors depending on the type of computing device 500 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 500, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented with Reduced Instruction Set Computing (RISC), or an x86 processor implemented with Complex Instruction Set Computing (CISC). The computing device 500 can contain one or more CPUs 506, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0118] In addition to or as an alternative to the one or more CPUs 506, the one or more GPUs 508 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 508 may be an integrated GPU (e.g., with one or more of the CPUs 506) and / or one or more of the GPUs 508 may be a discrete GPU. In embodiments, one or more of the GPUs 508 may be a coprocessor of one or more of the CPUs 506. The one or more GPUs 508 may be used by the computing device 500 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the one or more GPUs 508 may be used for general-purpose computing on GPUs (GPGPU).The one or more GPUs 508 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The one or more GPUs 508 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the one or more CPUs 506 received through a host interface). The one or more GPUs 508 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the 504 main memory. The one or more GPUs 508 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 508 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.
[0119] In addition to or as an alternative to the one or more CPUs 506 and / or the one or more GPUs 508, the one or more logic units 520 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to perform one or more of the methods and / or processes described herein. In embodiments, the one or more CPUs 506, the GPUs 508, and / or the one or more logic units 520 may discretely or jointly execute any combination of the methods, processes, and / or sections thereof. One or more of the logic units 520 may be part of and / or integrated into one or more of the CPUs 506 and / or one or more of the GPUs 508, and / or one or more of the logic units 520 may be discrete components or otherwise separate from the CPUs 506 and / or the GPUs 508.In embodiments, one or more of the logic units 520 can be a co-processor of one or more of the CPUs 506 and / or one or more of the GPUs 508.
[0120] Examples of the one or more Logic Units 520 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), a Programmable Vision Accelerator (PVA), and one or more systems with direct memory access. (Direct Memory Access, DMA),may contain one or more vision or vector processing units (VPUs), one or more pixel processing engines (PPEs), e.g., including a 2D array of processing elements, each communicating north, south, east, and west with one or more other processing elements in the array, one or more decoupled accelerators or units (e.g., decoupled lookup table accelerators or units, vision processing units (VPUs), optical flow accelerators (OFAs), field programmable gate arrays (FPGAs), neuromorphic chips, quantum processing units (QPUs), associative processing units (APUs), arithmetic logic units (ALUs),Application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or similar.
[0121] The Communications Interface 510 can include one or more receivers, transmitters, and / or transceivers that enable the Computing Device 500 to communicate with other computers over an electronic network, including wired and / or wireless communication. The Communications Interface 510 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 520 and / or the communication interface 510 may contain one or more data processing units (DPUs) to transfer data received via a network and / or via the connection system 502 directly to one or more GPUs 508 (e.g., a working memory thereof).
[0122] The I / O ports 512 enable the computing device 500 to be logically coupled with other devices, including the I / O components 514, one or more presentation components 518, and / or other components, some of which may be built into (e.g., integrated with) the computing device 500. Illustrative I / O components 514 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 514 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as further described below) associated with a display on the Computing Device 500. The Computing Device 500 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 500 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 500 to render immersive augmented reality or virtual reality.
[0123] The power supply 516 can include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 516 can power the computing device 500 to enable the operation of the computing device 500's components.
[0124] The one or more presentation components 518 can include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more presentation components 518 can receive data from other components (e.g., the one or more GPUs 508, the one or more CPUs 506, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0125] Fig. Figure 6 illustrates an exemplary data center 600 that can be used in at least one embodiment of the present disclosure. The data center 600 can include an infrastructure layer 610, a framework layer 620, a software layer 630, and / or an application layer 640.
[0126] As in Fig. As shown in Figure 6, the infrastructure layer 610 of the data center can contain a resource orchestrator 612, clustered compute resources 614 and node compute resources (“node RRs”) 616(1)-616(N), where “N” is any positive integer. In at least one embodiment, the Node-RRs 616(1)-616(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid-state memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules, and / or cooling modules, etc. In some embodiments, one or more Node-RRs mayThe node RRs 616(1)-616(N) correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node RRs 616(1)-616(N) may contain one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the node RRs 616(1)-616(N) may correspond to a virtual machine (VM).
[0127] In at least one embodiment, the grouped computing resources 614 can contain separate groupings of node RRs 616, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node RRs 616 within grouped computing resources 614 can contain grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs 616, including the CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide computing resources to support one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.
[0128] The resource orchestrator 612 can configure or otherwise control one or more node RRs 616(1)-616(N) and / or grouped compute resources 614. In at least one embodiment, the resource orchestrator 612 can include a software design infrastructure (SDI) management entity for the data center 600. The resource orchestrator 612 can include hardware, software, or a combination thereof.
[0129] In at least one embodiment, as in Fig. As shown in Figure 6, the framework layer 620 can contain a job scheduler 628, a configuration manager 634, a resource manager 636, and / or a distributed file system 638. The framework layer 620 can contain a framework that supports the software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. The software 632 or the one or more applications 642 can each contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 620 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use a distributed file system 638 for processing large amounts of data (e.g., "Big Data"), but is not limited to it.In at least one embodiment, the job scheduler 628 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 600. The configuration manager 634 can be capable of configuring different layers, such as the software layer 630 and the framework layer 620, which contains Spark and the distributed file system 638, to support the processing of large amounts of data. The resource manager 636 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 638 and the job scheduler 628. In at least one embodiment, the clustered or grouped compute resources can include the grouped compute resource 614 on the infrastructure layer 610 of the data center.The Resource Manager 636 can coordinate with the Resource Orchestrator 612 to manage these allocated or assigned computing resources.
[0130] In at least one embodiment, the software contained in software layer 630 may include software 632 that is used by at least sections of the node RRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 638 of framework layer 620. One or more types of software may include, among others, web page search software, email virus scanning software, database software, and streaming video content software.
[0131] In at least one embodiment, the applications 642 contained in the application layer 640 may include one or more types of applications used by at least sections of the node RRs 616(1)-616(N), the grouped compute resources 614, and / or the distributed file system 638 of the framework layer 620. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0132] In at least one embodiment, a configuration manager 634, resource manager 636, and resource orchestrator 612 can implement any number and type of self-modifying actions based on any amount and type of data acquired in any technically feasible manner. Self-modifying actions can relieve a data center operator of data center 600 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0133] The Data Center 600 may contain tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to the Data Center 600.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 600 by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.
[0134] In at least one embodiment, the data center can use 600 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0135] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the one or more computing devices. Fig. 5. Implemented - e.g., each device may contain similar components, features, and / or functionality to one or more computing devices 500. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 600, an example of which is given herein with reference to Fig. 6 is described in more detail.
[0136] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0137] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described here can be implemented on any number of client devices with reference to one or more servers.
[0138] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more applications of an application layer. The software or the one or more applications may each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that uses, for example, a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.
[0139] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can offload at least some functionality to the one or more edge servers. A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0140] The one or more client devices can have at least some of the components, features, and functions of the one or more mentioned here in relation to Fig.The 500 exemplary calculating devices described are included.As an example, and not as a limitation, a client device may be embodied as a personal computer (PC), laptop, mobile device, smartphone, tablet computer, smartwatch, portable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or global positioning device, video player, video camera, surveillance device or surveillance system, vehicle, boat, flying boat, virtual machine, drone, robot, hand-held communication device, hospital device, gaming device or gaming system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0141] The revelation can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract data types. The revelation can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc.The revelation can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected to each other via a network for communication.
[0142] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0143] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms “step” and / or “block” may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.
[0144] The disclosure of this application also contains the following numbered clauses: Clause 1 One or more processors, comprising: one or more circuits, for: Identifying text that corresponds to an image in an electronic document; Storing a representation of the text in association with an identifier of the image; Receiving an input prompt for a machine learning model and Generating a response to the input prompt using the machine learning model, where the response includes the image in response to identifying the representation of the text using a search function and an output of the machine learning model. Clause 2 One or more processors according to Clause 1, wherein the one or more circuits serve the following purpose: Identifying the text that corresponds to the image by extracting the text near the image in the electronic document. Clause 3 One or more processors according to Clause 2, wherein the one or more circuits serve the following purpose: Identifying the text that corresponds to the image by extracting a predetermined portion of the text near the image in the electronic document. Clause 4 One or more processors according to any of the preceding clauses, wherein the one or more circuits serve to: Generating the text representation by providing the text as input for an embedding model. Clause 5 One or more processors according to any of the preceding clauses, wherein the one or more circuits serve to: Saving the text representation in a vector database and Storing the image in an image database, where the image in the image database is identified by the image identifier. Clause 6 One or more processors according to any of the preceding clauses, wherein the one or more circuits serve to: Identifying a large number of images using the search function and the output of the machine learning model and Select at least one of the many images to include in the answer, based on at least one image selection parameter. Clause 7 One or more processors according to Clause 6, wherein the one or more circuits serve the following purpose: Receiving the image selection parameter with the input prompt for the machine learning model. Clause 8 One or more processors according to any of the preceding clauses, wherein the one or more circuits serve to: Presenting the output of the machine learning model with the image via a graphical user interface in response to the input prompt. Clause 9 One or more processors according to any of the preceding clauses, wherein the search function includes a vector similarity search function. Clause 10 The one or more processors according to any of the preceding clauses, wherein the one or more processors comprise at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing operations using conversational AI; a system for performing operations with generative AI using one or more multimodal language models; a system for performing operations with generative AI using a large language model (LLM); a system for performing operations with generative AI using a video language model (VLM); a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 11 System, comprehensive: one or more processors, for example: Receiving an input prompt for a machine learning model; Generating a response message using the input prompt and the machine learning model; Identifying coded text data using a search function and the response message, wherein the coded text data is stored in association with an image identifier; and Providing the reply message and image for display in response to the input prompt. Clause 12 system according to Clause 11, wherein the coded text data includes embedding data and wherein the search function is a vector search function. Clause 13 System according to Clause 12, wherein the one or more processors serve the following: Identifying a set of search results that contain the coded text data; and Selecting the coded text data based at least on a similarity between the coded text data and the response message. Clause 14 System according to one of Clauses 11 to 13, wherein the one or more processors serve the following purpose: Extracting text data from an electronic document, with the text data located near the image; Encoding the text data to generate the coded text data; and Storing the image identifier in a database in association with the encoded text data. Clause 15 System according to Clause 14, wherein the one or more processors serve the following: Encoding the text data using an embedding model that corresponds to the machine learning model. Clause 16 System according to one of Clauses 11 to 15, wherein the system includes at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing operations using conversational AI; a system for performing operations with generative AI using one or more multimodal language models; a system for performing operations with generative AI using a large language model (LLM); a system for performing operations with generative AI using a video language model (VLM); a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 17 Procedure, comprehensive: Identify, using one or more processors, text that corresponds to media in an electronic document; Storing, using one or more processors, a representation of the text in association with an identifier of the media; Receiving, using one or more processors, an input prompt for a machine learning model and Generate, using one or more processors, a response to the input prompt using the machine learning model, wherein the response includes the media in response to identifying the representation of the text using a search function and an output of the machine learning model. Clause 18 Procedure according to Clause 17, furthermore comprehensive: Identify, using one or more processors, the text corresponding to the media by extracting the text near the media in the electronic document. Clause 19 Procedure according to Clause 18, furthermore comprehensive: Identify, using one or more processors, the text corresponding to the media by extracting a predetermined portion of the text near the media in the electronic document. Clause 20 Procedure according to one of Clauses 17 to 19, furthermore comprehensive: Generate, using one or more processors, the representation of the text by providing the text as input for an embedding model.
[0145] It is understood that the aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of protection of the claims.
[0146] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.
[0147] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited non-patent literature
[0000] What is a graphics card? In response to the input prompt 202, an output response 204 (e.g., an output response 126
[0067] ) is displayed.
Claims
[1] One or more processors, comprising: one or more circuits, for: Identifying text that corresponds to an image in an electronic document; Storing a representation of the text in association with an identifier of the image; Receiving an input prompt for a machine learning model and Generating a response to the input prompt using the machine learning model, where the response includes the image in response to identifying the representation of the text using a search function and an output of the machine learning model. [2] One or more processors according to claim 1, wherein the one or more circuits serve to: Identifying the text that corresponds to the image by extracting the text near the image in the electronic document. [3] One or more processors according to claim 2, wherein the one or more circuits serve to: Identifying the text that corresponds to the image by extracting a predetermined portion of the text near the image in the electronic document. [4] One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: Generating the text representation by providing the text as input for an embedding model. [5] One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: Saving the text representation in a vector database and Storing the image in an image database, where the image in the image database is identified by the image identifier. [6] One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: Identifying a large number of images using the search function and the output of the machine learning model and Select at least one of the many images to include in the answer, based on at least one image selection parameter. [7] One or more processors according to claim 6, wherein the one or more circuits serve to: Receiving the image selection parameter with the input prompt for the machine learning model. [8] One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: Presenting the output of the machine learning model with the image via a graphical user interface in response to the input prompt. [9] One or more processors according to any of the preceding claims, wherein the search function comprises a vector similarity search function. [10] The one or more processors according to any one of the preceding claims, wherein the one or more processors comprise at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing operations using conversational AI; a system for performing operations with generative AI using one or more multimodal language models; a system for performing operations with generative AI using a large language model (LLM); a system for performing operations with generative AI using a video language model (VLM); a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [11] System, encompassing: one or more processors, for example: Receiving an input prompt for a machine learning model; Generating a response message using the input prompt and the machine learning model; Identifying coded text data using a search function and the response message, wherein the coded text data is stored in association with an image identifier; and Providing the reply message and image for display in response to the input prompt. [12] System according to claim 11, wherein the encoded text data comprise embedding data and wherein the search function is a vector search function. [13] System according to claim 12, wherein the one or more processors serve to: Identifying a set of search results that contain the coded text data; and Selecting the coded text data based at least on a similarity between the coded text data and the response message. [14] System according to any one of claims 11 to 13, wherein the one or more processors serve to: Extracting text data from an electronic document, with the text data located near the image; Encoding the text data to generate the coded text data; and Storing the image identifier in a database in association with the encoded text data. [15] System according to claim 14, wherein the one or more processors serve to: Encoding the text data using an embedding model that corresponds to the machine learning model. [16] System according to any one of claims 11 to 15, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing operations using conversational AI; a system for performing operations with generative AI using one or more multimodal language models; a system for performing operations with generative AI using a large language model (LLM); a system for performing operations with generative AI using a video speech model (VLM); a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [17] Procedures, including: Identify, using one or more processors, text that corresponds to media in an electronic document; Storing, using one or more processors, a representation of the text in association with an identifier of the media; Receiving, using one or more processors, an input prompt for a machine learning model and Generate, using one or more processors, a response to the input prompt using the machine learning model, wherein the response includes the media in response to identifying the representation of the text using a search function and an output of the machine learning model. [18] The method of claim 17, further comprising: Identify, using one or more processors, the text corresponding to the media by extracting the text near the media in the electronic document. [19] The method of claim 18, further comprising: Identify, using one or more processors, the text corresponding to the media by extracting a predetermined portion of the text near the media in the electronic document. [20] Method according to any one of claims 17 to 19, further comprising: Generate, using one or more processors, the representation of the text by providing the text as input for an embedding model.