Fused vector storage for efficient retrieval enhanced AI processing

The API system automatically coordinates local and cloud service calls, simplifies the indexing and query operations of RAG applications, solves the complexity problems of existing technologies, and achieves efficient RAG processing, which is suitable for non-expert users.

CN120653724APending Publication Date: 2025-09-16NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510306254.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-03-14
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing Retrieval Augmentation Generation (RAG) techniques come at a significant cost in terms of developer time and processing resources, and indexing and querying operations are complex, involving calls to multiple local and cloud services, which is beyond the capabilities of many users.

Method used

Provides an API system and technology that simplifies the development process of RAG applications by automatically coordinating local and cloud-based service calls, encapsulating indexing and query operations, and allowing users to select preset document processing pipelines (DPPs) for automated operations, reducing dependence on professional knowledge.

Benefits of technology

It implements efficient and automated RAG processing, reduces development complexity, frees up local computing and storage resources, and is suitable for non-expert users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653724A_ABST
    Figure CN120653724A_ABST
Patent Text Reader

Abstract

In various examples, systems and techniques are provided that encapsulate indexing and query operations as application programming interfaces (APIs) that automate and coordinate calls to various local and cloud-based services. When a user has documents to be added to a retrieval enhancement generation (RAG) database, an API may provide multiple document processing pipelines (DPPs) with a preset index configuration to the user. Similarly, when a user query is received, the API may generate a call to implement query processing that does not require the user to manually configure the embedded retrieval and processing. The API may further enable localization of relevant embedded stores and provision of stored embedding along with query embedding to calls of a search engine identifying the most relevant match. The API may then access the index embedded into the text and identify relevant text segments and documents to a hint generator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 566,105, filed on March 15, 2024, entitled “FUSED VECTOR STORE FOR EFFICIENT RETRIEVAL-AUGMENTED AI PROCESSING,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] At least one embodiment is directed to facilitating efficient data processing for artificial intelligence (AI) systems. For example, at least one embodiment is directed to utilizing stored vectors representing document content and other data related to AI operations to enhance AI processing. Background Art

[0004] Well-trained language models (e.g., large language models (LLMs)) are able to support natural language conversations, understand the speaker's intent and emotions, interpret complex topics, generate new text after receiving appropriate prompts, provide suggestions on topics of interest to the user, process images, audio, and / or other data types, and / or perform other functions. LLMs are typically self-supervised trained on large amounts of text data and / or other data types, depending on the embodiment, and learn to predict the next and / or missing tokens (which may correspond to subwords, symbols, words, etc.) in a phrase / sentence, detect the human speaker's intent and / or emotions, determine whether two sentences are related or unrelated, and / or perform other basic language tasks. After initial training, LLMs are typically subjected to guided (prompt-based) supervised fine-tuning, which enables the LLM to acquire deeper language capabilities and / or master more specialized tasks. Supervised fine-tuning includes using learning prompts (questions, hints, etc.) with example text (e.g., answers, example articles, etc.) as training ground truth. In enhanced fine-tuning, human evaluators assign grades that indicate how similar the generated text is to human-generated text. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a block diagram of an example computing architecture capable of implementing efficient retrieval augmentation generation (RAG) for augmenting input to an AI model, according to at least one embodiment;

[0006] Figure 2 An example computing device supporting efficient RAG for AI model processing is shown in accordance with at least one embodiment;

[0007] Figure 3shows an example data flow for the document indexing phase of efficient RAG for AI model processing according to at least one embodiment;

[0008] Figure 4 illustrates an example data flow for the query processing phase of efficient RAG for AI model processing according to at least one embodiment;

[0009] Figure 5 is a flow chart of an example method for an indexing phase of an efficient API-facilitated RAG for AI model processing according to at least one embodiment;

[0010] Figure 6 is a flow chart of an example method for the query processing phase of a RAG for efficient API facilitation of AI model processing in accordance with at least one embodiment;

[0011] Figure 7A illustrating reasoning and / or training logic according to at least one embodiment;

[0012] Figure 7B illustrating reasoning and / or training logic according to at least one embodiment;

[0013] Figure 8 illustrates the training and deployment of a neural network according to at least one embodiment;

[0014] Figure 9 is an example data flow diagram of a high-level computing pipeline according to at least one embodiment; and

[0015] Figure 10 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment. DETAILED DESCRIPTION

[0016] The training of language models (LMs), including large language models (LLMs) and / or visual language models (VLMs), typically involves a large amount of training data (e.g., human-generated text) and teaching the LM to generate responses to user queries, including questions, information requests, recommendations, explanations on a variety of general and specialized topics, images, videos, audio, digital assets, and / or the like. Because the number of topics that a user may be interested in is effectively infinite, LMs are often tasked with responding to queries about things and concepts that are not widely represented in the training data. Such queries can result in suboptimal responses that are incorrect and / or misleading.

[0017] Retrieval-augmented generation (RAG) is a technique that improves the output of a LM by augmenting the LM input (query) with additional information that may be relevant to the input (e.g., information including context, data, expertise on the subject of the query, etc.). This additional information can be stored in the form of embeddings (feature vectors or vectors) in a special N-dimensional embedding space that encode words, subwords, characters, etc. and their contextual connections. A trained encoder can encode text strings into embeddings that can be viewed as points in the embedding space. During training, the encoder learns to associate similar text strings with similar embeddings corresponding to points that are closer in the embedding space, and further learns to associate dissimilar text strings with points that are farther apart in the embedding space. Contextual information relevant to a particular LM query can be presented in the form of such embeddings, which can be previously generated for faster retrieval and stored in a suitable data store.

[0018] To utilize RAG techniques, received LM queries can similarly be converted into embeddings, each embedding encoding a portion of the input query. These query embeddings can then be compared with embeddings stored in a data store. For example, similarity factors (e.g., scalar products) of pairs of embeddings can be calculated, and the set of stored embeddings that most closely correlates with the query embedding can be identified. The document portions corresponding to the identified embeddings can then be included in user-generated queries as part of the contextual information to generate prompts that serve as input to the LM, causing the LM to produce more relevant and accurate responses. However, existing RAG techniques come at a significant cost in terms of developer time and processing resources.

[0019] More specifically, during the indexing phase of building a RAG-enhanced application, a user or developer can first select a document set related to a specific knowledge domain and then run code (e.g., code based on LangChang or Llamaindex) on a local (e.g., user) computer to split the documents into segments of a desired size (these segments can overlap). The document segments can be uploaded to, for example, a cloud-based embedding model, which converts the segments into embeddings that are then downloaded to the local computer. The embeddings can then be uploaded to the user's cloud space and stored in a vector database for subsequent user query access. The embeddings in the vector database can be indexed to the original text segments, which can be stored (e.g., also in the cloud) in a separate text store. During the query phase of RAG processing, the embeddings can be downloaded again to the local computer and compared with the embedding representing the user query. After identifying the most relevant (e.g., most similar to the query embedding) downloaded embeddings, the embeddings in the text index can be used to identify corresponding segments of the stored contextual information. These most relevant segments can be downloaded from the text store to the local computer, and LM hints can be generated, for example, using the original user query and the downloaded segments. The generated prompts can then be uploaded to the LM service for processing.

[0020] The operations described during the indexing and querying phases are complex and involve calls to multiple local and cloud services. Coding and optimizing such calls requires significant developer effort, coding experience, and knowledge of various RAG tools, which may be beyond the capabilities of many users.

[0021] Aspects and embodiments of the present disclosure address these and other challenges faced by users and developers of RAG-facilitated applications by providing systems and techniques that encapsulate indexing and query operations into an API that automates and coordinates calls to various local and cloud-based services. In some embodiments, when a user has a document (or multiple documents) to add to a RAG database, the API can provide the user with multiple document processing pipelines (DPPs) with preset indexing configurations. For example, a given DPP can have preset sizes of text segments (in characters, words, sentences, lines, and / or the like), the amount of segment overlap, the number K of top search matches to return, and / or the like. The DPP can further specify the embedding model to be used with the segment, including the cloud-based location of the embedding model, the embedding-to-text mapping scheme, the storage location of the embedding and / or text segment, and / or other relevant information. The user-selected (or default) DPP can automatically execute indexing instructions based on preprogrammed calls. In particular, the selected DPP can segment, embed, store in embedding and text data stores, and / or the like, the document. The generated embedding can then be indexed to the corresponding text segment. Similarly, when a user query is received, the API can generate calls to implement query processing that does not require the user to manually configure the retrieval and processing of embeddings. The API can further implement calls to locate relevant embedding stores and provide the stored embeddings along with the query embeddings to a search engine that identifies the most relevant matches. The API can then access the embeddings in the text index and identify the most relevant text snippets and documents for the hint generator.

[0022] A LM prompt generated using the original user's query and the recognized text snippets can be provided to the LM. In some embodiments, the K top matches (most relevant snippets) can be included in the prompt in an unranked form. In some embodiments, the K top matches can be provided to the LM as a ranked list. In some embodiments, before generating the LM prompt, an additional model can be used to rerank the top matches based on their relevance to the user's query. In some embodiments, before the LM model generates a final prompt using the reranked matches, a preliminary prompt for the same LM (or a separate LM) can be used to cause the LM to rerank the top matches. In some embodiments, any, some, or all indexing and / or query operations can be performed on the cloud, such as a user's cloud space. For example, a service that supports the RAG infrastructure and provides an API can also provide a fully functional and portable container (e.g., a Docker container) that can operate on a user's local computer or in the user's cloud space. This container can include any, some, or all of the APIs, API calls, DPP, segmentation engine, embedding model, search engine, reranking model, etc., and can facilitate the secure execution of RAG applications independently (in parallel) from other user-run applications, while maintaining its independence from other environments that can run in parallel.

[0023] Advantages of the disclosed embodiments include, but are not limited to, efficient and automated operation of indexing and querying RAG processing that coordinates calls to various processing and memory resources and does not have the complexity of traditional techniques that rely on user expertise and proficiency. Because the disclosed embodiments do not rely on RAG code and / or local processing, the intermixing of RAG code with application code that implements specific RAG facilitation is avoided. In some embodiments, the disclosed APIs and other techniques facilitate moving RAG operations to the cloud, which frees up local compute and memory resources for other tasks.

[0024] Figure 1 is a block diagram of an example computing architecture 100 that can implement efficient retrieval augmentation generation (RAG) for augmenting the input of an AI model, according to at least one embodiment. Figure 1 As depicted, computer architecture 100 may include a RAG client infrastructure 102, a RAG server 120, a data store 150, and an AI service 160 connected via a network 140. Network 140 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wireless network, a personal area network (PAN), combinations thereof, and / or other network types.

[0025] The RAG client infrastructure 102 may include one or more computing devices accessible to user 101 and facilitating user 101's efficient use of one or more AI models provided by AI services 160. In one example, non-limiting embodiment, the AI ​​model may include LM 162, but it should be understood that services associated with various other AI models, such as automatic speech recognition (ASR) models, computer vision (CV) models, text-to-speech models, anomaly detection models, action detection models, object detection models, and / or any other suitable generative or discriminative AI models, may similarly be improved using the disclosed techniques. In some embodiments, the RAG client infrastructure 102 may include client devices 104, which may be (or include) one or more computing devices under the control of user 101, such as desktop computers, laptops, smartphones, tablets, servers, wearable devices, virtual / augmented / mixed reality headsets or heads-up displays, digital avatar or chatbot kiosks, in-vehicle infotainment computing devices, and / or any other suitable computing devices capable of executing the techniques described herein. User 101 may be a person (e.g., an individual user) or an organization (e.g., a collective of users). The client device 104 may include a memory and one or more processors communicatively coupled to the memory (for simplicity, Figure 1 ), used to support local computing on the client device 104.

[0026] In some embodiments, the client device 104 may provide a user interface (UI) 106 for supporting receiving documents (or other data) and user queries (or other inputs to the AI ​​model) from the user 101 and providing responses to the user queries (or any other outputs of the AI ​​model) to the user 101. The UI 106 may include one or more devices of various modalities, such as a keyboard, a touch screen, a touchpad, a writing tablet, a graphical interface, a mouse, a stylus, and / or any other pointing device capable of selecting words / phrases displayed on the screen, and / or some other suitable device. In some embodiments, the UI 106 may include an audio device, such as a combination of a microphone and a speaker, a video device, such as a digital camera for capturing an image or a sequence of two or more images (video frames). In some embodiments, the text, voice, and / or video input devices may be integrated together (e.g., as part of a smartphone, tablet computer, desktop computer, and / or similar device).

[0027] Client device 104 may enable user 101 to access cloud server 112, which implements cloud-based computing, data storage, data authentication, and / or any other services that may be provided to user 101 as part of a paid or free subscription. Any suitable cryptographic protection technology may be used to protect the processing and storage of data on cloud server 112, including but not limited to symmetric and asymmetric key cryptography, digital certificates, and / or the like.

[0028] The cloud server 112 may deploy one or more computing devices, which may include memory 105 (e.g., one or more memory devices or units), which is communicatively coupled to one or more processing devices, such as one or more graphics processing units (GPUs) 110, one or more central processing units (CPUs) 130, one or more data processing units (DPUs), one or more parallel processing units (PPUs), and / or other processing devices (e.g., field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and / or the like). The memory 105 may include read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM), static memory (e.g., static random access memory (SRAM)), and / or some other memory capable of storing digital data. The cloud server 112 may support the execution of an application 107, which may be a text processing application, a video processing application, an audio processing application, a gaming application, an image or video rendering application, a computing application, a data processing application, a browsing application, and / or any other suitable application. Application 107 may be provided remotely to user 101 via UI 106 and client device 104 .

[0029] The RAG server 120 may deploy one or more processing devices ( Figure 1 100 ). In some embodiments, the RAG server 120 may be operated as part of the AI ​​service 160 (e.g., under the control of the same entity). In some embodiments, the RAG server 120 may be operated independently of the AI ​​service 160, but may use the AI ​​service 160 to support reasoning operations of clients (such as user 101).

[0030] In some embodiments, AI service 160 may deploy LM 162, which may be a large language model (e.g., a model with at least 100K learnable parameters) and / or a visual language model (VLM). LM 162 may be trained by training engine 164. In some embodiments, LM 162 may be trained in multiple stages. Initially, training engine 164 may train LM 122 to capture the syntax and semantics of human language, for example, by training it to predict the next, previous, and / or missing word in a sequence of words (e.g., one or more sentences of human speech or text). LM 162 may be further trained using training data containing a large amount of text (e.g., human conversations, newspaper text, magazine text, book text, web-based text, and / or any other text). Because the ground truth for such training is embedded in the text itself, training engine 164 may perform self-supervised training on LM 162 using such text. This teaches LM 162 how to converse with a user (human user or another computer) in natural language in a manner very similar to a conversation with a human speaker, including understanding the user's intent and responding in the manner the user expects from a conversation partner. After the initial self-supervised training, training engine 164 can perform supervised fine-tuning on LM 162 to teach LM 162 more specialized language skills, including expertise in a particular knowledge domain.

[0031] LM 162 can be implemented using a neural network with a large number of (e.g., billions of) artificial neurons. In at least one embodiment, LM 162 can be implemented as a deep learning neural network with multiple levels of linear and nonlinear operations. For example, LM 162 can include a convolutional neural network, a recurrent neural network, a fully connected neural network, a long short-term memory (LSTM) neural network, a neural network with an attention mechanism (e.g., a transformer neural network), a combination of a convolutional network and one or more transformers (conformer) and / or other types of neural networks. In at least one embodiment, LM 162 can include a plurality of neurons, wherein individual neurons receive their inputs from other neurons and / or from an external source, and generate an output by applying an activation function to the sum of weighted inputs (using trainable weights) and possible bias values. In at least one embodiment, LM 162 can include a plurality of neurons arranged in layers, including an input layer, one or more hidden layers, and / or an output layer. Neurons from adjacent layers can be connected by weighted edges.

[0032] Initially, some starting (e.g., random) values ​​may be assigned to the parameters of LM 162 (e.g., edge weights and biases). For each training input, the training engine 164 may cause LM 122 to generate a training output. The training engine 164 may then compare one or more training outputs with a desired target output. The resulting error or mismatch (e.g., the difference between one or more target outputs and one or more training outputs) may be back-propagated through the various neural layers of LM 162, and the weights and biases of LM 162 may be adjusted to bring the training output closer to the target output. In some embodiments, the training engine 164 may train multiple LMs 162 for multiple tasks (e.g., multiple different knowledge domains).

[0033] In some embodiments, the RAG server 120 may include one or more RAG APIs 108 that provide a user with a set of commands that a non-expert user can understand and that can implement any, some, or all of the operations of the document indexing phase and / or the query processing phase to enhance AI processing. The commands available via the RAG API 108 may include selecting a specific document processing pipeline (DPP) with a preset indexing and / or query processing configuration. The RAG service infrastructure may define any number of preset DPPs. Once selected, the DPP may implement various processing operations with minimal or no user input, including but not limited to issuing calls to operate the document segmentation engine 124 to segment the input document according to the DPP settings, operating the embedding model 126 to convert document fragments into feature vectors (embeddings), and operating the search engine 128 to perform a search in the embedding space to find relevant document fragments.

[0034] In some embodiments, an API package with the RAG API 108 can be downloaded to one of the computing devices of the cloud server 112. The downloaded API package can be used to install the RAG API 108 to enable the user 101 to deploy enhanced AI processing on, via, or using the cloud server 112. For example, the user 101 can identify a document stored on the client device 104 or elsewhere (e.g., stored in the data store 150) and select the DPP 122 to index the document. In response to receiving the DPP selection, the RAG API 108 can execute one or more pre-programmed calls to upload the identified document to the RAG server 120 for processing using the document segmentation engine 124 and the embedding model 126. The processed (e.g., segmented) document 152 can be stored in the data store 150 along with the embedding 154 associated with the segmented document 152. Similarly, a user query submitted by user 101, for example, via UI 106 of client device 104, can be forwarded to RAG server 120 for processing. RAG server 120 can apply embedding model 126 to the user query to generate a query embedding, and can further apply search engine 128 to identify stored embeddings 154 that are most similar to the query embedding. Hint generator 132 can then enhance the user query with fragments of stored documents 152, for example, using indexing 156 of embeddings 154 corresponding to the portion of document 152 represented by embedding 154. In some embodiments, RAG API 108 can be downloaded to client device 104 and can execute pre-programmed calls to RAG server 120 without using cloud server 112.

[0035] In some embodiments, any, some, or all operations of the RAG server 120 can be implemented on the cloud server 112 (or using the cloud server 112) using a container received from the RAG server 120. More specifically, various codes that implement any, some, or all of the RAG API 108, DPP 122, document segmentation engine 124, embedding model 126, search engine 128, and / or prompt generator 132 can be packaged into an image container (e.g., a lightweight executable software package that can be provided to the cloud server 112). The image container can further include various system tools, libraries, and settings. A container execution engine 114 (e.g., a Docker engine or similar container execution engine) running on the cloud server 112 can receive the container image and instantiate a RAG container based on the container image. The instantiated container can run on the cloud server 112 in an isolated environment.

[0036] Documents 152, embeddings 154, indexing 156, and / or other data stored in data store 150 may be accessed by RAG server 120, cloud server 112, client device 104, and / or Figure 1 104 . The data store 150 may include persistent storage and may be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tape or hard drives, network attached storage (NAS), storage area network (SAN), etc. Although depicted as being separate from the RAG server 120, cloud server 112, and / or client device 104, in at least some embodiments, the data store 150 may be part of the RAG server 120, cloud server 112, and / or client device 104. In at least some embodiments, the data store 150 may be a network attached file server, while in other embodiments, the data store 150 may be some other type of persistent storage, such as an object-oriented database, a relational database, etc., which may be hosted by the RAG server 120, cloud server 112, and / or client device 104 or by one or more different machines coupled to the RAG server 120, cloud server 112, and / or client device 104.

[0037] Figure 2 An example computing device 200 supporting efficient RAG for AI model processing is shown in accordance with at least one embodiment. In at least one embodiment, the computing device 200 can be part of the RAG server 120, part of the cloud server 112, and / or part of the client device 104. In at least one embodiment, one or more RAG APIs 108 can be run on the computing device 200. The RAG API 108 can be implemented, for example, by operating the document segmentation engine 124, the embedding model 126, the search engine 128, the hint generator 132, and / or at the client device 104. Figure 2 Other components not explicitly depicted in Figure 3 and / or Figure 4 public) to facilitate processing of input query / document 202.

[0038] Operations and calls of the RAG API 108 and various modules operating in conjunction with the RAG API 108 and / or other software / firmware instantiated on the computing device 200 may be performed using one or more GPUs 110, one or more CPUs 130, one or more parallel processing units (PPUs) or accelerators (e.g., deep learning accelerators), data processing units (DPUs), and / or the like. In at least one embodiment, the GPU 110 includes multiple cores 211, each capable of executing multiple threads 212. Each core can run multiple threads 212 concurrently (e.g., in parallel). In at least one embodiment, threads 212 can access registers 213. Registers 213 can be thread-specific registers, with access to registers limited to the corresponding thread. In addition, shared registers 214 can be accessible by one or more (e.g., all) threads of a core. In at least one embodiment, each core 211 can include a scheduler 215 for distributing computing tasks and processes among the different threads 212 of the core 211. Dispatch unit 216 may implement the scheduled task on the appropriate thread using the correct private registers 213 and shared registers 214. Computing device 200 may include input / output component 217 to facilitate exchanging information with one or more users or developers.

[0039] In at least one embodiment, the GPU 110 may have a (high-speed) cache 218, to which multiple cores 211 may share access. In addition, the computing device 200 may include a GPU memory 219, in which the GPU 110 may store intermediate and / or final results (outputs) of various computations performed by the GPU 110. After completing a particular task, the GPU 110 (or the CPU 130) may move the output to the (main) memory 105. In at least one embodiment, the CPU 130 may perform processes involving serial computational tasks, while the GPU 110 may perform tasks suitable for parallel processing (e.g., multiplying the input of a neural node by a weight and adding it to a bias).

[0040] The systems and methods described herein may be used for a variety of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, generative AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0041] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems for generating or presenting at least one of augmented reality content, virtual reality content, and mixed reality content, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing generative AI operations, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems implementing one or more language models (e.g., large language models (LLMs) that can process text, speech, images, and / or other data types to generate output in one or more formats), systems implementing one or more visual language models (VLMs), systems implemented at least in part using cloud computing resources, and / or other types of systems.

[0042] Although for the sake of brevity and clarity, the description throughout this disclosure often refers to text input, text documents, text queries, etc., it should be understood that the disclosed technology is also applicable to indexing and retrieval of images (including structured images), data (including structured data), and any other input / data, including multimodal data containing a combination of text and images, text and audio, etc. As used herein, "document" should be understood to be any digital data that can be segmented into semantically meaningful parts, such as a paragraph of text, a portion of an image, a table (or equivalent, such as a chart) of a dataset, one or more frames of a video file, a portion of an audio file, and / or the like.

[0043] Figure 3 An example data flow 300 is shown for the document indexing phase of efficient RAG for AI model processing, according to at least one embodiment. Figure 3 The operations shown in may be performed by the client device 104, the cloud server 112, the RAG server 120, and / or other suitable computing devices. Figure 3 The operations shown in can be performed to process, embed and index any suitable document 302, including text documents, images, audio files, video files, data files, etc. (Although for simplicity, the Figure 3Reference is made throughout the description of a single document 302, but any batches of multiple documents 302 may be processed similarly (e.g., sequentially or in parallel). In some implementations, the documents 302 may be associated with any general or specific area of ​​knowledge, such as medical diagnosis, computing technology, mathematics, computer games, art history, etc. The documents 302 may be provided using any suitable means, including by a user (e.g., Figure 1 The document 302 may be uploaded by or under the control of a user 101 in the document 302, automated data collection, database mining, and / or the like. In some cases, for example, when the document 302 is available as an image, the document 302 may first undergo optical character recognition (OCR). Prior to (or as part of) OCR, the document 302 may be denoised, filtered, sharpened, or enhanced using any suitable pre-processing tools or techniques.

[0044] The processing of the document 302 may be facilitated by an indexing API 310, which may be Figure 1 and Figure 2 The indexing API 310 may be part of the RAG API 108. The indexing API 310 may make available to the user who uploaded the document 302, or any other entity or software, a DPP selector 320 that enables selection from multiple DPPs 122. Different DPPs 122 may have different predefined settings (configurations), including the size of individual segments of the document 302. For example, in the case of a text document 302, the size S may be specified as a number of characters, words, sentences, lines, paragraphs, pages, etc., such as 1000 characters, 200 words, one page, etc. The DPP settings may further include the amount of segment overlap O between adjacent segments. This overlap O may also be specified as a number of words, sentences, etc. (e.g., 20 words, 100 characters, and / or the like), a percentage (e.g., 10%, 20%, and / or the like). The DPP settings may further include the number K of top search matches and / or the like to be returned. In some embodiments, the DPP 122 settings may also include a specific embedding model to be used to embed the segments. In some embodiments, the settings for the DPP 122 can specify the storage location of embedded and document fragments, the mapping (indexing) scheme of embeds to fragments, and / or other relevant information. In some embodiments, the DPP selector 320 can specify a default DPP 122 that can be used if the user lacks preferences, knowledge, or experience to make an informed DPP selection. The default DPP 122 can depend on the document type and can be different for text documents than for image documents.

[0045] An "embedding" ("feature vector" or simply vector) should be understood as any digital representation of a data unit (e.g., an alphanumeric string, an image or part of an image, a dataset or subset, etc.) in a space of dimension N (the number of components of the individual vectors) (the "embedding space") that can be set up as part of training a model (an "embedding model") for representing ("embedding") data via points (vectors) in the N-dimensional embedding space. The components of the embedding (vector) can have integer or floating point values. During training, the embedding model can learn to associate similar data units with similar embeddings (vectors) corresponding to points that are close together in the embedding space, and further learn to associate dissimilar data units with points that are far apart in the embedding space.

[0046] The user-selected (or default) DPP 122 can trigger the operation of the document segmentation engine 124, which segments the document 302 into a plurality of document segments 330 according to the settings of the selected DPP. These segments can be stored in a segment store 352, which can be implemented as part of the RAG store 350. The document segmentation engine 124 can also perform segment indexing 340 by assigning unique indicators to each segment. For example, individual segments can be indexed by a document identifier (ID), a number indicating the segment's location in the document, a storage location in the segment store 352, and / or any other relevant information. The embedding model 126 can process the document segments 330 to generate embeddings 360. The embeddings 360 can be vectors in an N-dimensional embedding space. The dimension N can be fixed as part of the architecture of the embedding model 126 or flexibly specified as part of the selected DPP 122. The generated embeddings can be stored in an embedding store 354, which can be implemented as part of the same RAG store 350 that hosts the fragment store 352, or as part of a different store. The stored embeddings 360 can also be fragment-indexed to uniquely identify each embedding 360 and map the stored embeddings 360 to stored document fragments 330, e.g., so that a given document fragment 330 can be used to identify a corresponding embedding 360 that encodes the same portion of a document 302. In some implementations, the RAG store 350 can store thousands or millions (or more) of documents 302 related to a single or multiple knowledge domains.

[0047] Figure 4 An example data flow 400 is shown for the query processing phase of an efficient RAG for AI model processing, according to at least one embodiment. Figure 4 The operations shown in can be performed by Figure 3 In some embodiments, query processing may be performed by the same RAG client infrastructure 102 (see Figure 1 ) is executed by different servers, for example, different cloud servers 112 or different client devices 104 that can access the stored document fragments and embedded RAG storage 350. Figure 4 The operations shown in can be performed to facilitate generating efficient queries to the LM or to enhance the input of some other AI model. For example, similar operations can be performed to generate an image based on a user's instruction (e.g., a description of an image), where one or more previously stored images will be recognized and used to enhance the user's instruction.

[0048] As shown, a document query 402, such as a text query, an audio query, and / or the like, may be received via a query API 410, which may be Figure 1 and Figure 2 In some embodiments, the query API 410 and ( Figure 3 The indexing API 310 can be implemented as a single API that facilitates both document retrieval and query processing. In some embodiments, the query API 410 and the indexing API 310 can be implemented as separate APIs. In some embodiments, the query API 410 can access multiple DPPs 122. The user (or process) generating the document query 402 can also select the DPP 122 to be used for query processing. In some embodiments, the user (or process) can select the same DPP 122 that was selected for processing (segmenting and embedding) the documents during the indexing phase. For example, a large document query 402 can be segmented into portions of the same size S and overlap O as those used to process documents associated with the same knowledge domain (during the indexing phase). In some embodiments, the user (or process) can select a different DPP 122 than the one used to process the documents during the indexing phase. In some embodiments, the selected DPP 122 can specify the number K of the best (most relevant) documents or document fragments to be returned.

[0049] The document query 402 (partitioned appropriately, if dictated by the size of the query) may be converted into a query embedding 420 (or, for large queries, multiple query embeddings 420). The generated query embedding 420 may be received by an embedding search module 440 of the search engine 128. The embedding search module 440 may further receive (e.g., from the embedding storage 354) the stored embeddings 360 (which may be combined with Figure 3The embedding search module 440 can identify the K stored embeddings 360 that are most similar to the query embedding 420 (and, therefore, most relevant to the document query 402). In an example embodiment, a cosine similarity function can be used as part of the embedding search 440, which calculates the K similarity between the query embedding 420 (QE) and each stored embedding 360 (SE j , where j represents any suitable ID of the stored embedding):

[0050]

[0051] Having identified the K best matches (e.g., the K stored embeddings 360 with the highest Similarity scores relative to the query embedding 420), the embedding search 460 can use the fragment indexing 340 to identify the corresponding K document fragments, referred to herein as the first fragment set 442. In some embodiments, the embedding search module 440 can perform a search of the relevant stored embeddings 360 in the embedding storage 354. In some embodiments, such an embedding search can first be performed sparsely on a set of stored embeddings, e.g., where one or more stored embeddings 360 are sampled (e.g., randomly or according to any suitable pattern, such as selecting one embedding every n embeddings) and compared to the query embedding 420. Then, a dense search of documents can be performed associated with those stored embeddings 360 with the highest similarity scores (e.g., the text fragments within a neighborhood of the most relevant hits).

[0052] In some embodiments, an additional document search 430 may be performed on document fragments 330 stored in their original format (e.g., textual form) (e.g., stored in fragment store 352). Document search 430 may be a sparse search. In some embodiments, document search 430 may be Elasticsearch or some other text search. Document search 430 may return a second set of fragments 432. The first set of fragments 442 and the second set of fragments 432 may partially (or completely) overlap. A joiner 450 may eliminate duplicate fragments and generate search results 460, e.g., a list of document fragments most similar to document query 402. In some embodiments, joiner 450 may limit search results 460 to a plurality of top results, e.g., the K most similar fragments. In some embodiments, some of search results 460 may be ranked, e.g., in order of decreasing similarity score for the first set of fragments 442 or in order of word matches for the second set of fragments 432. In some embodiments, fragments in search results 460 obtained through different searches may not be ranked. In some implementations, all segments in search results 460 can be ranked using some common ranking scheme. For example, connector 450 can retrieve embeddings for segments obtained using document search 430 but not captured by embedding search 440 and calculate corresponding similarity scores to rank all search results 460 using a common scheme based on the similarity scores.

[0053] In some embodiments, the search results 460 may be accepted as the final search results 480. In some embodiments, the search results 460 may be provided to a re-ranking model 470, which re-ranks the search results 460 based on the document query 402. In some embodiments, the re-ranking model 470 may be the same as the LM 162. In other embodiments, the re-ranking model 470 may be a separate model, such as a lightweight language model. The final search results 480 (originally ranked, re-ranked, or unranked) may be combined with the document query 402 to form a RAG enhanced prompt 490. The prompt 490 may then be provided for processing by the LM 162, which may return an appropriate response.

[0054] Figure 5 and Figure 6Example methods 500 and 600 for efficient API-facilitated retrieval enhancement generation for AI model processing are shown. Methods 500 and 600 can be used in the context of deployment and / or use of AI, including (but not limited to) language models, visual language models, computer vision models, text-to-speech models, speech-to-text models, and / or other AI models, where processing of an AI model's input (e.g., text, images, speech, audio, video, digital assets, CAD, and / or any other data) can be improved by enhancing that input with appropriate contextual information (e.g., instances of historical data, background data, sample data, and / or the like). In at least one embodiment, methods 500 and / or 600 can be used Figure 1 Client device 104, cloud server 112, RAG server 120, Figure 2 The method 500 and / or the method 600 may be performed by one or more processing units of the computing device 200 and / or some other computing device or combination of computing devices. The one or more processing units (e.g., CPU, GPU, accelerator, PPU, DPU, etc.) executing the method 500 and / or the method 600 may include (or communicate with) one or more memory devices.

[0055] In at least one embodiment, method 500 and / or method 600 may be performed by the same computing device. In at least one embodiment, method 500 and / or method 600 may be performed by different computing devices. In at least one embodiment, the processing unit executing method 500 and / or 600 may execute instructions stored on a non-transitory computer-readable storage medium. In at least one embodiment, method 500 and / or 600 may be executed using multiple processing threads (e.g., CPU threads and / or GPU threads), where each thread executes one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing any of methods 500 and / or 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing any of methods 500 and / or 600 may execute asynchronously with respect to each other. The individual operations of any of methods 500 and / or 600 may be performed in a manner similar to Figure 5 and Figure 6 Some operations of any of methods 500 and / or 600 may be performed concurrently with other operations. In at least one embodiment, Figure 5 and Figure 6 One or more of the operations shown may not always be performed.

[0056] Figure 5is a flow chart of an example method 500 for an indexing phase of an efficient API-facilitated RAG for AI model processing according to at least one embodiment. At block 510, the method 500 may include providing a plurality of document processing pipelines (DPPs) to a user interface using a processing device that executes an application programming interface (API). In some implementations, the user interface may be located on a device different from the device executing the method 500. For example, referring to Figure 1 , the method 500 may be performed by the cloud server 112, and the user interface may be on the client device 104 in remote communication with the cloud server 112 over a network (e.g., the network 140). At block 520, the method 500 may continue by receiving a selection (e.g., a user selection) of a DPP from the plurality of DPPs from the user interface via the API.

[0057] At block 530, method 500 may include segmenting the input document into a plurality of segments according to predetermined settings of the selected DPP. In some embodiments, the input document may be received from a remote device (e.g., client device 104, a processing device associated with data store 150, and / or the like). The input document may include text, one or more images, a table, a data set, an audio file, the like, or any combination thereof. In some embodiments, the predetermined settings may include the size of individual segments in the plurality of segments, the amount of overlap between adjacent segments in the plurality of segments, the selection of an embedding model, and / or any combination thereof.

[0058] At block 540, method 500 may continue by causing an embedding model (e.g., an embedding model identified by the DPP selection or a default embedding model) to process the plurality of segments to generate a plurality of embeddings. In some embodiments, the embedding model may be deployed by the same computing device that executes method 500. In some embodiments, the embedding model may be deployed by a different (e.g., remote) computing device (e.g., RAG server 120 and / or the like).

[0059] At block 550, method 500 may include causing the plurality of embeddings to be stored in a first data store. The first data store may be (or include) a memory device of a computing device executing method 500. In some embodiments, the first data store may be a data store (e.g., data store 150) remote from the computing device executing method 500. In some embodiments, the operations of block 570 may include storing indexed data mapping the plurality of embeddings to the plurality of segments. In some embodiments, as indicated by dashed block 560, method 500 may continue with causing the plurality of segments to be stored in at least the first data store or the second data store.

[0060] In some embodiments, the operations of method 500 may be performed using a container infrastructure, as shown by dashed blocks 502 and 504. More specifically, at block 502, method 500 may include receiving a container image from a remote computing device. The container image may include an API that facilitates execution of method 500. In some embodiments, the container image may further include a segmentation engine, an embedding model, and / or other tools and modules that segment an input document into multiple segments. At block 504, method 500 may include executing the API in a container instantiated using the container image. The segmentation engine, embedding model, and / or other tools and modules received with the container image may also be executed in the instantiated container.

[0061] Figure 6 is a flow diagram of an example method 600 for a query processing phase of a RAG facilitated by an efficient API for AI model processing according to at least one embodiment. At block 610, method 500 may include receiving a query. The query may include text, one or more images, a table, a data set, an audio file, and / or any combination thereof. The query may be received using a processing device that executes the API. In some embodiments, the API that facilitates the query processing phase of method 600 may be the same as the API that facilitates the indexing phase of method 500, e.g., a single API downloadable from a RAG server 120 that supports both phases of RAG-facilitated AI operations. In some embodiments, the API that facilitates the query processing phase of method 600 may be different from the API that facilitates the indexing phase of method 500 (e.g., Figure 4 In some implementations, in addition to receiving a query, the operations of block 610 may also include selecting a DPP from a plurality of DPPs provided by the API. The selected DPP may include a maximum number of document fragments to be identified in conjunction with the query. At block 620, method 600 may include causing an embedding model to process the query to generate one or more query embeddings.

[0062] At block 630, method 600 may continue to calculate a plurality of similarity scores that characterize the similarity of the one or more query embeddings to a plurality of (stored) embeddings associated with the one or more stored documents. At block 640, method 600 may continue to use the plurality of similarity scores to select one or more segments of the one or more stored documents (e.g., Figure 4 The one or more selected segments may be from a single stored document or from multiple stored documents. The number of segments selected from a given document need not be limited. In some embodiments, selecting one or more segments may include accessing stored indexed data (e.g., Figure 1156 ), which maps multiple (stored) embeddings (e.g., embedding 154 ) to one or more stored documents (e.g., document 152 ).

[0063] In some embodiments, selecting one or more segments may include: Figure 6 For example, at block 642, the operations of method 600 may include: using the plurality of similarity scores to identify one or more embeddings from a plurality of (stored) embeddings, the one or more identified embeddings corresponding to one or more segments associated with the query. At block 644, the operations of method 600 may include: using the plurality of similarity scores to rank the one or more segments according to their relevance to the query.

[0064] In some implementations, at block 646, operations of method 600 may include performing a document search (e.g., a text search) to identify one or more additional segments of one or more stored documents (e.g., Figure 4 ), the one or more additional segments having textual relevance to the query. At block 648, the operations of method 600 may further include ranking the segment set according to relevance to the query using a ranking model. In some implementations, the segment set being ranked may include one or more segments identified using embedded search (e.g., Figure 4 ) and one or more additional segments identified using the document search (e.g., Figure 4 The second set of fragments 442 in .

[0065] At block 650, method 600 may include generating a LM hint and processing the LM hint using LM to obtain a response to the query. The LM hint may be based on at least the query and one or more selected segments, e.g., segments identified using embedded search and / or additional segments identified using document search. In some embodiments, the LM hint may be generated using a ranked set of segments, e.g., by including segments and corresponding rankings. The segments included in the LM hint may include segments identified using embedded search and, in some embodiments, may also include additional segments identified using document search.

[0066] The systems and methods described herein may be used for a variety of purposes, such as, but not limited to, performing one or more operations related to machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0067] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., in-vehicle infotainment systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least in part using cloud computing resources, and / or other types of systems.

[0068] Reasoning and training logic

[0069] Figure 7A Inference and / or training logic 715 is shown for performing inference and / or training operations associated with one or more embodiments.

[0070] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequence, wherein weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs) or simply circuits). In at least one embodiment, code (such as graph code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores input / output data during training and / or inference using aspects of one or more embodiments and / or weight parameters during forward propagation of weight parameters for each layer of a neural network trained or used in conjunction with one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 may be included within other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0071] In at least one embodiment, any portion of code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 701 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or code and / or data storage 701 is internal or external to a processor, for example, or composed of DRAM, SRAM, flash memory, or some other type of storage, may depend on the available storage space on or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training of the neural network, or some combination of these factors.

[0072] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 705 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, during training and / or inference using aspects of one or more embodiments, the code and / or data storage 705 stores weight parameters and / or input / output data for each layer of the neural network trained or used in conjunction with one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software to control the timing and / or sequence in which weight and / or other parameter information is loaded to configure logic, which includes integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).

[0073] In at least one embodiment, code (such as graph code) causes weights or other parameter information to be loaded into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 705 is internal or external to the processor, for example, whether it is composed of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.

[0074] In at least one embodiment, code and / or data store 701 and code and / or data store 705 may be separate storage structures. In at least one embodiment, code and / or data store 701 and code and / or data store 705 may be a combined storage structure. In at least one embodiment, code and / or data store 701 and code and / or data store 705 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data store 701 and code and / or data store 705 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0075] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from a layer or neuron within a neural network) stored in activation storage 720, which are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 710 to generate activations stored in activation storage 720, wherein weight values ​​stored in code and / or data storage 705 and / or data storage 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 705 or code and / or data storage 701 or in another on-chip or off-chip storage.

[0076] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 710 may be external to the processor or other hardware logic devices or circuits that use them (e.g., coprocessors). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor or otherwise included in a group of ALUs accessible by the execution units of a processor, which may be within the same processor or distributed between different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 720 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0077] In at least one embodiment, activation storage 720 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 720 can be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, activation storage 720 can be internal or external to the processor, for example, or comprise DRAM, SRAM, flash memory, or some other storage type, depending on the available on-chip or off-chip storage, the latency requirements for performing training and / or inference functions, the batch size of data used in inferring and / or training neural networks, or some combination of these factors.

[0078] In at least one embodiment, Figure 7A The inference and / or training logic 715 shown in FIG may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7AThe illustrated inference and / or training logic 715 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0079] Figure 7B Inference and / or training logic 715 is shown in accordance with at least one embodiment. In at least one embodiment, inference and / or training logic 715 may include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown in FIG can be used in conjunction with an application specific integrated circuit (ASIC), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 7B In at least one embodiment shown in FIG, code and / or data storage 701 and code and / or data storage 705 are each associated with dedicated computing resources, such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, computing hardware 702 and computing hardware 706 each include one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) solely on the information stored in code and / or data storage 701 and code and / or data storage 705, respectively, with the results of the functions being stored in activation storage 720.

[0080] In at least one embodiment, each of the code and / or data stores 701 and 705 and the corresponding computing hardware 702 and 706 corresponds to a different layer of a neural network, such that activations from one storage / computation pair 701 / 702 of the code and / or data store 701 and computing hardware 702 are provided as input to the next storage / computation pair 705 / 706 of the code and / or data store 705 and computing hardware 706, reflecting the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) can be included in the inference and / or training logic 715 after or in parallel with the storage / computation pairs 701 / 702 and 705 / 706.

[0081] Neural network training and deployment

[0082] Figure 8 The training and deployment of a deep neural network according to at least one embodiment is shown. In at least one embodiment, an untrained neural network 806 is trained using a training dataset 802. In at least one embodiment, the training framework 804 is the PyTorch framework, while in other embodiments, the training framework 804 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 804 trains the untrained neural network 806 and enables it to be trained using the processing resources described herein to generate a trained neural network 808. In at least one embodiment, the weights can be randomly selected or pre-trained using a deep belief network. In at least one embodiment, the training can be performed in a supervised, partially supervised, or unsupervised manner.

[0083] In at least one embodiment, untrained neural network 806 is trained using supervised learning, where training dataset 802 includes inputs paired with expected outputs for the inputs, or where training dataset 802 includes inputs with known outputs and neural network 806 is manually graded for outputs. In at least one embodiment, untrained neural network 806 is trained in a supervised manner, processing inputs from training dataset 802 and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through untrained neural network 806. In at least one embodiment, training framework 804 adjusts the weights that control untrained neural network 806. In at least one embodiment, training framework 804 includes tools for monitoring the degree to which untrained neural network 806 converges toward a model (e.g., trained neural network 808) that is suitable for generating correct answers (e.g., results 814) based on input data (e.g., new dataset 812). In at least one embodiment, training framework 804 iteratively trains untrained neural network 806 while adjusting weights to improve the output of untrained neural network 806 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 804 trains the untrained neural network 806 until the untrained neural network 806 reaches a desired accuracy. In at least one embodiment, the trained neural network 808 can then be deployed to implement any number of machine learning operations.

[0084] In at least one embodiment, unsupervised learning is used to train an untrained neural network 806, which attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 802 will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 806 can learn the groupings within the training dataset 802 and can determine how the individual inputs relate to the untrained dataset 802. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in the trained neural network 808, which can perform operations useful for reducing the dimensionality of the new dataset 812. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for the identification of data points in the new dataset 812 that deviate from the normal pattern of the new dataset 812.

[0085] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mixture of labeled and unlabeled data is included in the training dataset 802. In at least one embodiment, the training framework 804 can be used to perform incremental learning, for example, through a transfer learning technique. In at least one embodiment, incremental learning enables the trained neural network 808 to adapt to the new dataset 812 without forgetting the knowledge that was infused into the trained neural network 808 during the initial training.

[0086] Reference Figure 9 , Figure 9 is an example data flow diagram of a process 900 for generating and deploying a processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 900 can be deployed to perform game name recognition analysis and inference on user feedback data at one or more facilities 902, such as a data center.

[0087] In at least one embodiment, process 900 may be performed within training system 904 and / or deployment system 906. In at least one embodiment, training system 904 may be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use with deployment system 906. In at least one embodiment, deployment system 906 may be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements at facility 902. In at least one embodiment, deployment system 906 may provide a pipeline platform for selecting, customizing, and implementing virtual instruments for use with computing devices at facility 902. In at least one embodiment, a virtual instrument may include a software-defined application for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in the pipeline may use or call services (e.g., reasoning, visualization, computation, AI, etc.) of deployment system 906 during application execution.

[0088] In at least one embodiment, some applications used in high-level processing and reasoning pipelines may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained at the facility 902 using feedback data 908 (e.g., imaging data) stored at the facility 902 or feedback data 908 from another or more facilities, or a combination thereof. In at least one embodiment, the training system 904 may be used to provide applications, services, and / or other resources to generate a working, deployable machine learning model for the deployment system 906.

[0089] In at least one embodiment, the model registry 924 can be backed by an object store that can support version control and object metadata. In at least one embodiment, the model registry 924 can be accessed from within the cloud platform through, for example, a cloud store (e.g., Figure 10 The object store is accessed through an application programming interface (API) compatible with the cloud 1026. In at least one embodiment, machine learning models within the model registry 924 can be uploaded, listed, modified, or deleted by developers or partners of the system interacting with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate a model with an application so that the model can be executed as part of the execution of a containerized instantiation of the application.

[0090] In at least one embodiment, the training pipeline 1004 ( Figure 10 ) may include scenarios where the facility 902 is training their own machine learning model, or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, feedback data 908 may be received from various channels (such as forums, web forms, etc.). In at least one embodiment, once the feedback data 908 is received, AI-assisted annotations 910 may be used to help generate annotations corresponding to the feedback data 908 to be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotations 910 may include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that may be trained to generate annotations corresponding to certain types of feedback data 908 (e.g., from certain devices), and / or certain types of anomalies in the feedback data 908. In at least one embodiment, the AI-assisted annotations 910 may then be used directly, or may be adjusted or fine-tuned using annotation tools to generate ground truth data. In at least one embodiment, in some examples, labeled data 912 may be used as ground truth data for training the machine learning model. In at least one embodiment, the AI-assisted annotations 910, labeled data 912, or a combination thereof may be used as ground truth data for training the machine learning model (e.g., via Figure 9-10 In at least one embodiment, the trained machine learning model can be referred to as an output model 916 and can be used by the deployment system 906, as described herein.

[0091] In at least one embodiment, the training pipeline 1004 ( Figure 10) may include situations where facility 902 requires a machine learning model for performing one or more processing tasks for one or more applications in deployment system 906, but facility 902 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from model registry 924. In at least one embodiment, model registry 924 may include machine learning models that have been trained to perform a variety of different inference tasks on imaging data. In at least one embodiment, the machine learning models in model registry 924 may be trained on imaging data from a different facility (e.g., a remotely located facility) than facility 902. In at least one embodiment, the machine learning model may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a specific location (which may be in the form of feedback data 908), the training may be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts its transfer off-site (e.g., to comply with HIPAA regulations, privacy regulations, etc.). In at least one embodiment, once a model is trained or partially trained at one location, the machine learning model can be added to the model registry 924. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in the model registry 924. In at least one embodiment, the machine learning model can then be selected from the model registry 924 (and referred to as the output model 916) and can be used in the deployment system 906 to perform one or more processing tasks for one or more applications of the deployment system.

[0092] In at least one embodiment, the training pipeline 1004 ( Figure 10) can be used in scenarios including a facility 902 that requires a machine learning model for performing one or more processing tasks for one or more applications in a deployment system 906, but the facility 902 may not currently possess such a machine learning model (or may not possess an optimized, efficient, or effective model for this purpose). In at least one embodiment, a machine learning model selected from the model registry 924 may not be fine-tuned or optimized for the feedback data 908 generated at the facility 902 due to population differences, genetic variation, robustness of the training data used to train the machine learning model, diversity of training data anomalies, and / or other issues with the training data. In at least one embodiment, AI-assisted annotation 910 can be used to help generate annotations corresponding to the feedback data 908 to serve as ground truth data for retraining or updating the machine learning model. In at least one embodiment, labeled data 912 can be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 914. In at least one embodiment, model training 914 (e.g., AI-assisted annotation 910, labeled data 912, or a combination thereof) can be used as ground truth data for retraining or updating the machine learning model.

[0093] In at least one embodiment, deployment system 906 may include software 918, services 920, hardware 922, and / or other components, features, and functionality. In at least one embodiment, deployment system 906 may include a software "stack" such that software 918 may be built on top of services 920 and may use services 920 to perform some or all processing tasks, and services 920 and software 918 may be built on top of hardware 922 and use hardware 922 to perform processing, storage, and / or other computing tasks of deployment system 906.

[0094] In at least one embodiment, the software 918 may include any number of different containers, each of which may execute an instantiation of an application. In at least one embodiment, each application may execute one or more processing tasks (e.g., reasoning, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and reasoning pipeline. In at least one embodiment, for each type of computing device, there may be any number of containers that may execute data processing tasks on the feedback data 908 (or other data types, such as those described herein). In at least one embodiment, in addition to receiving and configuring imaging data for use by each container and / or for use by the facility 902 after processing through the pipeline, a high-level processing and reasoning pipeline may be defined based on the selection of different containers desired or required to process the feedback data 908 (e.g., to convert the output back into a usable data type for storage and display at the facility 902). In at least one embodiment, the combination of containers within the software 918 (e.g., that comprise the pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize services 920 and hardware 922 to execute some or all of the processing tasks of the application instantiated in the container.

[0095] In at least one embodiment, data can be pre-processed as part of a data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing can be performed on the output of one or more inference tasks or other processing tasks in the pipeline to prepare output data for the next application and / or prepare the output data for transmission and / or use by a user (e.g., as a response to an inference request). In at least one embodiment, the inference task can be performed by one or more machine learning models, such as trained or deployed neural networks, which can include the output model 916 of the training system 904.

[0096] In at least one embodiment, the tasks of a data processing pipeline can be encapsulated in one or more containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment that can reference a machine learning model. In at least one embodiment, the containers or applications can be published to a private (e.g., limited access) area of ​​a container registry (described in more detail herein), and the trained or deployed models can be stored in the model registry 924 and associated with one or more applications. In at least one embodiment, an image of the application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, the image can be used to generate an instantiated container for the application for use by the user's system.

[0097] In at least one embodiment, a developer can develop, publish, and store applications (e.g., as containers) for performing processing and / or reasoning on provided data. In at least one embodiment, development, publishing, and / or storage can be performed using a software development kit (SDK) associated with the system (e.g., to ensure that the developed applications and / or containers conform to or are compatible with the system). In at least one embodiment, the developed applications can be tested locally (e.g., at a first facility, on data from the first facility) using the SDK, which serves as a system (e.g., Figure 10 The architecture 1000 in FIG. 1000 may support at least some of the services 920. In at least one embodiment, once validated by the architecture 1000 (e.g., for accuracy, etc.), the application is made available in the container registry for selection and / or implementation by a user (e.g., a hospital, clinic, laboratory, healthcare provider, etc.) to perform one or more processing tasks on data at the user's facility (e.g., a second facility).

[0098] In at least one embodiment, the developer can then share the application or container over a network for use by a system (e.g., Figure 10 924). In at least one embodiment, completed and validated applications or containers can be stored in a container registry, and associated machine learning models can be stored in a model registry 924. In at least one embodiment, a requesting entity (providing an inference or image processing request) can browse the container registry and / or model registry 924 for applications, containers, datasets, machine learning models, etc., select the desired combination of elements to include in a data processing pipeline, and submit a processing request. In at least one embodiment, the request can include the input data necessary to execute the request and / or can include a selection of the application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request can then be passed to one or more components of the deployment system 906 (e.g., a cloud) to perform processing in the data processing pipeline. In at least one embodiment, the processing performed by the deployment system 906 can include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 924. In at least one embodiment, once results are generated by the pipeline, the results can be returned to the user for reference (e.g., for viewing in a viewing application suite executed locally, on a local workstation, or on a terminal).

[0099] In at least one embodiment, to assist in processing or executing applications or containers in the pipeline, services 920 may be utilized. In at least one embodiment, services 920 may include computing services, collaborative content creation services, simulation services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, services 920 may provide functionality that is common to one or more applications in software 918, and thus may abstract functionality into services that can be called or utilized by applications. In at least one embodiment, the functionality provided by services 920 may run dynamically and more efficiently, while also enabling faster processing of data by allowing applications to process data in parallel (e.g., using Figure 10 In at least one embodiment, rather than requiring each application that shares the same functionality provided by service 920 to have a corresponding instance of service 920, service 920 can be shared between and among various applications. In at least one embodiment, as non-limiting examples, the service may include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included that can provide machine learning model training and / or retraining capabilities.

[0100] In at least one embodiment, where the service 920 includes an AI service (e.g., an inference service), as part of the execution of an application, one or more machine learning models associated with an application for anomaly detection (e.g., tumor, growth abnormality, scarring, etc.) can be executed by calling (e.g., as an API call) the inference service (e.g., an inference server) to execute one or more machine learning models or processing thereof. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can call the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, the software 918 implementing the high-level processing and inference pipeline can be pipelined in that each application can call the same inference service to perform one or more inference tasks.

[0101] In at least one embodiment, hardware 922 may include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX TMIn at least one embodiment, different types of hardware 922 can be used to provide efficient, purpose-built support for the software 918 and services 920 in the deployment system 906. In at least one embodiment, GPU processing can be used to perform local processing (e.g., at the facility 902) within the AI / deep learning system, in the cloud system, and / or in other processing components of the deployment system 906 to improve the efficiency, accuracy, and effectiveness of game name recognition.

[0102] In at least one embodiment, software 918 and / or services 920 may be optimized for GPU processing, as non-limiting examples, with respect to deep learning, machine learning, and / or high performance computing, simulation, and visual computing. In at least one embodiment, at least some of the computing environments of deployment system 906 and / or training system 904 may be run on GPUs with GPU-optimized software (e.g., NVIDIA DGX TM In at least one embodiment, the cloud platform may include GPU-optimized execution of deep learning tasks, GPU processing of machine learning tasks, or other computing tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as in NVIDIA's DGX TM system) as a hardware abstraction and scaling platform to execute cloud platforms (e.g., NVIDIA's NGC TM In at least one embodiment, the cloud platform can integrate an application container cluster system or coordination system (e.g., Kubernetes) on multiple GPUs to achieve seamless scaling and load balancing.

[0103] Figure 10 is a system diagram of an example architecture 1000 for generating and deploying a deployment pipeline according to at least one embodiment. In at least one embodiment, the architecture 1000 can be used to implement Figure 9 The process 900 and / or other processes of the present invention include high-level processing and reasoning pipelines. In at least one embodiment, the architecture 1000 may include a training system 904 and a deployment system 906. In at least one embodiment, the training system 904 and the deployment system 906 may be implemented using software 918, services 920, and / or hardware 922, as described herein.

[0104] In at least one embodiment, architecture 1000 (e.g., training system 904 and / or deployment system 906) can be implemented in a cloud computing environment (e.g., using cloud 1026). In at least one embodiment, architecture 1000 can be implemented locally (with respect to a facility) or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1026 can be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocols can include a network token, which can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, the APIs of the virtual instrument (described herein) or other instances of architecture 1000 can be restricted to a set of public Internet Service Providers (ISPs) that have been vetted or authorized for interaction.

[0105] In at least one embodiment, the various components of the architecture 1000 can communicate with each other and among themselves using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communications between facilities and components of the architecture 1000 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be transmitted via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0106] In at least one embodiment, similar to the present disclosure regarding Figure 9 As described, the training system 904 can execute a training pipeline 1004. In at least one embodiment, where the deployment system 906 will use one or more machine learning models in a deployment pipeline 1010, the training pipeline 1004 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1006 (e.g., without retraining or updating). In at least one embodiment, as a result of the training pipeline 1004, an output model 916 can be generated. In at least one embodiment, the training pipeline 1004 can include any number of processing steps, AI-assisted annotation 910, labeling of feedback data 908 or annotation to generate labeled data 912, selecting a model from a model registry, model training 914, training, retraining, or updating a model, and / or other processing steps. In at least one embodiment, different training pipelines 1004 can be used for different machine learning models used by the deployment system 906. In at least one embodiment, similar to the description regarding Figure 9 The training pipeline 1004 of the first example described can be used for a first machine learning model, similar to the one described with respect to Figure 9The second example training pipeline 1004 described can be used for a second machine learning model, similar to the one described with respect to Figure 9 The third example training pipeline 1004 is described as being usable for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 904 may be used, depending on the requirements of each respective machine learning model. In at least one embodiment, one or more machine learning models may already be trained and ready for deployment, so the training system 904 may not perform any processing on the machine learning model, and the machine learning model may be implemented by the deployment system 906.

[0107] In at least one embodiment, one or more output models 916 and / or pre-trained models 1006 may include any type of machine learning model, depending on the embodiment. In at least one embodiment and without limitation, the machine learning model used by the architecture 1000 may include linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., autoencoder, convolution, recursion, perceptron, long / short term memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machine, etc.), and / or other types of machine learning models.

[0108] In at least one embodiment, the training pipeline 1004 may include AI-assisted annotation. In at least one embodiment, the labeled data 912 may be generated by any number of techniques (e.g., traditional annotation). In at least one embodiment, in some examples, labels or other annotations may be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of application suitable for generating ground-truth annotations or labels, and / or may be hand-drawn. In at least one embodiment, the ground-truth data may be synthetically generated (e.g., generated from a computer model or rendering), realistically generated (e.g., designed and generated from real-world data), machine-generated (e.g., using feature analysis and learning to extract features from the data and then generate labels), manually annotated (e.g., a labeler or annotation expert defines the location of the labels), and / or a combination thereof. In at least one embodiment, for each instance of feedback data 908 (or other data type used by the machine learning model), there may be corresponding ground-truth data generated by the training system 904. In at least one embodiment, AI-assisted annotation may be performed as part of the deployment pipeline 1010; in addition to or in place of the AI-assisted annotation included in the training pipeline 1004. In at least one embodiment, architecture 1000 may include a multi-layer platform that may include a software layer (eg, software 918 ) of a diagnostic application (or other application type) that may perform one or more medical imaging and diagnostic functions.

[0109] In at least one embodiment, the software layer can be implemented as a secure, encrypted, and / or authenticated API through which an application or container can be invoked (e.g., called) from an external environment (e.g., facility 902). In at least one embodiment, the application can then call or execute one or more services 920 to perform computational, AI, or visualization tasks associated with the respective application, and the software 918 and / or services 920 can utilize hardware 922 to perform the processing tasks in an effective and efficient manner.

[0110] In at least one embodiment, the deployment system 906 can execute a deployment pipeline 1010. In at least one embodiment, the deployment pipeline 1010 can include any number of applications that can be sequential, non-sequential, or otherwise applied to feedback data (and / or other data types) - including AI-assisted annotations, as described above. In at least one embodiment, the deployment pipeline 1010 for an individual device, as described herein, can be referred to as a virtual instrument for the device. In at least one embodiment, there can be more than one deployment pipeline 1010 for a single device, depending on the information desired from the data generated by the device.

[0111] In at least one embodiment, applications that can be used to deploy pipeline 1010 can include any application that can be used to perform processing tasks on feedback data or other data from the device. In at least one embodiment, because various applications can share common image operations, in some embodiments, a data enhancement library (e.g., as one of the services 920) can be used to accelerate these operations. In at least one embodiment, in order to avoid the bottlenecks of traditional processing methods that rely on CPU processing, parallel computing platform 1030 can be used for GPU acceleration of these processing tasks.

[0112] In at least one embodiment, the deployment system 906 may include a user interface 1014 (e.g., a graphical user interface, a web interface, etc.) that may be used to select applications to be included in the deployment pipeline 1010, to arrange applications, to modify or change applications or their parameters or configuration, to use and interact with the deployment pipeline 1010 during setup and / or deployment, and / or to otherwise interact with the deployment system 906. In at least one embodiment, although not shown with respect to the training system 904, the UI 1014 (or a different user interface) may be used to select models for use in the deployment system 906, to select models for training or retraining in the training system 904, and / or to otherwise interact with the training system 904. In at least one embodiment, the training system 904 and the deployment system 906 may include DICOM adapters 1002A and 1002B.

[0113] In at least one embodiment, in addition to the application orchestration system 1028, a pipeline manager 1012 can be used to manage the interactions between applications or containers of the deployment pipeline 1010 and services 920 and / or hardware 922. In at least one embodiment, the pipeline manager 1012 can be configured to facilitate interactions from application to application, from application to service 920, and / or from application or service to hardware 922. In at least one embodiment, although shown as included in the software 918, this is not intended to be limiting, and in some examples, the pipeline manager 1012 can be included in the service 920. In at least one embodiment, the application orchestration system 1028 (e.g., Kubernetes, Docker, etc.) can include a container orchestration system that can group applications into containers as a logical unit for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from the deployment pipeline 1010 (e.g., rebuilding an application, partitioning an application, etc.) with individual containers, each application can execute in a self-contained environment (e.g., at the kernel level) to improve speed and efficiency.

[0114] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed separately (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer), which can allow for focus and attention on the tasks of a single application and / or container without being hindered by the tasks of other applications or containers. In at least one embodiment, the pipeline manager 1012 and the application coordination system 1028 can facilitate communication and collaboration between different containers or applications. In at least one embodiment, the application coordination system 1028 and / or the pipeline manager 1012 can facilitate communication and resource sharing between and among each application or container, as long as the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the configuration of the application or container). In at least one embodiment, because one or more applications or containers in the deployment pipeline 1010 can share the same services and resources, the application coordination system 1028 can coordinate, load balance, and determine the sharing of services or resources between and among the various applications or containers. In at least one embodiment, the scheduler can be used to track resource requirements of applications or containers, current or planned usage of those resources, and resource availability. Thus, in at least one embodiment, the scheduler can allocate resources to different applications and distribute resources between and among applications, taking into account the needs and availability of the system. In some examples, the scheduler (and / or other components of the application coordination system 1028) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0115] In at least one embodiment, the services 920 utilized and shared by the applications or containers in the deployment system 906 may include compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, and / or other service types. In at least one embodiment, an application may call (e.g., execute) one or more services 920 to perform processing operations for the application. In at least one embodiment, an application may utilize the compute services 1016 to perform supercomputing or other high performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1016 may be utilized to perform parallel processing (e.g., using a parallel computing platform 1030) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, a parallel computing platform 1030 (e.g., NVIDIA's ) can implement general-purpose computing on a GPU (GPGPU) (e.g., GPU 1022). In at least one embodiment, the software layer of the parallel computing platform 1030 can provide access to the GPU's virtual instruction set and parallel computing elements to execute computational kernels. In at least one embodiment, the parallel computing platform 1030 can include memory, and in some embodiments, the memory can be shared between and among multiple containers, and / or between and among different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls can be generated for multiple containers and / or multiple processes within a container to use the same data from a shared memory segment of the parallel computing platform 1030 (e.g., where multiple different stages of an application or multiple applications are processing the same information). In at least one embodiment, rather than copying and moving data to different locations in memory (e.g., read / write operations), the same data in the same memory location can be used by any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, because data is used to generate new data as a result of processing, information about the new location of the data can be stored and shared between the various applications. In at least one embodiment, the location of the data and the location of updated or modified data can be part of the definition of how the payload in the container is understood.

[0116] In at least one embodiment, AI service 1018 can be utilized to perform inference services for executing machine learning models associated with an application (e.g., tasked with executing one or more processing tasks of the application). In at least one embodiment, AI service 1018 can utilize AI system 1024 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, an application in deployment pipeline 1010 can use one or more output models 916 from training system 904 and / or other models of the application to perform inference on imaging data (e.g., DICOM data, RIS data, CIS data, REST-compliant data, RPC data, raw data, etc.). In at least one embodiment, two or more instances of inference using application coordination system 1028 (e.g., a scheduler) can be available. In at least one embodiment, the first category can include high-priority / low-latency paths that can achieve higher service level agreements, such as for performing inference on urgent requests in emergency situations or for radiologists during diagnostic procedures. In at least one embodiment, the second category may include a standard priority path that may be used for requests that may not be urgent or where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1028 may allocate resources (e.g., services 920 and / or hardware 922) for different reasoning tasks of the AI ​​service 1018 based on the priority path.

[0117] In at least one embodiment, shared memory can be installed into the AI ​​service 1018 in the architecture 1000. In at least one embodiment, the shared memory can operate as a cache (or other storage device type) and can be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of the deployment system 906 can receive the request and select one or more instances (e.g., for best fit, load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request can be entered into a database, and if it is not already in the cache, the machine learning model can be located from the model registry 924. A validation step can ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model can be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., a scheduler of the pipeline manager 1012) can be used to start the application referenced in the request. In at least one embodiment, if an inference server has not yet been started to execute the model, an inference server can be started. In at least one embodiment, any number of inference servers can be started per model. In at least one embodiment, in a pull model where the inference servers are clustered, the model can be cached whenever load balancing is beneficial. In at least one embodiment, the inference servers can be statically loaded into the corresponding distributed servers.

[0118] In at least one embodiment, inference can be performed using an inference server running in a container. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally with multiple versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, a model can be passed to the inference server when it is started, so that the same container can be used to serve different models as long as the inference server is running as a different instance.

[0119] In at least one embodiment, during application execution, an inference request for a given application may be received, a container (e.g., an instance hosting an inference server) may be loaded (if not already loaded), and a launcher may be called. In at least one embodiment, pre-processing logic in the container may load, decode, and / or perform any additional pre-processing on the incoming data (e.g., using the CPU and / or GPU). In at least one embodiment, once the data is ready for inference, the container may perform inference on the data as needed. In at least one embodiment, this may include a single inference call for a single image (e.g., a hand X-ray), or may request inference on hundreds of images (e.g., a chest CT scan). In at least one embodiment, the application may summarize the results before completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generated visualizations, or generated text summarizing the results. In at least one embodiment, different models or applications may be assigned different priorities. For example, some models may have real-time priority (turnaround time less than 1 minute), while other models may have a lower priority (e.g., turnaround time less than 10 minutes). In at least one embodiment, model execution time may be measured from the requesting mechanism or entity and may include collaborative network traversal time as well as execution time of the inference service.

[0120] In at least one embodiment, the transmission of requests between the service 920 and the inference application can be hidden behind a software development kit (SDK) and can provide robust transport via queues. In at least one embodiment, requests are placed in a queue for individual application / tenant ID combinations via an API, and the SDK pulls the request from the queue and delivers it to the application. In at least one embodiment, the name of the queue from which the SDK picks up the request can be provided. In at least one embodiment, asynchronous communication via queues can be useful because it allows any instance of the application to pick up work when it is available. In at least one embodiment, results can be transmitted back through the queue to ensure no data is lost. In at least one embodiment, queues can also provide the ability to partition work, as the highest priority work can enter a queue connected to the majority of instances of the application, while the lowest priority work can enter a queue connected to a single instance, which processes the tasks in the order they are received. In at least one embodiment, the application can run on a GPU-accelerated instance generated in the cloud 1026, and the inference service can perform inference on the GPU.

[0121] In at least one embodiment, visualization services 1020 can be utilized to generate visualizations for viewing application and / or deployment pipeline 1010 outputs. In at least one embodiment, visualization services 1020 can utilize GPU 1022 to generate visualizations. In at least one embodiment, visualization services 1020 can implement rendering effects such as ray tracing or other light transport simulation techniques to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, and the like. In at least one embodiment, a virtualized environment can be used to generate virtual interactive displays or environments (e.g., virtual environments) for system users (e.g., doctors, nurses, radiologists, etc.) to interact with. In at least one embodiment, visualization services 1020 can include internal visualizers, movies, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).

[0122] In at least one embodiment, hardware 922 may include a GPU 1022, an AI system 1024, a cloud 1026, and / or any other hardware for executing the training system 904 and / or the deployment system 906. In at least one embodiment, the GPU 1022 (e.g., NVIDIA and / or QUADRO GPUs) may include any number of GPUs that can be used to perform processing tasks for any feature or functionality of compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, other services, and / or software 918. For example, for AI services 1018, GPU 1022 may be used to perform pre-processing on imaging data (or other data types used by machine learning models), post-processing on the output of machine learning models, and / or perform inference (e.g., to execute machine learning models). In at least one embodiment, cloud 1026, AI system 1024, and / or other components of architecture 1000 may use GPU 1022. In at least one embodiment, cloud 1026 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1024 may use GPUs, and cloud 1026 (or at least part of a task that is deep learning or inference) may be executed using one or more AI systems 1024. Likewise, while hardware 922 is shown as discrete components, this is not intended to be limiting, and any component of hardware 922 may be combined with or utilized by any other component of hardware 922 .

[0123] In at least one embodiment, AI system 1024 may include a purpose-built computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPU, RAM, storage, and / or other components, features, or functions, AI system 1024 (e.g., NVIDIA's DGX TM ) can also include GPU-optimized software (e.g., a software stack) that can execute GPU-optimized software using multiple GPUs 1022. In at least one embodiment, one or more AI systems 1024 can be implemented in a cloud 1026 (e.g., in a data center) to perform some or all of the AI-based processing tasks of architecture 1000.

[0124] In at least one embodiment, cloud 1026 may include GPU-accelerated infrastructure (e.g., NVIDIA's NGC TM ), which can provide a GPU-optimized platform for executing processing tasks of the architecture 1000. In at least one embodiment, the cloud 1026 can include an AI system 1024 for executing one or more AI-based tasks of the architecture 1000 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the cloud 1026 can be integrated with an application coordination system 1028 that utilizes multiple GPUs to enable seamless scaling and load balancing between and among applications and services 920. In at least one embodiment, the cloud 1026 can be responsible for executing at least some of the services 920 of the architecture 1000, including the compute services 1016, AI services 1018, and / or visualization services 1020, as described herein. In at least one embodiment, the cloud 1026 can perform inference of large and small batches (e.g., executing NVIDIA's TensorRT TM ), providing accelerated parallel computing APIs and platforms 1030 (e.g., NVIDIA's ), execute application coordination system 1028 (e.g., Kubernetes), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher quality cinematic effects), and / or may provide other functionality for architecture 1000.

[0125] In at least one embodiment, to protect patient confidentiality (e.g., in situations where patient data or records are used off-site), the cloud 1026 may include a registry - such as a deep learning container registry. In at least one embodiment, the registry may store containers for instantiating applications that may perform pre-processing, post-processing, or other processing tasks on the patient data. In at least one embodiment, the cloud 1026 may receive data that includes patient data as well as sensor data in containers, perform the requested processing only on the sensor data in those containers, and then forward the resulting output and / or visualization to the appropriate parties and / or devices (e.g., local medical devices for visualization or diagnosis) without extracting, storing, or otherwise accessing the patient data. In at least one embodiment, the confidentiality of the patient data is preserved in accordance with HIPAA and / or other data regulations.

[0126] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.

[0127] Unless otherwise noted or clearly contradicted by the context, the use of the terms "a" and "an" and "the" and similar references in the context of describing the disclosed embodiments (particularly in the context of the appended claims) should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include," "have," "include," and "contain" should be interpreted as open-ended terms (meaning "including but not limited to"), unless otherwise noted. The term "connected" (when unmodified, refers to a physical connection) should be interpreted as partially or completely contained within, attached to, or connected together, even if there is some intervention. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were separately recited herein. In at least one embodiment, unless otherwise noted or contradicted by the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.

[0128] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C" are understood in context to generally refer to an item, clause, or the like, which may be A or B or C, or any non-empty subset of the set of A, B, and C. For example, in the illustrative example of a set having three members, the conjunctions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctions are not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless expressly indicated otherwise or contradicted by context, the term "plurality" refers to a plurality (e.g., "a plurality of items" refers to a plurality of items). In at least one embodiment, the number of items in the plurality of items is at least two, but may be more if expressly indicated or indicated by context. Further, unless stated otherwise or clear from context, the phrase "based on" means "based at least in part on" rather than "based solely on."

[0129] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are collectively executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of, for example, a computer program that includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) having executable instructions stored thereon, which, when executed by one or more processors of a computer system (i.e., as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all of the code, but rather the plurality of non-transitory computer-readable storage media collectively store all of the code. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.

[0130] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system comprising multiple devices operating in different ways such that the distributed computer system performs the operations described herein and such that no single device performs all of the operations.

[0131] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0132] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0133] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0134] Unless expressly stated otherwise, it is understood that throughout this specification, terms such as “process,” “calculate,” “compute,” “determine,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.

[0135] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process may refer to multiple processes to execute instructions continuously or intermittently, sequentially, or in parallel. In at least one embodiment, the terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.

[0136] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in a variety of ways, such as by receiving data as parameters of a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting data as input or output parameters of a function call, an application programming interface, or an interprocess communication mechanism.

[0137] Although the description herein sets forth example embodiments of the described technology, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities are defined above for descriptive purposes, the various functions and responsibilities may be assigned and divided in different ways depending on the circumstances.

[0138] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A method comprising: providing, using a processing device executing an application programming interface (API), a visual representation of a plurality of document processing pipelines (DPPs) for presentation within a user interface; receiving, from the user interface via the API, a selection of a DPP from the plurality of DPPs; Segment the input document into multiple segments according to the predetermined settings of the selected DPP; causing an embedding model to process the plurality of segments to generate a plurality of embeddings; as well as The plurality of embeddings is caused to be stored in a data store.

2. The method of claim 1 , wherein the predetermined setting comprises one or more of: the size of individual fragments in the plurality of fragments; an amount of overlap between adjacent segments of the plurality of segments; or Selection of the embedding model.

3. The method of claim 1, further comprising: The plurality of fragments are caused to be stored in at least the data store or a second data store.

4. The method of claim 3, further comprising: Indexed data mapping the plurality of embeddings to the plurality of segments is stored.

5. The method of claim 1, wherein the input document is received from a client device in remote communication with the processing device over a network.

6. The method of claim 1, further comprising: receiving a container image including the API from a remote computing device; as well as The API is executed in a container instantiated using the container image.

7. The method of claim 6, wherein the container image further comprises at least one of the following: a segmentation engine that segments the input document into the plurality of segments, or The embedding model.

8. The method of claim 1, further comprising: Receive queries using the API; causing the embedding model to process the query to generate one or more query embeddings; calculating a plurality of similarity scores characterizing similarity between the one or more query embeddings and the plurality of embeddings; selecting one or more of the plurality of segments using the plurality of similarity scores; as well as Hints for a language model LM are generated, wherein the hints are based on at least the query and the selected one or more segments.

9. A method comprising: receiving a query using a processing device executing an application programming interface (API); causing an embedding model to process the query to generate one or more query embeddings; calculating a plurality of similarity scores characterizing a similarity of the one or more query embeddings to a plurality of embeddings associated with one or more stored documents; selecting one or more segments from the one or more stored documents using the plurality of similarity scores; as well as A LM hint is processed using a language model LM to obtain a response to the query, wherein the LM hint is based on at least the query and the selected one or more segments.

10. The method of claim 9, further comprising: A selection of a DPP from among a plurality of document processing pipeline DPPs provided using the API is received, the selected DPP including a maximum number of segments to be identified.

11. The method of claim 9, wherein selecting the one or more segments comprises: identifying one or more embeddings from the plurality of embeddings using the plurality of similarity scores, the identified one or more embeddings corresponding to the one or more segments associated with the query; as well as The one or more segments are ranked according to relevance to the query using the plurality of similarity scores.

12. The method of claim 9, wherein selecting the one or more segments comprises: performing a document search to identify one or more additional segments of the one or more stored documents, the one or more additional segments having a textual relevance to the query; as well as A ranking model is used to rank a set of snippets according to relevance to the query, wherein the set of snippets comprises: the one or more fragments, and the one or more additional segments; as well as The LM hints are generated using a ranked set of fragments.

13. The method of claim 9, wherein selecting the one or more segments comprises: Stored indexed data is accessed that maps the plurality of embeddings to the one or more stored documents.

14. The method of claim 9, further comprising: Receive a container image from a remote computing device, the container image comprising one or more of: the API, or the embedding model; as well as The one or more of the API or the embedded model is executed in a container instantiated using the container image.

15. The method of claim 9, wherein individual documents of the one or more stored documents are stored using operations comprising: receiving, for the individual document, a selection of a DPP from among a plurality of document processing pipelines (DPPs) provided using the API; Segmenting the individual document into a plurality of segments according to a predetermined setting of the selected DPP; causing the embedding model to process the plurality of segments to generate a set of embeddings for the individual documents; as well as The embedding set is caused to be stored in a data store.

16. The method of claim 15, wherein the predetermined settings include one or more of the following: the sizes of individual fragments of the plurality of fragments, an amount of overlap between adjacent segments in the plurality of segments, or Selection of the embedding model.

17. A system comprising: Processing equipment for: Receive queries using an application programming interface (API); causing an embedding model to process the query to generate one or more query embeddings; calculating a plurality of similarity scores characterizing a similarity of the one or more query embeddings to a plurality of embeddings associated with one or more stored documents; selecting one or more segments of the one or more stored documents using the plurality of similarity scores; as well as A LM hint is processed using a language model LM to obtain a response to the query, wherein the LM hint is based on at least the query and the selected one or more segments.

18. The system of claim 17, wherein to select the one or more segments, the processing device is configured to: identifying one or more embeddings from the plurality of embeddings using the plurality of similarity scores, the identified one or more embeddings corresponding to the one or more segments associated with the query; and The one or more segments are ranked according to relevance to the query using the plurality of similarity scores.

19. The system of claim 17, wherein to select the one or more segments, the processing device is configured to: performing a document search to identify one or more additional segments of the one or more stored documents, the one or more additional segments having a textual relevance to the query; as well as A ranking model is used to rank a set of snippets according to relevance to the query, wherein the set of snippets comprises: the one or more fragments, and the one or more additional segments; as well as The LM hints are generated using a ranked set of fragments.

20. The system of claim 17, wherein the system is included in at least one of: In-vehicle infotainment systems for autonomous or semi-autonomous machines; a system for performing simulation operations; Systems for performing digital twin operations; a system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems implemented using edge devices; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; Systems implemented using robots; Systems for performing conversational AI operations; Systems for generating synthetic data; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.