PROMPT SUITABILITY ANALYSIS FOR LANGUAGE MODEL-BASED AI SYSTEMS AND APPLICATIONS

By using a prompt analyzer to evaluate the suitability of user prompts for large language models, the challenges of hallucinations, fractal inaccuracies, and malicious prompts are addressed, resulting in improved quality and security of LLM responses.

DE102024136304A1Pending Publication Date: 2025-06-12NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024136304
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Large language models (LLMs) face challenges in generating accurate and reliable responses due to hallucinations, fractal inaccuracies, and the inability to handle prompts that exceed their training data scope, leading to potential malicious prompt injection attacks and policy violations.

Method used

The implementation of a prompt analyzer that evaluates the suitability of user prompts for LLM processing by determining Promptverifizierungsscores, which indicate the probability of token transitions within the prompt. This analysis helps determine whether the prompt can be processed accurately and reliably by the LLM.

Benefits of technology

The prompt analysis technique reduces the need for evaluating LLM outputs, improves the quality and security of LLM responses by identifying improper or malicious prompts, and maintains high processing speed by front-end prompt verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Apparatus, systems, and techniques are disclosed that evaluate the suitability of prompts for language model (LM) processing for improved quality and security of LM outputs. The techniques include: determining prompt verification score(s) including a first subset of tokens and a second subset of tokens, and obtaining, using an LM, the individual prompt verification score indicating a probability that the second subset of tokens appears in the prompt along with the first subset of tokens. The techniques further include: determining, using the prompt verification score(s), whether to provide the prompt to the LM.
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF THE INVENTION

[0001] At least one embodiment relates to computational resources used to perform and facilitate natural language technologies. For example, at least one embodiment relates to systems and techniques that facilitate and improve the processing and quality of user questions and queries through language models. GENERAL STATE OF THE ART

[0002] Well-trained language models—such as large language models (LLMs)—are capable of supporting natural language conversations, understanding the speaker's intent and emotions, explaining complex topics, receiving new text upon appropriate prompts, writing and debugging software code, providing users with advice on topics of interest, processing image, audio, and / or other data types, and / or performing other functions. Depending on the implementation, LLMs typically undergo self-supervised training on massive amounts of text and / or other data types and learn to identify next and / or missing tokens (which correspond to subwords, symbols, words, etc.).may correspond) in an expression / sentence, detect the intent and / or emotion of a human speaker, determine whether two sentences are related or unrelated, and / or perform other basic language tasks. After initial training, LLMs often undergo instructional (prompt-based), supervised fine-tuning that prompts LLMs to acquire a more detailed level of language proficiency and / or master more specialized tasks. Supervised fine-tuning includes learning prompts (questions, hints, etc.) accompanied by sample texts (e.g., answers, sample essays, etc.) that serve as training ground truth. As the fine-tuning intensifies, a human evaluator assigns ranks that indicate how closely the generated text resembles human-produced texts. OVERVIEW

[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not be within the scope of the claims.

[0004] Apparatus, systems, and techniques are disclosed that evaluate the suitability of prompts for language model (LM) processing for improved quality and security of LM outputs. The techniques include: determining prompt verification score(s) including a first subset of tokens and a second subset of tokens, and obtaining, using an LM, the individual prompt verification score indicating a probability that the second subset of tokens appears in the prompt along with the first subset of tokens. The techniques further include: determining, using the prompt verification score(s), whether to provide the prompt to the LM.

[0005] The disclosure extends to any novel aspects or features described and / or illustrated herein.

[0006] Further features of the disclosure are characterized by the independent and dependent claims.

[0007] Any feature in one aspect of the disclosure may be applied to other aspects of the disclosure in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.

[0008] Furthermore, features implemented in hardware may also be implemented in software, and vice versa. Any reference to software and hardware features in this specification should be construed accordingly.

[0009] Any system or device feature described herein may also be provided as a method feature, and vice versa. Functionally described system and / or device aspects (including means-plus-function features) may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.

[0010] It should also be understood that certain combinations of the various features described and defined in any aspects of the disclosure may be implemented and / or provided and / or used independently of one another.

[0011] This disclosure also provides computer programs and computer program products comprising software code adapted, when executed on a computing device, to perform any of the methods described herein and to embody any of the device and system features described herein, including any or all of the substeps of a method.

[0012] The disclosure also provides a computer or computing system (including networked or distributed systems) having an operating system that supports a computer program for performing any of the methods described herein and / or embodying any of the device or system features described herein.

[0013] The disclosure also provides a computer-readable medium having stored thereon one or more of the aforementioned computer programs.

[0014] The disclosure also provides a signal carrying one or more of the aforementioned computer programs.

[0015] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0016] Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram of an exemplary computing system, according to at least one embodiment, capable of implementing suitability analysis of user prompts in the context of language model (LM) processing for improved quality and security of LM outputs; Fig. 2 illustrates an example computing device 200, according to at least one embodiment, that supports live suitability analyses of prompts prior to LM processing; Fig. 3 illustrates an architecture of an exemplary prompt analyzer, according to at least one embodiment, used to improve the quality and security of LM outputs; Fig. 4 illustrates an architecture of an example system, according to at least one embodiment, that uses prompt analysis to select one of multiple LMs for prompt processing; Fig. 5 illustrates a data flow of exemplary operations of an LM with prompt analysis processing according to at least one embodiment; Fig. 6 is a flow diagram of an exemplary method, according to at least one embodiment, that performs a suitability analysis of prompts prior to LM processing for improved quality and security of LM outputs; Fig. 7A illustrates inference and / or training logic according to at least one embodiment; Fig. 7B illustrates inference and / or training logic according to at least one embodiment; Fig. 8 illustrates the training and deployment of a neural network according to at least one embodiment; Fig. 9 is an exemplary dataflow diagram for an advanced compute pipeline according to at least one embodiment; and Fig. 10 is a system diagram for an example system for training, adapting, instantiating, and deploying machine learning models in an advanced compute pipeline, according to at least one embodiment. DETAILED DESCRIPTION

[0017] LLM capabilities are determined, at least in part, by the quantity and variety of training data used to train the LLM. However, even LLMs trained on enormous amounts of data can be affected by hallucinations (recalling irrelevant or random information), factual inaccuracies, and / or other similar situations that diminish user experiences and the reliability of LLM output. Such situations arise due to a virtually unlimited number of topics that can be the subjects of user prompts (questions, queries, instructions, etc.) and / or ways in which user prompts can be phrased, with the range of topics / phrasing inevitably exceeding the volume of training data, however enormous.In such situations, an LLM tasked with responding to a user prompt is unlikely to produce a quality response that would be helpful to the user. However, evaluating LLM responses for appropriateness and factual accuracy (often referred to as "model alignment") is a very challenging task. Automating such a task may require training one or more arbiter models (e.g., models trained on a specific topic of the prompts) to at least the same level of sophistication as the LLM being supervised. This significantly increases the cost and duration of model training.

[0018] Furthermore, some of the prompts may occur as part of a malicious prompt injection attack, during which an attacker attempts to manipulate an LLM to produce a desired (e.g., offensive, deceptive, fraudulent, etc.) output. For example, during a user-chatbot conversation, an attacker may attempt to intercept user prompts and supplement them with additional information and / or instructions, thereby stealing the chatbot's personality and causing the chatbot to respond in an inappropriate manner. In still other situations, an LLM may receive a prompt requesting specific knowledge the LLM lacks (e.g., "What are the side effects of zidovudine?"), a prompt requesting assistance in manufacturing an illegal device or substance, and / or another request that conflicts with the public policy or scope of a user agreement with the LLM services.As is the case with LLM hallucinations, for example, detecting such situations based on analysis of LLM outputs can be a difficult and costly process.

[0019] Aspects and embodiments of the present disclosure address these and other technological challenges by providing systems and techniques that analyze the suitability of user prompts for LLM processing before such processing is attempted. Prompt analysis reduces the need to evaluate the quality and accuracy of actual LLM outputs by leveraging the self-training of the LLMs to determine the boundaries of the LLM level. In particular, a prompt analyzer may analyze a sequence of prompt tokens (numerical representations of words of the prompt) {T j} = T1, T2, ... T Nevaluate, and obtain a statistical distribution and an evaluation metric M that characterizes a probability that the sequence of tokens {T j} corresponds to a type encountered during LLM training, as opposed to an unseen type of prompt, a prompt requiring specialized knowledge, a prompt likely to be used by a malicious attacker, a request for information that is against the LLM service policy, and / or the like. The prompt analyzer may use the LLM's learned ability to determine a likely next token and follow a sequence of known tokens. In some embodiments, the LLM may be used to determine a probability of a missing token T m in a sequence of given tokens, T1, ..., T m -1, T m+1 ... T jinstead of (or in addition to) determining the probability of the next token(s).

[0020] In particular, the prompt analyzer can form a verification prompt that contains n first tokens of the user prompt, T1, T2, ... T j , where j can be 1, 2, ... N - 1, depending on a specific iteration of the prompt verification. The verification prompt T1, T2, ... T j can then be used as input to the LLM, thereby determining the probability P j determines that the next word of the prompt, T j+1 , will follow the first j tokens of the prompt. After a sequence of N - 1 such probabilities P1, P2 ... P N-1 obtained, the evaluation metric M can be calculated, e.g., using the product, M=∏j=1N−1Pj (which can be normalized by the number of tokens in the prompt to accommodate prompts of varying length, e.g. M=N∏j=1N−1Pj, The evaluation metric M estimates the probability that the user prompt is of a type represented in the corpus of training data used in LLM training.

[0021] A low metric M indicates that some token transitions T j → T j+1 in the user prompt correspond to unusual word combinations that the LLM did not learn during training (or did not learn with sufficient reliability). Provided that the metric M is below a threshold metric, M T(which can be determined empirically based on field tests), the prompt analyzer may determine that the LLM is unable to generate an accurate and reliable response to the user prompt, or that the prompt is of a type that the LLM is not intended to process. In such situations, the prompt may be returned to the user with an explanation that the LLM cannot process the prompt, a recommendation that the prompt be expressed differently, and / or the like. In situations where M ≥ M T , the prompt analyzer can determine that the LLM is likely to generate an acceptable, reliable, and accurate response and can pass the user prompt to the LLM for regular processing.

[0022] In some embodiments where multiple LLMs (e.g., specialized knowledge models) are available, the prompt analyzer may perform multiple prompt verifications (e.g., in parallel) for the multiple models and then select the model with the highest metric M, indicating the corresponding model's best knowledge of the subject matter of the user prompt. In some embodiments, instead of selecting from multiple LLMs, the system may select from different prompt matching models trained to adapt the prompt to a specific domain for which the prompt matching model is configured. As such, if the LLM itself does not have a high M-score, the LLM in combination with a specific prompt matching model may have an acceptable M-score, making the combination suitable for addressing the query or a particular task at hand.

[0023] The advantages of the disclosed techniques include, among others, the identification of inappropriate or malicious prompts, prompts that exceed the learned capabilities of the model, and / or the like, without the need for specialized arbiter models (or human arbiters). This provides quality and security for language model processing while maintaining high speed for such processing. Front-end prompt verification reduces the number of confusing, misleading, and / or factually incorrect language model responses, improves user satisfaction, and reduces situations of language model misuse. SYSTEM ARCHITECTURE

[0024] Fig. 1 is a block diagram of an exemplary computing system 100, according to at least one embodiment, capable of implementing a suitability analysis of user prompts in the context of language model (LM) processing for improved quality and security of LM outputs. As in Fig. 1, the computing system 100 may include a computing device 102, a data store 150, and an LM training server 160 connected to a network 140. The network 140 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wireless network, a personal area network (PAN), a combination thereof, and / or another network type.

[0025] Computing device 102 may include a desktop computer, a laptop computer, a smartphone, a tablet computer, a server, a wearable device, a virtual / augmented / mixed reality headset or head-up display, a digital avatar or chatbot kiosk, an in-vehicle infotainment computing device, and / or any suitable computing device capable of performing the techniques described herein. Computing device 102 may be configured to receive a prompt 101, which may be any human-generated or machine-generated data capable of being used directly or after appropriate preprocessing as input to an LM 122. In some embodiments, LM 122 may be an LLM, e.g., a model with hundreds of millions or one or more billion learned parameters. Prompt 101 may be text (e.g.,a sequence of one or more typed words), speech (e.g., a sequence of one or more spoken words), an image (e.g., a drawing or picture), a video or series of images, a tokenized sequence of sensor data, a sequence of sensor data, and / or a combination thereof. The prompt 101 may be generated as part of an interaction of a user (including a user of a computing program or other machine user) with an entity that deploys, uses, and / or interfaces with the LM 122. Such an entity may include a chatbot, a digital avatar, a digital assistant, an in-vehicle communication system, a gaming application, and / or another digital agent capable of using a natural language prompt.The interaction may include a private conversation, a customer-agent session, a browsing session, an information gathering session, a research request, a multi-speaker conversation, a public conversation, a work session, and / or a combination thereof. The prompt 101 may be formulated as a statement, query, question, request for explanation / instruction, request for advice, expression of emotion, narrative (or as part of a narrative), and / or any type of input requiring a response from the LM 122, and / or the like. In some embodiments, the prompt 101 may include text and / or speech generated by a computer, e.g., a language model different from the LM 122.

[0026] The prompt 101 may be received via any suitable user interface (UI) 106, which may include one or more devices of various modalities. For example, the prompt 101 may be received via a keyboard, a touchscreen, a touchpad, a writing pad, a graphical interface, a mouse, a stylus, and / or using any other pointing device capable of selecting words / phrases displayed on a screen and / or any other suitable device. In some embodiments, the UI 106 may include an audio device, e.g., a microphone, a video device, such as a digital camera, to capture a video prompt, a sequence of two or more images (video frames). In some embodiments, text, voice, and / or video input devices may be integrated together (e.g.,in a smartphone, tablet computer, desktop computer and / or the like).

[0027] The prompt 101 may be captured (e.g., typed, recorded, etc.) live using one or more peripheral devices connected to the computing device 102 (e.g., keyboard, microphone, digital camera), retrieved from a memory 104 of the computing device 102, and / or received from an external computing device via a local network connection (e.g., via the network 140). The text prompt(s) 101 may be in any suitable text format, e.g., plain text format, document format (DOC), rich text format (RTF), hypertext markup language (HTML) format, extensible markup language (XML) format, portable document format (PDF), and / or the like. The voice prompt(s) 101 may be in any suitable audio format, e.g., audio format, or the like. B. WAV, AIFF, MP3, AAC, WMA or any other compressed or uncompressed audio format.Likewise, the video prompt(s) 101 may be in any suitable video format, such as raw video data format, MPEG-4 format, MOV format, WMV format, AVI format, or other compressed or uncompressed video format.

[0028] Computing device 102 may include a memory 104 (e.g., one or more memory devices or units) communicatively coupled to one or more processing devices, such as one or more graphics processing units (GPU) 110, one or more central processing units (CPU) 130, one or more data processing units (DPUs), one or more parallel processing units (PPUs), and one or more other processing units (e.g., field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and / or the like). Memory 104 may include read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM), static memory, such as static random access memory (SRAM), and / or other memory capable of storing digital data.The memory 104 may store a prompt analyzer 120, an LM 122, and an application 125. The application 125 may include software that utilizes the capabilities of the LM 122 to provide natural language services to users. For example, the application 125 may include a chatbot application, a browser application, a digital avatar application, a digital assistant application, a chatbot application, and / or the like. The prompt analyzer 120 may implement various techniques of the present disclosure. For example, the prompt analyzer 120 may parse the prompt 101 into individual words or tokens and construct one or more verification prompts that include some of the words / tokens of the prompt 101 while excluding some other words / tokens of the prompt 101.The prompt analyzer 120 may feed the constructed verification prompts to the LM 122 and receive from the LM 122 various probabilities (verification scores) that indicate the probabilities that one or more words / tokens not included in the verification prompts may occur together with words / tokens from verification prompts as part of the same prompt 101. Based on the received probabilities, the prompt analyzer 120 may determine whether the prompt 101 is a valid prompt or an invalid prompt. A valid prompt is a prompt of a type that the LM 122 has been trained to process and whose processing does not contradict any relevant (public and / or private) policy. An invalid prompt is a prompt that the LM 122 has not been sufficiently trained to process (e.g.,a prompt that requires specialized knowledge not learned by the LM 122), a prompt that violates a relevant policy, a prompt that is likely to generate a response that would violate a relevant policy, an unusual prompt, a prompt that may have been generated by a malicious attacker, and / or the like. Prompt 101 determined to be valid may be forwarded to the LLM 122 for regular processing. Prompt 101 determined to be invalid may be returned to a user (human or machine) who generated prompt 101 with a recommendation to express prompt 101 differently or a notification that prompt 101 cannot be processed.

[0029] In some embodiments, the LM 122 may be a model trained and deployed by an external (relative to the computing device 102) entity, e.g., language model service 170, which may be a cloud service, a subscription service, and / or a combination thereof. In some embodiments, the LM 122 may be trained by the LM training server 160. In some embodiments, the LM 122 (and / or other deployed language models) may be or include a large language model (LLM). The LM 122 may be trained to capture syntax and semantics of human language, e.g., by predicting a next, a previous, and / or a missing word in a sequence of words (e.g., one or more sentences of human language or text, for example).The LM 122 may further be trained using training data containing a large number of texts, such as human dialogues, newspaper texts, magazine texts, book texts, web-based texts, and / or other texts. Because ground truth for such training is embedded in the texts themselves, the LM training server 160 may use such texts for self-supervised training of the LM 122. This teaches the LM 122 how to conduct a conversation with a user (a human user or a computer) in a natural language in a manner that closely resembles a dialogue with a human speaker, including understanding the user's intent and responding in a manner the user expects from a conversation partner.

[0030] Following the initial self-supervised training, the LM training server 160 may implement instructional (e.g., prompt-based) training of the LM 122, e.g., supervised fine-tuning, which prompts the LM 122 to acquire a more detailed level of language proficiency and / or master more specialized tasks. The supervised fine-tuning may include learning prompts (questions, hints, etc.) accompanied by sample texts (e.g., answers, sample essays, etc.) that serve as training ground truth. In some embodiments, the LM training server 160 may further implement intensification training, where a human evaluator assigns ranks indicating how closely the generated text resembles a human-produced text.

[0031] In some embodiments, the LM training server 160 may train multiple LMs 122, which may be models that differ by the number of neurons, the number of neuron layers, specific neural architecture, and / or the like. Different trained LMs 122 may be trained (e.g., fine-tuned) using specialized texts, e.g., medical texts, mathematical texts, scientific texts, computer technology texts, and / or the like.

[0032] The LM 122 may be implemented using neural networks with a large number (e.g., billions) of artificial neurons. In at least one embodiment, the LM 122 and / or other deployed models may be implemented as deep learning neural networks having multiple layers of linear and nonlinear operations. For example, the LM 122 may include convolutional neural networks, recurrent neural networks, fully connected neural networks, long-short-term memory (LSTM) neural networks, neural networks with attention, e.g., transformer neural networks, a combination of a convolutional network and one or more transformers (a conformer), and / or neural networks of other types.In at least one embodiment, the LM 122 may include multiple neurons, where an individual neuron receives its input from other neurons and / or from an external source and produces an output by applying an activation function to the sum of weighted (using trainable weights) inputs and possibly a bias value. In at least one embodiment, the LM 122 may include multiple neurons arranged in layers, including an input layer, one or more hidden layers, and / or an output layer. Neurons from adjacent layers may be connected by weighted edges.

[0033] Initially, parameters (e.g., edge weights and biases) of LM(s) 122 may be assigned initial values ​​(e.g., randomized values). For different training inputs, LM training server 160 may cause LM(s) 122 to generate one or more training outputs. LM training server 160 may then compare the training output(s) with the desired target output(s). The resulting error or discrepancy, e.g., the difference between the target output(s) and the training output(s), may be backpropagated across different neural layers of LM(s) 122, and the weights and biases of LM(s) 122 may be adjusted to bring the training outputs closer to the target (ground truth) outputs. This adjustment may be repeated until the output error for a given training input satisfies a predetermined condition (e.g., B. falls below a predetermined value).A different training input can then be selected, a new training output generated, and a new series of adjustments implemented until LM(s) 122 is / are trained to a target accuracy level or until LM(s) 122 converges to a limit of accuracy.

[0034] In some embodiments, operations of the prompt analyzer 120 may generate additional training data that may be used to retrain the LM 122 (or multiple LMs 122). For example, the prompts 101 determined to be invalid by the prompt analyzer 120 may be stored in the data store 150. In addition, various probability distributions (e.g., next token probabilities calculated for various prompts 101) may also be stored in the data store 150. Stored prompts 152 and / or stored probability distributions 154 may then be used by the LM training server 160 to retrain the LM 122. For example, stored prompts 152 may be used (e.g.,using a keyword search of stored prompts 152) to identify subject areas in which the LM 122 lacks expertise and to identify an additional corpus of text and / or other data for retraining the LM 122. Similarly, stored probability distributions 154 may be used to predict next-word transitions (e.g., transitions between tokens T1...T. j and Token T j+1 ) in stored prompts 152 that have been determined to be low-probability transitions (thereby causing the prompt analyzer 120 to classify the corresponding stored prompts 152 as invalid). The LM training server 160 can then identify texts (e.g., using a suitable search engine crawler) that include such and / or similar transitions and retrain the LM 122 using these identified texts.

[0035] The data store 150 can be accessed by the LM training server 160 and / or the computing device 102 directly (e.g., via a bus, an interconnect, and / or the like) or (as in Fig. 1) over the network 140. The data store 150 may be persistent storage capable of storing audio data as well as metadata for the stored audio files. The data store 150 may be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tapes or hard drives, network-attached storage (NAS), storage area network (SAN), and so on. Although illustrated as separate from the LM training server 160 and / or the computing device 102, in at least some embodiments, the data store 150 may be part of the LM training server 160 and / or the computing device 102.In at least some embodiments, the data store 150 may be a network-attached file server, while in other embodiments, the data store 150 may be another type of persistent storage, such as an object-oriented database, a relational database, etc., that may be hosted by the LM training server 160 and / or the computing device 102 or one or more different machines coupled to the LM training server 160 and / or the computing device 102 via the network 140.

[0036] In at least one embodiment, the LM training server 160 and the computing device 102 may be implemented on a single computing device. The LM training server 160 and / or the computing device 102 may be (and / or include) a rackmount server, a router computer, a personal computer, a laptop computer, a tablet computer, a desktop computer, a media center, or a combination thereof.

[0037] Fig. 2 illustrates an example computing device 200, according to at least one embodiment, that supports live suitability analyses of prompts prior to LM processing. In at least one embodiment, computing device 200 may be part of computing device 102. In at least one embodiment, computing device 200 may be part of LM training server 160. In at least one embodiment, computing device 200 supports prompt analyzer 120, which includes (but is not limited to) a verification prompt generator 210, a prompt scorer 220, and a language model 122. Prompt analyzer 120 may be capable of processing prompt 101 and generating a response 230. Response 230 may be a text response, a speech (voice) response, e.g., a text response additionally converted to speech via text-to-speech processing, and / or the like.

[0038] Operations of the prompt analyzer 120 and / or one or more LMs 122 may be performed using one or more GPUs 110, one or more CPUs 130, one or more parallel management units (PPUs) or accelerators, such as a deep learning accelerator, data processing units (DPUs), and / or the like. In at least one embodiment, a GPU 110 may include multiple cores 211, each core capable of executing multiple threads 212. Each core may run multiple threads 212 concurrently (e.g., in parallel). In at least one embodiment, threads 212 may have access to registers 213. The registers 213 may be thread-specific registers with access to a register restricted to a respective thread. Furthermore, shared registers 214 may be accessed by one or more (e.g., all) threads of the core.In at least one embodiment, each core 211 may include a scheduler 215 to distribute computational tasks and processes among various threads 212 of the respective core 211. A dispatcher 216 may implement scheduled tasks on corresponding threads using proper private registers 213 and shared registers 214. The computing device 200 may include one or more input / output components 217 to facilitate the exchange of information with one or more users or developers.

[0039] In at least one embodiment, the GPU 110 may include a (high-speed) cache 218, wherein multiple cores 211 may share access to it. Furthermore, the computing device 200 may include a GPU memory 219, wherein the GPU 110 may store intermediate or final results (outputs) of various computations performed by the GPU 110. Upon completion of a respective task, the GPU 110 (or the CPU 130) may move the output to the (main) memory 104. In at least one embodiment, the CPU 130 may execute processes that include serial computational tasks, wherein the GPU 110 may perform tasks (such as multiplying inputs of a neural node by weights and adding biases) that are amenable to parallel processing.In at least one embodiment, the prompt analyzer 120 and / or the LMs 122 may determine which processes to execute on the GPU 110 and which processes to execute on the CPU 130. In other embodiments, the CPU 130 may determine which processes to execute on the GPU 110 and which processes to execute on the CPU 130.

[0040] The systems and methods described herein may be used for a variety of purposes, including, but not limited to, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, generative AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.

[0041] Disclosed embodiments may be included in a number of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, antenna systems, medical systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twinning operations, systems implemented using an edge device, systems for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems,that are implemented at least partially in a data center, systems for performing generative AI operations, systems for performing conversational AI operations, systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems that implement one or more language models, such as large language models (LLMs) (which can process text, voice, image, and / or other data types to generate output in one or more formats), systems that are implemented at least partially using cloud computing resources, and / or other types of systems. SYSTEMS AND TECHNIQUES FOR PROMPT ANALYSIS

[0042] Fig. 3 illustrates an architecture 300 of an exemplary prompt analyzer, according to at least one embodiment, used to improve the quality and security of LM outputs. In its embodiment, the Fig. 3 illustrated prompt analyzer of the prompt analyzer 120 of Fig. 1 and Fig. 2. Various in Fig. 3 with the same reference numerals as the blocks of Fig. 1 and / or Fig. 2 marked blocks implement the same (or similar) functionality. Different blocks of Fig. 3 may correspond to modules and components located on a single computing device or distributed across multiple computing devices. For example, the LM 124 may be located on the same server (e.g., computing device 102 of Fig. 1) such as the prompt analyzer 120 or on another server (e.g., one or more computing devices of the language model service 170 of Fig. 1). The prompt 101 can be implemented using the UI 106, e.g., as in connection with Fig. 1 described, can be received.

[0043] A prompt extension 305 may extend the prompt 101 with any suitable information provided by the application 125. For example, the prompt extension 305 may include a personality assigned to the LM 124 by the application 125, various contextual information related to the prompt 101, identification of the type (e.g., subject matter) of the prompts 101 to which the LM 124 should or should not respond, definitions of response formats, and / or the like.

[0044] The prompt 101 may be tokenized 302 to represent a sequence of words W1, W2... of the prompt 101 via a sequence of tokens recognizable by the LM 124. A set of tokens understood by the LM 124 may be LM-specific (different for different models and model builders) and fixed during training of the LM 124. The set of tokens may include a suitable representation of units of language (e.g., syllables, words, etc.) as numbers. In an example of GPT-4 tokens, the word "die" may be represented via the token "280," the word "import" may be represented via the token "476," the word "description" may be represented via the token "4097," etc. In some embodiments, individual words may be represented using any number of tokens or word transitions.For example, a long word or a word containing multiple words may be represented using multiple tokens, e.g., with one token representing an initial portion of the word and another token(s) representing a middle or final portion of the word. In some situations, even a long / compound word may be represented by a single token. As such, tokenization 302 may be performed in any manner suitable for input to the LM 124.

[0045] In some embodiments, appropriate preprocessing (in Fig. 1 not shown) prior to tokenization 302. For example, in the situation of speech prompts 101, preprocessing may include audio filtering, denoising, amplification, dereverberation, segmentation, audio signal enhancement, and speech-to-text processing to identify the sequence of spoken words W1, W2... in the prompt 101. Speech-to-text processing may be performed using a trained automatic speech recognition (ASR) model and may further include removing portions of the audio recording that correspond to filler words / sounds or that do not have speech content (pauses, noises, etc.).

[0046] Tokenization 302 transforms the sequence of words into a sequence of tokens W1, W2, ... → {T j} = T1, T2, ... T N. The number of tokens N may differ from the number of written or spoken words in the prompt 101 (e.g., be greater or lesser). Instead of providing the received sequence of tokens {Tj} directly to the LM 124, as would be the case with conventional systems, tokens {T j} must first be provided to the prompt analyzer 120. The prompt analyzer 120 may use a verification prompt generator 210 to generate one or more verification prompts 310 consisting of some of the tokens {T j} and provided to the LM 124 with instructions to predict tokens not included in verification prompts 310.

[0047] In an exemplary embodiment, N−1 verification prompts 310 may be generated. The first verification prompt may include a first token T1, a second token T2, and an instruction to the LM 124 to predict a conditional probability P(T2|T1) that the second token T2 follows the first token T1. Similarly, the second verification prompt may include a first token T1, a second token T2, a third token T2, and an instruction to the LM 124 to predict a conditional probability P(T3|T1, T2) that the third token T3 follows the sequence of the first token T1 and the second token T2. ​​The jth verification prompt may include a sequence of tokens T1, T2, ... T j together with the next token T j+1 and include an instruction to the LM 124 to calculate a conditional probability P(T j+1 |T 1, T2, ... T j ) that the j+1st token T j+1 the sequence of tokens T1, T2, ... Tj to predict what follows.

[0048] When generating and receiving responses to verification prompts 310, the prompt analyzer 120 utilizes the learned capability of the LM 124 to determine a likely next token following a sequence of given tokens. The LM 124 may be trained to process an input string of tokens and generate probabilities (e.g., using a softmax classifier) ​​for various tokens of the token corpus (vocabulary) of the LM 124 (which may include thousands of tokens). The prompt analyzer 120 may then receive, from the LM 124, the conditional probability 320 for the target token T j+1 , the next token Pj ≡ P(T j+1 |T1, T2, ... T j ) for the target token T j+1For example, the prompt analyzer 120 may access a neuron in the output layer (softmax layer) of the LM 124 that provides a next token probability for the target token T j+1 After various generated verification prompts 310 have been processed and a set of N - 1 such token probabilities {P j} = P1, P2 ... P N-1 has been received, the prompt evaluation 220 may evaluate the received probabilities to determine a usability of the prompt 101. In particular, prompts 101 corresponding to the level domain of LM 124 may be evaluated by the set of probabilities {P j} that have relatively uniform values, e.g., without a significant reduction in one or more probabilities, indicating unexpected transitions between tokens (words). Such unexpected transitions may signal a malicious attack or a lack of learned capabilities of the LM. An appropriate response to such unexpected prompts 101 may be to refuse to process the prompts.

[0049] In some embodiments, decision making regarding the processing of the prompt 101 may be performed by calculating a suitable evaluation metric M that is suitable for the set of probabilities {P j} is representative. In one embodiment, the evaluation metric may be calculated as a product (which is calculated for the number of probabilities (e.g., N - 1 or N) to account for prompts 101 with variable lengths. M=N∏j=1N−1Pj,

[0050] The presence of unexpected token-to-token transitions is captured by the evaluation metric M, since one or more unusually small probabilities make the overall evaluation metric M low.

[0051] In some embodiments, the evaluation metric M may be defined by a minimum probability of the distribution of probabilities, M=min(P1,P2,…,PN−1).

[0052] In other embodiments, the evaluation metric M may be calculated using a different function for the probabilities {P j} be calculated, e.g. an arithmetic mean, a harmonic mean, an average or product of a predetermined number of the lowest probabilities P j and / or the like.

[0053] In some embodiments, verification prompts 310 may require the LM 124 to verify probabilities of multiple (two or more) tokens, e.g., tokens Tj +1and T j+2 after a set of given tokens T1, T2, ... T j , to predict.

[0054] In the above autoregressive prompt generation scheme, the LM 124 is given the task of calculating a probability of the next token T j+1, that corresponds to a set of given tokens T1, T2, ... T j In some embodiments, a fill-mask prompt generation scheme may be used instead of or in addition to the autoregressive schedule. In particular, the verification prompt generator 210 may instruct the LM 124 to calculate a probability of a missing token T m tokens issued in a sequence, T1, ..., T m-1 , T m+1 ... T j , to predict. The corresponding probability can be determined from a corresponding (to the missing token T m ) neuron of the output layer of LM 122.

[0055] In some embodiments, the evaluation metric M may be generated by a trained prompt evaluation machine learning model (MLM) employed as part of the prompt evaluation 220. In particular, an input to the prompt evaluation MLM may be a distribution of token probabilities {P j} and an output M can indicate a degree of suitability of the LM 124 to process the prompt. {Pj}→MLM→M.

[0056] During training of the prompt evaluation MLM, a training prompt may be selected, e.g., generated by a developer or obtained from a public database of user-submitted queries. A set of verification prompts 310 may then be generated as disclosed above and used to determine a distribution of token probabilities {P j} from the LM 124. The training prompt may also be processed by the LM 124, and a human developer may assign a metric M to a response generated by the LM 124, e.g., based on factual accuracy, topic suitability, conformity to various public and / or private policies, and / or the like. The distribution of token probabilities {P j} can then be used as training input to the prompt evaluation MLM and the metric M can be used as ground truth for training the prompt evaluation MLM.

[0057] The prompt evaluation 220 may compare the evaluation metric M with a threshold metric M T compare. The threshold metric M T can be determined empirically, e.g., as part of testing the prompt analyzer 120. Based on a result of the comparison, the prompt evaluation 220 can determine whether the prompt 101 is valid, e.g., whether M ≥ M T (or M > M T), or invalid, e.g., whether M < M T (or M ≤ M T). A low metric M indicates that the LM 124 is not capable of generating an accurate and reliable response to the prompt 101, or that the prompt is not processable by the LM 124 (e.g., for policy reasons). In such situations, the prompt evaluation 220 may determine that the prompt 101 is an invalid prompt and return the prompt 101 to the UI 106, e.g., along with a recommendation to express the prompt differently, an explanation that the prompt cannot be processed in its current form, and / or the like. A high metric M indicates that the LM 124 is likely to generate an acceptable response. In such situations, the prompt evaluation 220 may determine that the prompt 101 is valid and provide the prompt 101 to the LM 124. After processing the prompt 101, the LM 124 may return a generated response 330 to a user via the UI 106.

[0058] Fig. 4 illustrates an architecture 400 of an exemplary system, according to at least one embodiment, that uses prompt analysis to select one of multiple LMs for prompt processing. Various blocks of Fig. 4, which have the same reference numerals as the respective blocks of Fig. 3 can implement the same (or similar) functionality. As described in Fig. 4, several trained LMs, e.g., LM 424-1, LM 424-2, LM 424-3 and / or the like (three LMs are shown in Fig. 4 for illustration purposes, however, any number of LMs may be employed), verification prompts 310 generated by the verification prompt generator 210 may be provided. Different LMs 424-k may be models trained using different specialized data, e.g., LM 424-1 may be a model trained using medical data, LM 424-2 may be a model trained using financial information, LM 424-3 may be a model trained using mathematical texts, and / or the like. Individual LMs 424-k may return separate distributions of token probabilities that may be used by the prompt evaluator 220 to determine separate evaluation metrics M kfor separate LMs. Prompt evaluation 220 may then select a model with the highest evaluation metric and further conditionally, that the highest evaluation metric be above (or at) the threshold metric M T is located, identify prompt 101 as a valid prompt, and provide prompt 101 to the model (which is expected to be the optimal model for handling the prompt). Fig. 4 illustrates a situation in which the prompt 101 is provided to the LM 424-2. After processing the prompt 101 and generating a response 430, the model (e.g., LM 424-2) can deliver the response 430 to a user via the UI 106. If no evaluation metric M k at or above the threshold metric M T it can be determined that the prompt 101 is invalid and it can be returned to the user via the UI 106, as described in connection with Fig. 3 described.

[0059] Fig. 5 illustrates a data flow of example operations 500 of an LM with prompt analysis processing according to at least one embodiment. At block 510, the operations 500 may include receiving a prompt, e.g., from a human user or machine user. The prompt may be in a natural language. At block 520, the prompt may be tokenized by representing words of the prompt via tokens of a set of natural language tokens used by an LM. At block 530, a prompt analyzer may generate a verification prompt and, at block 540, provide the verification prompt to the LM, where a probability for a next or missing token is requested from the LM. At block 550, the prompt analyzer may receive the requested probability. The operation of blocks 530-550 may be repeated multiple times to obtain a distribution of such probabilities for the prompt.At block 560, the next / missing token probabilities may be used to compute one or more evaluation metrics for the prompt. The evaluation metrics may be computed for one or more LMs available for prompt processing. At decision-making block 565, the prompt analyzer may determine whether the evaluation metric is met for at least one of the available LMs, e.g., by comparing the computed evaluation metrics to a threshold metric. If no evaluation metric matches or exceeds the threshold metric, operations 500 may continue to block 570, where an appropriate prompt rejection protocol consistent with the policy of the language model services is implemented.For example, the protocol may include a notification to the user that the prompt cannot be processed, a recommendation to the user to express the prompt differently, and / or the like. If at least one evaluation metric matches or exceeds the threshold metric, the prompt analyzer may, at block 580, select a target LM to process the prompt. For example, the LM with the highest evaluation metric may be selected as the target LM. At block 590, the prompt may be provided to the target LM. After the LM processes the prompt and generates a response, the response may be provided to the user.

[0060] Fig. 6 is a flow diagram of an exemplary method 600, according to at least one embodiment, that performs a suitability analysis of prompts prior to LM processing for improved quality and security of LM outputs. The method 600 may be performed using one or more processing units (e.g., CPUs, GPUs, accelerators, PPUs, DPUs, etc.) that may include (or communicate with) one or more memory devices. In at least one embodiment, the method 600 may be performed using processing units of the computing device 102 of Fig. 1 and / or computing device 200 of Fig. 2. In at least one embodiment, the processing units performing method 600 may execute instructions stored on a non-transient computer-readable storage medium. In at least one embodiment, method 600 may be performed using multiple processing threads (e.g., CPU threads and / or GPU threads), with individual threads executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, processing threads implementing method 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, processing threads implementing method 600 may be executed asynchronously with respect to one another. Various operations of method 600 may be compared to the Fig. 6 in a different order. Some operations of method 600 may be performed concurrently with other operations. In at least one embodiment, one or more operations in Fig. 6 operations shown are not always performed.

[0061] The method 600 may include prompts produced by entering (e.g., typing) the text of a prompt, speaking the words of the prompt, e.g., as part of a question, a request, a conversation, a dialogue, and / or other suitable interaction of a human user with an AI system, in an example embodiment.

[0062] At block 610, the method 600 may include obtaining a plurality of tokens associated with a prompt (e.g., by processing the prompt 101 using the tokenizer 302 of Fig. 3). The prompt may include text (e.g., a digital representation of letters, syllables, words, phrases, sentences, and / or the like), speech (e.g., a digital audio representation of spoken words), video (e.g., a sequence of video frames), or other plurality of images (e.g., coherent through context rather than representing a temporal video sequence).

[0063] At blocks 620-630, method 600 may include determining one or more prompt verification scores to evaluate the suitability of the prompt for LM processing. In particular, at block 620, method 600 may include generating a verification prompt. The verification prompt may include a first subset of the plurality of tokens and a second subset of the plurality of tokens. The first subset may include one or more tokens to be provided to the LM as unconventional (known) tokens. The second subset may include one or more tokens to be provided to the LM as conditional tokens, the occurrence or non-occurrence of which, along with the first subset, may be probabilistic.In some embodiments, the second subset of tokens may include a token that follows the first subset of tokens or may include a token that occurs within the first subset of tokens. In some embodiments, the second subset may include a next token (e.g., token T). j+1 ) corresponding to the first subset of tokens (e.g., tokens T1 ... T j ) and exclude a token that follows the next token (e.g., token T j+2 exclude).

[0064] At block 630, the method 600 includes obtaining, using a first LM, an individual prompt verification score (e.g., P j ), which indicates a probability that the second subset of one or more tokens in the prompt occurs together with the first subset of tokens (e.g., that token T j+1 Tokens T1 ... T j follows).

[0065] In some embodiments, operations of blocks 620 and 630 may be performed over multiple iterations (e.g., N-1 iterations or another number of iterations). In one illustrative example, a pair of consecutive iterations may be performed as follows: each iteration—first or second—including a respective (first or second) verification prompt. The terms “first” and / or “second” should be understood merely as identifiers and do not imply that the first verification prompt is chronologically the earliest verification prompt generated and / or processed. In particular, the first subset of tokens of a first verification prompt may include first k tokens (e.g., T1 ... T k) of the plurality of tokens, where k is an integer less than N-1, where N is the total number of tokens in the prompt. The second subset of the first verification prompt may include the k+1st token (e.g., token T k+1 ) of the prompt (but may also include other tokens). Similarly, the first subset of the second verification prompt may contain the first k+1 tokens (e.g., T1 ... T k+1 ) of the prompt and the second subset of tokens of the second verification prompt may include the k+2nd token (e.g., T k+2 ) of the prompt (but can also include other tokens).

[0066] In some embodiments, operations of blocks 620-630 may also be performed for additional LMs (e.g., a second, third, etc.), e.g., to determine one or more additional prompt verification scores. Determining an individual additional prompt verification score of the one or more additional prompt verification scores may include obtaining, using a second (third, etc.) LM, the individual prompt verification score that indicates a probability that the second subset of one or more tokens occurs in the prompt along with the first subset of tokens. In such embodiments, e.g., as described below in connection with blocks 648 and 652, the one or more additional prompt verification scores may be used to determine whether to provide the prompt to the first (second, etc.) LM.

[0067] At block 640, the method 600 may continue with determining, using the one or more prompt verification scores (and in some embodiments, the additional prompt verification scores), whether to provide the prompt to the first LM. In some embodiments, operations of block 640 may include operations associated with highlighting Fig. 6. In particular, at block 642, the method 600 may include calculating, using the one or more prompt verification scores (e.g., P1 ... P N-1), an evaluation metric for the prompt. In some embodiments, calculating the evaluation metric for the prompt may include aggregating (e.g., calculating a product, a sum, and / or some other suitable function) the one or more prompt verification scores to obtain the evaluation metric. In some embodiments, method 600 may include calculating, using the one or more prompt verification scores, a first evaluation metric (e.g., M1) for the prompt (indicating suitability of the prompt for the first LM), and may further include calculating, using the one or more additional prompt verification scores, a second (third, etc.) evaluation metric (e.g., M2, M3, etc.) for the prompt (indicating suitability of the prompt for the second LM, third LM, etc.).At block 644, method 600 may continue by comparing the calculated evaluation metric to a threshold metric. In some embodiments, multiple evaluation metrics (calculated for multiple LMs) may be compared to the threshold metric.

[0068] In some situations, as indicated by block 646, to determine whether to provide the prompt to the first LM, the processing units performing method 600 may determine that the evaluation metric (e.g., M) is below the threshold metric (e.g., M T ) (or that each evaluation metric calculated for individual available LMs is less than M T At block 650, the method 600 may continue by generating a response (e.g., to the user who produced the prompt) that includes a request to modify the prompt or a notification that the prompt cannot be processed.

[0069] In other situations, as indicated by block 648, to determine whether to provide the prompt to the first LM, the processing units performing method 600 may determine that the evaluation metric is above the threshold metric (e.g., that M1 > M T ), In such situations, the method 600 may continue with providing the prompt to the first LM. In some embodiments, when two or more LMs are available to process the prompt, the method 600 may include providing the prompt to the LM responsive to the first evaluation metric being above the second evaluation metric (M1 > M2). In those situations where the first evaluation metric is below the second evaluation metric (M1 < M2), the method 600 may include providing the prompt to the second LM.

[0070] Furthermore, the systems and methods described herein may be used for a number of purposes, for example, but not limited to, performing one or more operations related to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.

[0071] Disclosed embodiments may be included in a number of different systems, such as automotive systems (e.g., an in-vehicle infotainment system for an autonomous or semi-autonomous machine), systems implemented using a robot, antenna systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twinning operations, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing a light transport simulation,Systems for performing collaborative content creation for 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems. INFERENCE AND TRAINING LOGIC

[0072] Fig. 7A illustrates inference and / or training logic 715 used to perform inference and / or training operations associated with one or more embodiments.

[0073] In at least one embodiment, the inference and / or training logic 715 may include, among other things, code and / or data storage 701 to store feedforward and / or output weights and / or input / output data and / or other parameters for configuring neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 to store graph code or other software for controlling the timing and / or order in which weight and / or other parameter information is loaded for configuring logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs) or simply circuits).In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data from each layer of a neural network trained or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.

[0074] In at least one embodiment, any portion of code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 may be cache memory, dynamic randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, a choice of whether code and / or data storage 701 is, for example, internal or external to a processor or includes DRAM, SRAM, flash, or other storage type may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.

[0075] In at least one embodiment, the inference and / or training logic 715 may include, among other things, code and / or data storage 705 to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 705 stores weight parameters and / or input / output data from each layer of a neural network trained or used in connection with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, the training logic 715 may include or be coupled to a code and / or data storage 705 for storing graph code or other software for controlling the timing and / or order in which weight and / or other parameter information for configuring logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)) is loaded.

[0076] In at least one embodiment, code, such as graph code, causes weight or other parameter information to be loaded into processor ALUs based on a neural network architecture to which such code conforms. In at least one embodiment, any portion of code and / or data storage 705 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 may be internal to or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, a choice of whether the code and / or data storage 705 is, for example, internal or external to a processor or includes DRAM, SRAM, flash, or other storage type may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.

[0077] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be a combined storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.

[0078] In at least one embodiment, the inference and / or training logic 715 may include, among other things, one or more arithmetic logic unit(s) ("ALU(s)") 710, including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or specified by training and / or inference code (e.g., graph code), a result of which may produce activations (e.g., output values ​​of layers or neurons within a neural network) stored in an activation storage 720 that are functions of input / output and / or weight parameter data stored in the code and / or data storage 701 and / or the code and / or data storage 705.In at least one embodiment, activations stored in activation storage 720 are generated according to linear algebraic and / or matrix-based mathematics performed by ALU(s) 710 in response to the execution of instructions or other code, using weight values ​​stored in code and / or data storage 705 and / or data storage 701 as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 705 or code and / or data storage 701 or other on- or off-chip storage.

[0079] In at least one embodiment, ALU(s) 710 are included within one or more processors or other hardware logic devices or circuits, while in another embodiment, ALU(s) 710 may be external to a processor or other hardware logic device or circuitry utilizing them (e.g., a coprocessor). In at least one embodiment, ALU(s) 710 may be included within execution units of a processor or otherwise within a series of ALUs accessible by execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and enable storage 720 may share a processor or other hardware logic device or circuit, while in another embodiment, they may reside in different processors or other hardware logic devices or circuits, or a combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of enable storage 720 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.Furthermore, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry, and may be retrieved and / or processed using fetch, decode, scheduling, execution, idle, and / or other logic circuitry of a processor.

[0080] In at least one embodiment, the activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 720 may be located entirely or partially within or external to one or more processors or other logic circuits. In at least one embodiment, a choice of whether the activation storage 720 is, for example, internal or external to a processor or comprises DRAM, SRAM, flash, or another storage type may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inferencing functions performed, the batch size of data used in inferring and / or training a neural network, or a combination of these factors.

[0081] In at least one embodiment, the Fig. 7A may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a Tensorflow® Processing Unit from Google, an Inference Processing Unit (IPU) from Graphcore™, or a Nervana® processor (e.g., “Lake-Crest” processor) from Intel Corp. In at least one embodiment, the inference and / or training logic 715 illustrated in Fig. 7A may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware, such as field programmable gate arrays (“FPGAs”).

[0082] Fig. 7B illustrates inference and / or training logic 715, according to at least one embodiment. In at least one embodiment, the inference and / or training logic 715 may include, among other things, hardware logic in which computational resources are dedicated or otherwise used exclusively in connection with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the inference and / or training logic 715 may Fig. 7B may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as Google’s Tensorflow® Processing Unit, a Graphcore™ Inference Processing Unit (IPU), or a Nervana® processor (e.g., “Lake-Crest” processor) from Intel Corp. In at least one embodiment, the inference and / or training logic 715 illustrated in Fig. 7B may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware, such as field programmable gate arrays (“FPGAs”). In at least one embodiment, the inference and / or training logic 715 includes, among other things, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment Fig. 7B, code and / or data storage 701 and code and / or data storage 705 are each associated with a dedicated computing resource, such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, computing hardware 702 and computing hardware 706 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in code and / or data storage 701 and code and / or data storage 705, respectively, the results of which are stored in activation storage 720.

[0083] In at least one embodiment, the code and / or data storage 701 and / or data storage 701 and the corresponding computing hardware 702 and 706 correspond to different layers of a neural network, such that a resulting activation from one storage / computing pair 701 / 702 of the code and / or data storage 701 and computing hardware 702 is provided as input to the next storage / computing pair 705 / 706 of the code and / or data storage 705 and computing hardware 706 to mirror the conceptual organization of a neural network. In at least one embodiment, each of the storage / computing pairs 701 / 702 and 705 / 706 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computing pairs (not shown) may be included after or in parallel with storage / computing pairs 701 / 702 and 705 / 706 in the inference and / or training logic 715. TRAINING AND DEPLOYING A NEURAL NETWORK

[0084] Fig. 8 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training dataset 802. In at least one embodiment, the training framework 804 is a PyTorch framework, while in other embodiments, the training framework 804 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 804 trains an untrained neural network 806 and allows it to be trained using processing resources described herein to generate a trained neural network 808. In at least one embodiment, weights may be selected arbitrarily or by pre-training using a deep belief network.In at least one embodiment, the training may be conducted in either a supervised, semi-supervised, or unsupervised manner.

[0085] In at least one embodiment, an untrained neural network 806 is trained using supervised learning, where a training data set 802 includes an input paired with a desired output, or where the training data set 802 includes an input having a known output, and an output of the neural network 806 is manually ranked. In at least one embodiment, an untrained neural network 806 is trained in a supervised manner and processes inputs from a training data set 802 and compares resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are thereafter backpropagated through the untrained neural network 806. In at least one embodiment, the training framework 804 adjusts weights that control the untrained neural network 806.In at least one embodiment, the training framework 804 includes tools to monitor how well the untrained neural network 806 converges to a model, such as a trained neural network 808, capable of generating correct answers, such as result 814, based on input data, such as a new data set 812. In at least one embodiment, the training framework 804 repeatedly trains an untrained neural network 806 while adjusting weights to refine an output of the untrained neural network 806 using a loss function and an adaptation algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 804 trains an untrained neural network 806 until the untrained neural network 806 achieves a desired accuracy.In at least one embodiment, the trained neural network 808 may then be used to implement any number of machine learning operations.

[0086] In at least one embodiment, an untrained neural network 806 is trained using unsupervised learning, while the untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, the training data set 802 for unsupervised learning will include input data without associated output data, or ground truth data. In at least one embodiment, an untrained neural network 806 may learn groupings within the training data set 802 and determine how individual inputs relate to the untrained data set 802. In at least one embodiment, unsupervised training may be used to generate a self-organizing mapping in the trained neural network 808, which may perform operations useful for reducing the dimensionality of a new data set 812.In at least one embodiment, unsupervised training may also be used to perform anomaly detection, allowing identification of data points in the new data set 812 that deviate from normal patterns of the new data set 812.

[0087] In at least one embodiment, semi-supervised learning may be used, which is a technique in which the training dataset 802 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 804 may be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning allows a trained neural network 808 to adapt to a new dataset 812 without forgetting knowledge introduced to the trained neural network 808 during initial training.

[0088] With reference to Fig. 9 is Fig. 9 illustrates an exemplary dataflow diagram for a process 900 for generating and deploying a processing and inference pipeline, according to at least one embodiment. In at least one embodiment, the process 900 may be deployed to perform game name recognition analysis and inference on user feedback data at one or more facilities 902, such as a data center.

[0089] In at least one embodiment, process 900 may be performed within a training system 904 and / or a deployment system 906. In at least one embodiment, training system 904 may be used to perform training, deployment, and execution of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in deployment system 906. In at least one embodiment, deployment system 906 may be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements at facility 902. In at least one embodiment, deployment system 906 may provide a simplified platform for selecting, customizing, and implementing virtual instruments for use with computing resources at facility 902.In at least one embodiment, virtual instruments may include software-defined applications for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in a pipeline may use or invoke services (e.g., inference, visualization, computation, AI, etc.) of the deployment system 906 during application execution.

[0090] In at least one embodiment, some applications used in advanced processing and inferencing pipelines may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models may be trained at device 902 using feedback data 908 (such as imaging data) stored at device 902, or feedback data 908 from another device or devices, or a combination thereof. In at least one embodiment, training system 904 may be used to provide applications, services, and / or other resources for generating working, deployable machine learning models for deployment system 906.

[0091] In at least one embodiment, the model registry 924 may be secured by object storage, which may support version and object metadata. In at least one embodiment, the object storage may be provided, for example, via cloud storage (e.g., a cloud 1026 of Fig. 10) be accessible in a cloud platform compatible with an application programming interface (API). In at least one embodiment, machine learning models may be uploaded, listed, modified, or deleted within the model registry 924 by developers or partners of a system interacting with an API. In at least one embodiment, an API may provide access to methods that allow users with appropriate permissions to associate models with applications so that models may be executed as part of an execution of containerized instantiations of applications.

[0092] In at least one embodiment, a training pipeline 1004 ( Fig. 10) may include a scenario where a device 902 is training its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, feedback data 908 may be received from various channels, such as forums, web forums, or the like. In at least one embodiment, after the feedback data 908 is received, AI-assisted annotation 910 may be used to assist in generating annotations corresponding to the feedback data 908 to be used as ground truth data for a machine learning model. In at least one embodiment, AI-assisted annotation 910 may include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that may be trained to generate annotations corresponding to certain types of feedback data 908 (e.g.,of certain devices) and / or certain types of anomalies in feedback data 908. In at least one embodiment, AI-assisted annotations 910 may then be used directly or adapted or fine-tuned using an annotation tool to generate ground truth data. In at least one embodiment, in some examples, labeled data 912 may be used as ground truth data for training a machine learning model. In at least one embodiment, AI-assisted annotations 910, labeled data 912, or a combination thereof may be used as ground truth data for training a machine learning model, e.g., via model training 914 in FIG. Fig. 9-10. In at least one embodiment, a trained machine learning model may be referred to as output model 916 and used by deployment system 906 as described herein.

[0093] In at least one embodiment, the training pipeline 1004 ( Fig. 10) include a scenario where a device 902 requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 906, but the device 902 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, an existing machine learning model may be selected from a model registry 924. In at least one embodiment, a model registry 924 may include machine learning models trained to perform a number of different inference tasks on imaging data. In at least one embodiment, machine learning models in the model registry 924 may be trained on imaging data from devices other than the device 902 (e.g.,remote facilities). In at least one embodiment, machine learning models may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a specific location, which may be a form of feedback data 908, the training may take place at that location, or at least in a manner that protects the confidentiality of imaging data or restricts the transfer of imaging data off-premises (e.g., to comply with HIPAA regulations, data privacy regulations, etc.). In at least one embodiment, after a model has been trained—or partially trained—at a location, a machine learning model may be added to the model registry 924.In at least one embodiment, a machine learning model may then be retrained or updated at any number of other facilities, and a retrained or updated model may be made available in model registry 924. In at least one embodiment, a machine learning model may then be selected from model registry 924—referred to as output model 916—and may then be used in deployment system 906 to perform one or more processing tasks for one or more deployment system applications.

[0094] In at least one embodiment, the training pipeline 1004 ( Fig. 10) may be used in a scenario involving a device 902 that requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 906, but the device 902 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, a machine learning model selected from the model registry 924 may not be fine-tuned or optimized for feedback data 908 generated at the device 902 due to differences in populations, genetic variations, robustness of training data used to train a machine learning model, diversity of irregularities in training data, and / or other issues with training data.In at least one embodiment, AI-assisted annotation 910 may be used to assist in generating annotations corresponding to feedback data 908 to be used as ground truth data for retraining or updating a machine learning model. In at least one embodiment, labeled data 912 may be used as ground truth data for training a machine learning model. In at least one embodiment, retraining or updating a machine learning model may be referred to as model training 914. In at least one embodiment, model training 914—e.g., AI-assisted annotations 910, labeled data 912, or a combination thereof—may be used as ground truth data for retraining or updating a machine learning model.

[0095] In at least one embodiment, deployment system 906 may include software 918, services 920, hardware 922, and / or other components, features, and functionality. In at least one embodiment, deployment system 906 may include a software "stack" such that software 918 may be built on top of services 920 and may use services 920 to perform some or all of the processing tasks, and services 920 and software 918 may be built on top of hardware 922 and may use hardware 922 to perform processing, storage, and / or other computational tasks of deployment system 906.

[0096] In at least one embodiment, software 918 may include any number of different containers, each container capable of executing an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks in an advanced processing and inference pipeline (e.g., inferencing, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, for each type of computing device, there may be any number of containers capable of performing a data processing task with respect to feedback data 908 (or other data types such as those described herein).In at least one embodiment, an advanced processing and inference pipeline may be defined based on selections of different containers desired or required for processing feedback data 908, in addition to containers that receive and configure imaging data for use by each container and / or for use by a device 902 after processing by a pipeline (e.g., for converting outputs back to a usable data type for storage and display at the device 902). In at least one embodiment, a combination of containers within software 918 (e.g., constituting a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and a virtual instrument may utilize services 920 and hardware 922 to perform some or all of the processing tasks of applications instantiated in containers.

[0097] In at least one embodiment, data may be subjected to preprocessing as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, postprocessing may be performed on an output from one or more inference tasks or other processing tasks of a pipeline to prepare output data for a next application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks may be performed by one or more machine learning models, such as trained or deployed neural networks, which may include output models 916 of the training system 904.

[0098] In at least one embodiment, data processing pipeline tasks may be encapsulated in one or more containers, each representing a discrete, fully functional instantiation of an application and virtualized computing environment capable of referencing machine learning models. In at least one embodiment, containers or applications may be published in a private area (e.g., limited access area) by a container registry (described in more detail herein), and trained or deployed models may be stored in the model registry 924 and associated with one or more applications. In at least one embodiment, images of applications (e.g.,Container images) must be available in a container registry, and after being selected by a user from a container registry for use in a pipeline, an image can be used to generate a container for an instantiation of an application for use by a user system.

[0099] In at least one embodiment, developers may develop, publish, and store applications (e.g., as containers) for performing processing and / or inferencing on supplied data. In at least one embodiment, the development, publishing, and / or storage may be performed using a software development kit (SDK) associated with a system (e.g., to ensure that a developed application and / or container is consistent or compatible with a system). In at least one embodiment, an application being developed with an SDK that can support at least some services 920 may be deployed locally (e.g., at a first device, to data from a first device) as a system (e.g., system 1000 of Fig. 10). In at least one embodiment, after being validated by system 1000 (e.g., for accuracy, etc.), an application may be available in a container registry for selection and / or execution by a user (e.g., a hospital, clinic, laboratory, healthcare provider, etc.) to perform one or more processing tasks on data at a user's facility (e.g., a second facility).

[0100] In at least one embodiment, developers can then deploy applications or containers over a network for access and use by users of a system (e.g., System 1000 of Fig. 10). In at least one embodiment, completed and validated applications or containers may be stored in a container registry, and associated machine learning models may be stored in the model registry 924. In at least one embodiment, a requesting entity providing an inference or image processing request may search a container registry and / or model registry 924 for an application, container, dataset, machine learning model, etc., select a desired combination of elements for inclusion in a data processing pipeline, and submit a processing request. In at least one embodiment, a request may include input data necessary to perform a request and / or may include a selection of applications and / or machine learning models to be executed in processing a request.In at least one embodiment, a request may be forwarded to one or more components of the deployment system 906 (e.g., a cloud) to perform processing of the data processing pipeline. In at least one embodiment, processing by the deployment system 906 may include referencing selected elements (e.g., applications, containers, models, etc.) from a container registry and / or model registry 924. In at least one embodiment, after results are generated by a pipeline, results may be returned to a user for reference (e.g., for viewing in a viewing application suite executing on a local on-premises workstation or terminal).

[0101] In at least one embodiment, services 920 may be used to support the processing or execution of applications or containers in pipelines. In at least one embodiment, services 920 may include computing services, collaborative content creation services, simulation services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, services 920 may provide functionality common to one or more applications in software 918, so that functionality may be abstracted to a service that may be invoked or consumed by applications. In at least one embodiment, functionality provided by services 920 may run more dynamically and efficiently while also scaling well by allowing applications to process data in parallel, e.g., using a parallel computing platform 1030 ( Fig. 10). In at least one embodiment, a service 920 may be shared between and among different applications, rather than each application sharing the same functionality offered by a service 920 requiring a respective instance of the service 920. In at least one embodiment, services may include an inference server or an inference engine, which, as non-limiting examples, may be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included, which may provide capabilities for machine learning model training and / or retraining.

[0102] In at least one embodiment, where a service 920 includes an AI service (e.g., an inference service), one or more machine learning models associated with an application for anomaly detection (e.g., tumors, growth abnormalities, scarring, etc.) may be executed by invoking (e.g., as an API call) an inference service (e.g., an inference server) to execute one or more machine learning models, or processing thereof, as part of application execution. In at least one embodiment, where another application includes one or more machine learning models for segmentation tasks, an application may invoke an inference service to execute machine learning models to perform one or more processing operations associated with segmentation tasks.In at least one embodiment, software 918 implementing an advanced processing and inference pipeline may be simplified because each application may invoke the same inference service to perform one or more inference tasks.

[0103] In at least one embodiment, hardware 922 may include GPUs, CPUs, graphics cards, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX™ supercomputer system), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 922 may be used to provide efficient, dedicated support for software 918 and services 920 in the deployment system 906. In at least one embodiment, the use of GPU processing may be implemented for local processing (e.g., at device 902), within an AI / deep learning system, in a cloud system, and / or in other processing components of the deployment system 906 to improve the efficiency, accuracy, and effectiveness of game name recognition.

[0104] In at least one embodiment, software 918 and / or services 920 may be optimized for GPU processing related to deep learning, machine learning, and / or high-performance computing, simulation, and visual computing, as non-limiting examples. In at least one embodiment, at least a portion of the computing environment of deployment system 906 and / or training system 904 may be executed in a data center or one or more supercomputers or high-performance computing systems with GPU-optimized software (e.g., hardware and software combination of NVIDIA's DGX™ system). In at least one embodiment, hardware 922 may include any number of GPUs that may be invoked to perform the processing of data in parallel, as described herein.In at least one embodiment, a cloud platform may further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other compute tasks. In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC™) may be executed using one or more AI / deep learning supercomputers and / or GPU-optimized software (e.g., as deployed on NVIDIA's DGX™ systems) as a hardware abstraction and scaling platform. In at least one embodiment, the cloud platform may integrate an application container clustering system or orchestration system (e.g., KUBERNETES) across multiple GPUs to enable seamless scaling and load balancing.

[0105] Fig. 10 is a system diagram for an exemplary system 1000 for generating and deploying a deployment pipeline, according to at least one embodiment. In at least one embodiment, the system 1000 may be used to implement the process 900 of Fig. 9 and / or other processes including advanced processing and inference pipelines. In at least one embodiment, system 1000 may include a training system 904 and deployment system 906. In at least one embodiment, training system 904 and deployment system 906 may be implemented using software 918, services 920, and / or hardware 922 as described herein.

[0106] In at least one embodiment, system 1000 (e.g., training system 904 and / or deployment system 906) may be implemented in a cloud computing environment (e.g., using cloud 1026). In at least one embodiment, system 1000 may be implemented locally with respect to a facility or as a combination of cloud and local computing resources. In at least one embodiment, access to APIs in cloud 1026 may be restricted to authorized users through enforced security measures or protocols. In at least one embodiment, a security protocol may include web tokens that may be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and may include appropriate authorization.In at least one embodiment, virtual instrument APIs (described herein), or other instantiations of system 1000, may be restricted to a set of public IPs that have been vetted or authorized for interaction.

[0107] In at least one embodiment, various components of system 1000 may communicate between themselves and with each other using a variety of different network types, including, but not limited to, local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communications between devices and components of system 1000 (e.g., for transmitting inference requests, for receiving results of inference requests, etc.) may be communicated via a data bus or via data buses, wireless data protocols (WiFi), wired data protocols (e.g., Ethernet), etc.

[0108] In at least one embodiment, the training system 904 may include training pipelines 1004 similar to those described herein with respect to Fig. 9. In at least one embodiment, training pipelines 1004 may be used when one or more machine learning models in deployment pipelines 1010 are to be used by deployment system 906 to train or retrain one or more (e.g., pre-trained) models and / or implement one or more of the pre-trained models 1006 (e.g., without the need for retraining or updating). In at least one embodiment, output model(s) 916 may be generated as a result of training pipelines 1004.In at least one embodiment, the training pipelines 1004 may include any number of processing steps, AI-assisted annotation 910, labeling or annotating feedback data 908 to generate labeled data 912, model selection from a model registry, model training 914, training, retraining, or updating models, and / or other processing steps. In at least one embodiment, different training pipelines 1004 may be used for different machine learning models used by the deployment system 906. In at least one embodiment, a training pipeline 1004 similar to a first one described with respect to FIG. Fig. 9, for a first machine learning model, a training pipeline 1004, similar to a second one with respect to Fig. 9, for a second machine learning model, and may include a training pipeline 1004 similar to a third in terms of Fig. 9, for a third machine learning model. In at least one embodiment, any combination of tasks may be used within a training system 904 depending on what is required for each machine learning model. In at least one embodiment, one or more machine learning models may already be trained and ready for deployment, such that machine learning models may not be subjected to any processing by the training system 904 and may be implemented by the deployment system 906.

[0109] In at least one embodiment, the output model(s) 916 and / or pre-trained model(s) 1006 may include any type of machine learning model, depending on the embodiment. In at least one embodiment, machine learning models used by system 1000 may include, but are not limited to, machine learning model(s) using linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., auto-encoders, convolutional, recurrent, perceptrons, long / short-term / memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machine, etc.), and / or other types of machine learning models.

[0110] In at least one embodiment, training pipelines 1004 may include AI-assisted annotation. In at least one embodiment, labeled data 912 (e.g., traditional annotation) may be generated through a variety of techniques. In at least one embodiment, labels or other annotations may be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of program suitable for generating annotations or labels for ground truth, and / or, in some examples, hand-drawn. In at least one embodiment, ground truth data may be produced synthetically (e.g., generated from computer models or renderings), real (e.g., designed and produced from real-world data), machine-automated (e.g.,using feature analysis and learning to extract features from data and subsequently generate labels), annotated by a human (e.g., labeler, or annotation expert, defines the location of the labels), and / or a combination thereof. In at least one embodiment, for each instance of feedback data 908 (or type of data used by machine learning models), there may be corresponding ground truth data generated by the training system 904. In at least one embodiment, AI-assisted annotation may be performed as part of deployment pipelines 1010; either in addition to or instead of AI-assisted annotation included in training pipelines 1004. In at least one embodiment, the system 1000 may include a multi-layer platform that includes a software layer (e.g.,Software 918) for diagnostic applications (or other types of applications) that can perform one or more medical imaging and diagnostic functions.

[0111] In at least one embodiment, a software layer may be implemented as a secure, encrypted, and / or authenticated API through which applications or containers may be invoked (e.g., invoked) from external environments (e.g., device 902). In at least one embodiment, applications may then invoke or execute one or more services 920 to perform computational, AI, or visualization tasks associated with respective applications, and software 918 and / or services 920 may utilize hardware 922 to perform processing tasks in an effective and efficient manner.

[0112] In at least one embodiment, deployment system 906 may execute deployment pipelines 1010. In at least one embodiment, deployment pipelines 1010 may include any number of applications that may be applied sequentially, non-sequentially, or otherwise to feedback data (and / or other data types), including AI-assisted annotation as described above. In at least one embodiment, a deployment pipeline 1010, as described herein, for an individual device may be referred to as a virtual instrument for a device. In at least one embodiment, there may be more than one deployment pipeline 1010 for a single device, depending on the information desired from data generated by a device.

[0113] In at least one embodiment, applications available to deployment pipelines 1010 may include any application that may be used to perform processing tasks on feedback data or other data from devices. In at least one embodiment, since different applications share common image operations in some embodiments, a data extension library (e.g., as one of services 920) may be used to accelerate these operations. In at least one embodiment, to avoid bottlenecks of traditional processing approaches that rely on CPU processing, a parallel computing platform 1030 may be used for GPU acceleration of these processing tasks.

[0114] In at least one embodiment, the deployment system 906 may include a user interface ("UI") 1014 (e.g., a graphical user interface, a web interface, etc.) that may be used to select applications for inclusion in deployment pipelines 1010, arrange applications, modify or change applications or parameters or constructs thereof, use and interact with deployment pipelines 1010 during setup and / or deployment, and / or otherwise interact with the deployment system 906. In at least one embodiment, the UI 1014 (or other user interface), although not illustrated with respect to the training system 904, may be used to select models for use in the deployment system 906, to select models for training, or retraining, in the training system 904, and / or to otherwise interact with the training system 904.In at least one embodiment, the training system 904 and the deployment system 906 may include DICOM adapters 1002A and 1002B.

[0115] In at least one embodiment, a pipeline manager 1012 may be used, in addition to an application orchestration system 1028, to manage the interaction between applications or containers of deployment pipelines 1010 and services 920 and / or hardware 922. In at least one embodiment, the pipeline manager 1012 may be configured to enable application-to-application, application-to-service 920, and / or application or service-to-hardware 922 interactions. In at least one embodiment, although illustrated as being included in software 918, which is not to be construed as limiting, in some examples, the pipeline manager 1012 may be included in services 920. In at least one embodiment, the application orchestration system 1028 (e.g., Kubernetes, DOCKER, etc.)) may include a container orchestration system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from deployment pipelines 1010 (e.g., a reconstruction application, a segmentation application, etc.) with individual containers, each application can be executed in a self-contained environment (e.g., at a kernel level) to increase speed and efficiency.

[0116] In at least one embodiment, each application and / or container (or image thereof) may be individually developed, modified, and deployed (e.g., a first user or developer may develop, modify, and deploy a first application, and a second user or developer may develop, modify, and deploy a second application separately from a first user or developer), which may allow focus and attention to be directed to a task from a single application and / or container(s) without being hindered by tasks from other applications or containers. In at least one embodiment, communication and cooperation between different containers or applications may be supported by the pipeline manager 1012 and the application orchestration system 1028.In at least one embodiment, as long as an expected input and / or output from each container or application is known to a system (e.g., based on application or container constructs), the application orchestration system 1028 and / or the pipeline manager 1012 may facilitate communication among and between, and sharing of resources among, each of the applications or containers. In at least one embodiment, since one or more of the applications or containers in deployment pipelines 1010 may share the same services and resources, the application orchestration system 1028 may orchestrate, load balance, and determine the sharing of services or resources between and among different applications or containers.In at least one embodiment, a scheduler may be used to track resource requests from applications or containers, current or planned usage of those resources, and resource availability. Therefore, in at least one embodiment, the scheduler may allocate resources to different applications and distribute resources between and among applications with respect to the needs and availability of a system. In some examples, the scheduler (and / or another component of the application orchestration system 1028) may determine resource availability and distribution based on constraints imposed on a system (e.g., user constraints), such as quality of service (QoS), urgency of the need for data outputs (e.g., to determine whether to perform real-time or deferred processing), etc.

[0117] In at least one embodiment, services 920 used and shared by applications or containers in the deployment system 906 may include compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, and / or other service types. In at least one embodiment, applications may invoke (e.g., execute) one or more services 920 to perform processing operations for an application. In at least one embodiment, compute services 1016 may be used by applications to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, compute service(s) 1016 may be used to perform parallel processing (e.g.,using a parallel computing platform 1030) to process data across one or more applications and / or perform one or more tasks of a single application substantially simultaneously. In at least one embodiment, the parallel computing platform 1030 (e.g., NVIDIA's CUDA) may enable general-purpose computing on GPUs (GPGPU) (e.g., GPUs 1022). In at least one embodiment, a software layer of the parallel computing platform 1030 may provide access to virtual instruction sets and parallel computing elements of GPUs for executing compute kernels. In at least one embodiment, the parallel computing platform 1030 may include memory, and in some embodiments, memory may be shared between and among multiple containers and / or between and among different processing tasks within a single container.In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or for multiple processes within a container to use the same data from a shared segment of the memory of the parallel computing platform 1030 (e.g., when multiple different stages of an application or multiple applications process the same information). In at least one embodiment, instead of making a copy of data and moving data to different locations in memory (e.g., a read / write operation), the same data in the same memory location may be used for any number of processing tasks (e.g., at the same time, at different times, etc.).In at least one embodiment, when data is used to generate new data as a result of processing, this information can be stored about a new location and shared between different applications. In at least one embodiment, the location of data and a location of updated or modified data can be part of a definition of how a payload is understood within containers.

[0118] In at least one embodiment, AI services 1018 may be utilized to perform inference services for executing machine learning models associated with applications (e.g., tasked with performing one or more processing tasks of an application). In at least one embodiment, AI services 1018 may utilize AI system 1024 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, applications of deployment pipelines 1010 may utilize one or more output models 916 from training system 904 and / or other models of applications to perform inference on imaging data (e.g., DICOM data, RIS data, CIS data, RESTful data, RPC data, raw data, etc.).In at least one embodiment, two or more examples of inferencing using application orchestration system 1028 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can achieve higher service level agreements, such as performing inference on urgent requests during emergencies or for a radiologist during diagnosis. In at least one embodiment, a second category may include a default priority path that can be used for requests that may not be urgent or when analysis can be performed at a later time. In at least one embodiment, application orchestration system 1028 may distribute resources (e.g., services 920 and / or hardware 922) based on priority paths for different inference tasks of AI services 1018.

[0119] In at least one embodiment, shared storage may be connected to AI services 1018 within system 1000. In at least one embodiment, shared storage may operate as a cache (or other storage device type) and be used to process inference requests from applications. In at least one embodiment, when an inference request is sent, a request may be received by a set of API instances of deployment system 906, and one or more instances may be selected (e.g., for best fit, for load balancing, etc.) to process a request.In at least one embodiment, to process a request, a request may be entered into a database, a machine learning model may be found from a model registry 924 if not already in a cache, a validation step may ensure that the appropriate machine learning model is loaded into a cache (e.g., shared storage), and / or a copy of a model may be stored in a cache. In at least one embodiment, the scheduler (e.g., a pipeline manager 1012) may be used to launch an application referenced in a request if no application is already running or if there are not enough instances of an application. In at least one embodiment, an inference server may be launched if no inference server has already been launched to execute a model. In at least one embodiment, a number of inference servers may be launched per model.In at least one embodiment, in a pull model where inference servers are clustered, models may be cached when load balancing is advantageous. In at least one embodiment, inference servers may be statically loaded in corresponding distributed servers.

[0120] In at least one embodiment, inference may be performed using an inference server running in a container. In at least one embodiment, an instance of an inference server may be associated with a model (and optionally, a plurality of versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance may be loaded. In at least one embodiment, a model may be submitted to an inference server upon startup, so that a same container may be used to serve different models as long as the inference server runs as a different instance.

[0121] In at least one embodiment, during application execution, an inference request for a given application may be received, and a container (e.g., hosting an instance of an inference server) may be loaded (if not already loaded), and a startup operation may be invoked. In at least one embodiment, preprocessing logic in a container may perform loading, decoding, and / or other additional preprocessing on incoming data (e.g., using a CPU(s) and / or GPU(s)). In at least one embodiment, after data has been prepared for inference, a container may perform inference on data as needed. In at least one embodiment, this may include a single inference call on an image (e.g., a hand x-ray) or may require inference on hundreds of images (e.g., a chest CT).In at least one embodiment, an application may summarize results prior to completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generating a visualization, or generating text to summarize findings. In at least one embodiment, different models or applications may be assigned different priorities. For example, some models may have a real-time (turnaround time less than one minute) priority, while others may have a lower priority (e.g., turnaround less than 10 minutes). In at least one embodiment, model execution times may be measured by a requesting institution or entity and may include partner network traversal time as well as execution to an inference service.

[0122] In at least one embodiment, the transfer of requests between services 920 and inference applications may be hidden behind a software development kit (SDK), and robust transport may be provided by a queue. In at least one embodiment, a request is placed in a queue via an API for a single application / tenant ID combination, and an SDK fetches a request from a queue and passes a request to an application. In at least one embodiment, a queue name may be provided in an environment from which an SDK picks up the request. In at least one embodiment, asynchronous communication through a queue may be useful because this may allow each instance of an application to pick up work as it becomes available.In at least one embodiment, results may be transferred back via a queue to ensure that no data is lost. In at least one embodiment, queues may also provide the ability to segment work, where highest priority work may go to a queue with the most instances of an application connected to it, while lowest priority work may go to a queue with a single instance connected to it, processing tasks in the order received. In at least one embodiment, an application may run on a GPU-accelerated instance generated in the cloud 1026, and an inference service may perform inference on a GPU.

[0123] In at least one embodiment, visualization services 1020 may be used to generate visualizations for viewing outputs from applications and / or deployment pipelines 1010. In at least one embodiment, GPUs 1022 may be used by visualization services 1020 to generate visualizations. In at least one embodiment, rendering effects, such as ray tracing or other light transport simulation techniques, may be implemented by visualization services 1020 to generate higher quality visualizations. In at least one embodiment, visualizations may include, among other things, 2D image renderings, 3D volume renderings, 3D volume reconstruction, 2D tomography slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, virtualized environments may be used to create a virtual interactive display or environment (e.g.,a virtual environment) for interaction by users of a system (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, visualization services 1020 may include an internal visualizer, cinematography, and / or other rendering or image processing capabilities or functionality (e.g., ray tracing, rasterization, internal optics, etc.).

[0124] In at least one embodiment, hardware 922 may include GPUs 1022, AI system 1024, cloud 1026, and / or other hardware used to execute training system 904 and / or deployment system 906. In at least one embodiment, GPUs 1022 (e.g., NVIDIA's TESLA ® and / or QUADRO ®GPUs) may include any number of GPUs used to perform processing tasks of compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, other services, and / or any features or functionality of software 918. For example, with respect to AI services 1018, GPUs 1022 may be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on outputs of machine learning models, and / or to perform inferencing (e.g., to execute machine learning models). In at least one embodiment, cloud 1026, AI system 1024, and / or other components of system 1000 may use GPUs 1022. In at least one embodiment, the cloud 1026 may include a GPU-optimized platform for deep learning tasks.In at least one embodiment, AI system 1024 may utilize GPUs, and cloud 1026—or at least a portion dedicated to deep learning or inference—may be executed using one or more AI systems 1024. As such, although hardware 922 is illustrated as discrete components, this is not to be construed as limiting, and any components of hardware 922 may be combined with or utilized by other components of hardware 922.

[0125] In at least one embodiment, the AI ​​system 1024 may include a dedicated computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, the AI ​​system 1024 (e.g., NVIDIA's DGX™) may include GPU-optimized software (e.g., a software stack) that may be executed using a plurality of GPUs 1022, in addition to CPUs, RAM, storage, and / or other components, features, or functionality. In at least one embodiment, one or more AI systems 1024 may be implemented in the cloud 1026 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1000.

[0126] In at least one embodiment, cloud 1026 may include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC™), which may provide a GPU-optimized platform for performing processing tasks of system 1000. In at least one embodiment, cloud 1026 may include AI system(s) 1024 for performing one or more AI-based tasks of system 1000 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1026 may be integrated with an application orchestration system 1028 utilizing multiple GPUs to enable seamless scaling and load balancing between and among applications and services 920. In at least one embodiment, the cloud 1026 may be provided for running at least some services 920 of the system 1000, including compute services 1016, AI services 1018, and / or visualization services 1020, as described herein.In at least one embodiment, the cloud 1026 may perform small and large batch inference (e.g., running NVIDIA's TensorRT™), an accelerated parallel computing API and platform 1030 (e.g., NVIDIA's CUDA. ® ), run an application orchestration system 1028 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques for producing higher quality cinematography), and / or provide other functionality for the system 1000.

[0127] In at least one embodiment, in an effort to preserve patient confidentiality (e.g., when patient data or records are to be used externally), cloud 1026 may include a registry, such as a deep learning container registry. In at least one embodiment, a registry may store containers for instantiations of applications that can perform pre-processing, post-processing, or other processing tasks on patient data. In at least one embodiment, cloud 1026 may receive data including both patient data and sensor data in containers, perform requested processing only on sensor data in those containers, and thereafter provide resulting output and / or visualizations to appropriate parties and / or devices (e.g.,on-premises medical devices used for visualization or diagnosis), all without having to extract, store, or otherwise access patient data. In at least one embodiment, the confidentiality of patient data is maintained in compliance with HIPAA and / or other data regulations.

[0128] Other variations are within the spirit of the present disclosure. Therefore, while the disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are illustrated in the drawings and have been described in detail above. It should be understood, however, that there is no intention to limit the disclosure to any specific form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents that are within the spirit and scope of the disclosure as defined in the appended claims.

[0129] The use of the terms "a" and "an" and "the" and reference designations in the context of describing the disclosed embodiments (particularly in the context of the claims below) are to be construed to cover both singular and plural, unless otherwise stated herein or clearly contradicted by context, and not as a definition of any term. The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (meaning "including, but not limited to") unless otherwise noted. "Connected," when unmodified and referring to physical connections, is to be construed as partially or wholly contained within, attached to, or joined with, even if something in between.The repetition of ranges of values ​​herein is intended as a quick way to individually refer to each separate value falling within the range, unless otherwise noted herein, and each separate value is incorporated into the specification as if individually recited herein. In at least one embodiment, unless otherwise noted or contradicted by context, the use of the term "set" (e.g., "a set of things") or "subset" is to be construed as a non-empty collection comprising one or more items. Furthermore, unless otherwise noted or contradicted by context, the term "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set; however, the subset and corresponding set may be the same.

[0130] Conjunctive expression, such as expressions of the form "at least one of A, B, and C" or "at least one of A, B, and C," unless specifically indicated otherwise or otherwise clearly contradicted by the context, is otherwise to be understood with the context as it is generally used to present that a thing, concept, etc., can be either A or B or C, or a non-empty subclause of the set of A and B and C. For example, in an illustrative example of a set having three elements, the conjunctive expressions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Such conjunctive language is therefore generally not intended to imply that certain embodiments require that at least one of A, at least one of B, and at least one of C be present.Furthermore, unless otherwise noted or contradicted by context, the term "plurality" indicates that a state is plural (e.g., "a plurality of things" indicates multiple things). In at least one embodiment, a number of things in a plurality is at least two, but may be more if indicated either explicitly or by context. Further, unless otherwise noted or clear from context, the phrase "based on" means "at least partially based on" and not "solely based on."

[0131] Operations of processes described herein may be performed in any suitable order, unless otherwise specified herein or otherwise clearly contradicted by context. In at least one embodiment, a process such as those described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed collectively on one or more processors, by hardware, or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmission) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within transceivers of transient signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media having stored thereon (or on other memory to store executable instructions) executable instructions that, when executed by one or more processors of a computer system (i.e., as a result of being executed), cause a computer system to perform the operations described herein.In at least one embodiment, a set of non-transitory computer-readable storage media comprises a plurality of non-transitory computer-readable storage media, and one or more individual non-transitory storage media of the plurality of non-transitory computer-readable storage media lack all code, while a plurality of non-transitory computer-readable storage media together store all code. In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors—for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions.In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of instructions.

[0132] Accordingly, in at least one embodiment, computing systems are configured to implement one or more services that alone or collectively perform operations of processes described herein, and such computing systems are configured with applicable hardware and / or software that enable operations to be performed. Further, a computing system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computing system that includes multiple devices that operate differently such that distributed computing systems perform operations described herein and such that a single device does not perform all operations.

[0133] The use of any example or exemplary language (e.g., "such as") provided herein is intended only to clarify embodiments of the disclosure and is not intended to limit the scope of the disclosure unless otherwise claimed. No language in the specification should be construed to imply that any unclaimed element is essential to the practice of the disclosure.

[0134] All references, including publications, patent applications, and patents cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically incorporated by reference and set forth in its entirety herein.

[0135] Throughout the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms may not be intended as synonyms. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but nevertheless cooperate or interact with each other.

[0136] Unless specifically stated otherwise, it may be recognized throughout the specification that terms such as "processing," "computing," "calculating," "determining," or the like, refer to operations and / or processes of a computer or computing system, or similar electronic computing device, that manipulate data represented as physical, such as electronic, quantities within the registers and / or memories of the computing system and / or transform data into other data similarly represented as physical quantities within the memories, registers, or other such information storage, transmission, or display devices of the computing system.

[0137] Similarly, the term "processor" may refer to a device or part of a device that processes electronic data from registers and / or memory and transforms that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, the "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, "software" processes may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, any process may refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently.In at least one embodiment, the terms "system" and "method" are used interchangeably herein in that a system may include one or more methods and methods may be considered a system.

[0138] References herein may be made to obtaining, acquiring, receiving, or inputting analog or digital data to a subsystem, computing system, or computer-implemented machine. In at least one embodiment, a process of obtaining, acquiring, receiving, or inputting analog and digital data may be implemented in various ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the processes of obtaining, acquiring, receiving, or inputting analog and digital data may be implemented by transferring data over a serial or parallel interface.In at least one embodiment, the processes of obtaining, capturing, receiving, or inputting analog or digital data may be implemented by transferring data over a computer network from a providing entity to a capturing entity. In at least one embodiment, references may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, processes for providing, outputting, transmitting, sending, or presenting analog or digital data may be implemented by transferring data as an input or output parameter of a function call, a parameter of an application programming interface, or an inter-process communication mechanism.

[0139] Although descriptions herein set forth exemplary embodiments of described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. Furthermore, although specific distributions of responsibilities may be defined above for the purpose of description, various functions and responsibilities may be distributed and divided in different ways depending on the circumstances.

[0140] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, specific features and acts are disclosed as exemplary forms for implementing the claims.

[0141] The disclosure of this application also includes the following numbered clauses: Clause 1. A procedure which includes: Obtaining a plurality of tokens associated with a prompt; Determining one or more prompt verification scores by at least, for an individual prompt verification score of the one or more prompt verification scores: Generating a verification prompt, comprising: a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), the individual prompt verification score, which indicates a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determine, using the one or more prompt verification scores, whether to provide the prompt to the LM. Clause 2. The method of clause 1, wherein the prompt comprises at least one of text, speech, video, or a plurality of images. Clause 3. The method of any preceding clause, wherein the second subset of one or more tokens comprises at least one or more of the following: a token that follows the first subset of one or more tokens, or a token that occurs within the first subset of one or more tokens. Clause 4. The method of any preceding clause, wherein the second subset of one or more tokens includes a next token following the first subset of one or more tokens and excludes a token following the next token. Clause 5. Procedure under any preceding clause, where: the first subset of one or more tokens of a first verification prompt comprises first k tokens of the plurality of tokens, where k is an integer less than N-1, where N is a number of the plurality of tokens, the second subset of one or more tokens of the first verification prompt comprises k+1th token of the plurality of tokens, the first subset of one or more tokens of a second verification prompt comprises k+1 tokens of the plurality of tokens and the second subset of one or more tokens of the second verification prompt comprises k+2th token of the plurality of tokens. Clause 6. The method of any preceding clause, wherein determining whether to provide the prompt to the LM comprises: Calculating, using the one or more prompt verification scores, an evaluation metric for the prompt; and Comparing the evaluation metric with a threshold metric. Clause 7. The method of Clause 6, wherein calculating the evaluation metric for the prompt comprises: Aggregating the one or more prompt verification scores to obtain the evaluation metric. Clause 8. The method of Clause 6 or 7, wherein determining whether to provide the prompt to the LM further comprises: Determining that the evaluation metric is above the threshold metric; and Providing the prompt for the LM. Clause 9. The method of any of clauses 6-8, wherein determining whether to provide the prompt to the LM further comprises: Determining that the evaluation metric is below the threshold metric; and the method further comprising: Generating a response comprising at least one of the following: a request to modify the prompt or a message that the prompt cannot be processed. Clause 10. Procedure under any preceding clause, further comprising: Determining one or more additional / supplemental prompt verification scores, wherein determining an individual additional prompt verification score of the one or more additional prompt verification scores comprises: Obtaining, using a second language model LM, the individual prompt verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and wherein determining whether to provide the prompt to the LM comprises: Using the one or more additional prompt verification scores. Clause 11. The method of Clause 10, wherein determining whether to provide the prompt to the LM further comprises: Calculating, using the one or more prompt verification scores, a first evaluation metric for the prompt; and Calculating, using the one or more additional prompt verification scores, a second evaluation metric for the prompt. Clause 12. The method of Clause 11, further comprising performing at least one of the following: Providing, in response to the first evaluation metric being above the second evaluation metric, the prompt for the LM or Provide, responsive to the first evaluation metric being below the second evaluation metric, the prompt for the second LM. Clause 13. A system comprising: one or more processing units for: Obtaining a plurality of tokens associated with a prompt; Determine at least one prompt verification score by at least: Generating a verification prompt comprising a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), at least one prompt verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determining, using the at least one prompt verification score, whether to provide the prompt to the LM or to present output generated using the prompt. Clause 14. A system according to Clause 13, wherein the second subset of one or more tokens comprises at least one of the following: a token that follows the first subset of one or more tokens, or a token that occurs within the first subset of one or more tokens. Clause 15. A system as claimed in clause 13 or 14, wherein, in order to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units are to: Calculating, using the at least one prompt verification score, an evaluation metric for the prompt; and Comparing the evaluation metric with a threshold metric. Clause 16. The system of Clause 15, wherein, to determine whether to provide the prompt to the LM or to present the output generated using the prompt, the one or more processing units are further operable to perform at least one of the following: Providing, responsive to determining that the evaluation metric is above the threshold metric, the prompt for the LM or Responsive to determining that the evaluation metric is below the threshold metric, generating a response comprising at least one of: (i) a request to modify the prompt or (ii) a notification that the prompt cannot be processed. Clause 17. The system of Clauses 13-16, wherein the one or more processing units further serve to: Determining one or more additional / supplementary prompt verification scores, wherein in order to determine an individual additional prompt verification score of the one or more additional prompt verification scores, the one or more processing units are to: Obtaining, using a second language model LM, the individual prompt verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and wherein, to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units are to: Using the one or more additional prompt verification scores. Clause 18. The system of Clause 17, wherein, in order to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units are further for: Calculating, using the one or more prompt verification scores, a first evaluation metric for the prompt; and Calculating, using the one or more additional prompt verification scores, a second evaluation metric for the prompt; and wherein the one or more processing units are further operable to perform at least one of the following: Providing, in response to the first evaluation metric being above the second evaluation metric, the prompt for the LM or Provide, responsive to the first evaluation metric being below the second evaluation metric, the prompt for the second LM. Clause 19. A system as defined in any of Clauses 13-18, which system consists of at least one of the following: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for carrying out one or more digital twinning operations; a system for performing a light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system that implements one or more large language models (LLMs); a system that implements one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is implemented at least in part in a data center; or a system that is implemented at least in part using cloud computing resources. Clause 20. One or more processors for: Determine at least one prompt verification score by at least: Generating a verification prompt comprising a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), at least one verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determine, using the at least one prompt verification score, whether to provide the prompt to the LM or to present at least one output generated based on the prompt.

[0142] It is to be understood that aspects and embodiments described above are merely exemplary and that modifications in detail may be made within the scope of the claims.

[0143] Each device, method, and feature disclosed in the description and (if applicable) in the claims and drawings may be provided independently or in any suitable combination.

[0144] Reference signs appearing in the claims are for illustrative purposes only and are not intended to have a limiting effect on the scope of the claims.

Claims

[1] A process comprising: Obtaining a plurality of tokens associated with a prompt; Determining one or more prompt verification scores by at least, for an individual prompt verification score of the one or more prompt verification scores: Generating a verification prompt, comprising: a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), the individual prompt verification score, which indicates a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determine, using the one or more prompt verification scores, whether to provide the prompt to the LM. [2] The method of claim 1, wherein the prompt comprises at least one of text, speech, video, or a plurality of images. [3] A method according to any preceding claim, wherein the second subset of one or more tokens comprises at least one or more of the following: a token that follows the first subset of one or more tokens, or a token that occurs within the first subset of one or more tokens. [4] A method according to any preceding claim, wherein the second subset of one or more tokens comprises a next token following the first subset of one or more tokens and excludes a token following the next token. [5] A method according to any preceding claim, wherein: the first subset of one or more tokens of a first verification prompt comprises first k tokens of the plurality of tokens, where k is an integer less than N-1, where N is a number of the plurality of tokens, the second subset of one or more tokens of the first verification prompt comprises k+1th token of the plurality of tokens, the first subset of one or more tokens of a second verification prompt comprises k+1 tokens of the plurality of tokens and the second subset of one or more tokens of the second verification prompt comprises k+2th token of the plurality of tokens. [6] A method according to any preceding claim, wherein determining whether to provide the prompt to the LM comprises: Calculating, using the one or more prompt verification scores, an evaluation metric for the prompt; and Comparing the evaluation metric with a threshold metric. [7] The method of claim 6, wherein calculating the evaluation metric for the prompt comprises: Aggregating the one or more prompt verification scores to obtain the evaluation metric. [8] The method of claim 6 or 7, wherein determining whether to provide the prompt to the LM further comprises: Determining that the evaluation metric is above the threshold metric; and Providing the prompt for the LM. [9] The method of any of claims 6-8, wherein determining whether to provide the prompt to the LM further comprises: Determining that the evaluation metric is below the threshold metric; and wherein the method further comprises: Generating a response comprising at least one of the following: a request to modify the prompt or a message that the prompt cannot be processed. [10] A method according to any preceding claim, further comprising: Determining one or more additional / supplemental prompt verification scores, wherein determining an individual additional prompt verification score of the one or more additional prompt verification scores comprises: Obtaining, using a second language model LM, the individual prompt verification score, which characterizes a probability that the second subsentence of one or more tokens, in the prompt, together with the first subsentence one or more tokens occurs; and wherein determining whether to provide the prompt to the LM comprises: Using the one or more additional prompt verification scores. [11] The method of claim 10, wherein determining whether to provide the prompt to the LM further comprises: Calculating, using the one or more prompt verification scores, a first evaluation metric for the prompt; and Calculating, using the one or more additional prompt verification scores, a second evaluation metric for the prompt. [12] The method of claim 11, further comprising performing at least one of the following: Providing, in response to the first evaluation metric being above the second evaluation metric, the prompt for the LM or Provide, responsive to the first evaluation metric being below the second evaluation metric, the prompt for the second LM. [13] A system comprising: one or more processing units for: Obtaining a plurality of tokens associated with a prompt; Determine at least one prompt verification score by at least: Generating a verification prompt comprising a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), at least one prompt verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determining, using the at least one prompt verification score, whether to provide the prompt to the LM or to present output generated using the prompt. [14] The system of claim 13, wherein the second subset of one or more tokens comprises at least one of the following: a token that follows the first subset of one or more tokens, or a token that occurs within the first subset of one or more tokens. [15] The system of claim 13 or 14, wherein to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units are to: Calculating, using the at least one prompt verification score, an evaluation metric for the prompt; and Comparing the evaluation metric with a threshold metric. [16] The system of claim 15, wherein to determine whether to provide the prompt to the LM or to present the output generated using the prompt, the one or more processing units are further operable to perform at least one of the following: Providing, responsive to determining that the evaluation metric is above the threshold metric, the prompt for the LM or Responsive to determining that the evaluation metric is below the threshold metric, generating a response comprising at least one of: (i) a request to modify the prompt or (ii) a notification that the prompt cannot be processed. [17] The system of claims 13-16, wherein the one or more processing units further serve to: Determining one or more additional / supplementary prompt verification scores, wherein in order to determine an individual additional prompt verification score of the one or more additional prompt verification scores, the one or more processing units are to: Obtaining, using a second language model LM, the individual prompt verification score, which characterizes a probability that the second subsentence of one or more tokens, in the prompt, together with the first subsentence one or more tokens occurs; and wherein, to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units are to: Using the one or more additional prompt verification scores. [18] The system of claim 17, wherein to determine whether the prompt should be provided to the LM or whether the output generated using the prompt should be presented, the one or more processing units further serve to: Calculating, using the one or more prompt verification scores, a first evaluation metric for the prompt; and Calculating, using the one or more additional prompt verification scores, a second evaluation metric for the prompt; and wherein the one or more processing units are further operable to perform at least one of the following: Providing, in response to the first evaluation metric being above the second evaluation metric, the prompt for the LM or Provide, responsive to the first evaluation metric being below the second evaluation metric, the prompt for the second LM. [19] A system according to any one of claims 13-18, wherein the system consists of at least one of the following: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for carrying out one or more digital twinning operations; a system for performing a light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system that implements one or more large language models (LLMs); a system that implements one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least in part using cloud computing resources. [20] One or more processors for: Determine at least one prompt verification score by at least: Generating a verification prompt comprising a first subset of one or more tokens of the plurality of tokens and a second subset of one or more tokens of the plurality of tokens; and Obtaining, using a language model (LM), at least one verification score indicating a probability that the second subset of one or more tokens occurs in the prompt together with the first subset of one or more tokens; and Determine, using the at least one prompt verification score, whether to provide the prompt to the LM or to present at least one output generated based on the prompt.