AUTOMATIC TRANSCRIPT-ASSISTED SPEECH LANGUAGE TRANSLATION USING LANGUAGE MODELS
The TAT system addresses inaccuracies in AST by using a two-stage process with ASR-generated text transcriptions to enhance language model translations, improving translation accuracy and reducing errors.
Patent Information
- Application Number
- DE102025137262
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-17
- Filing Date
- 2025-09-16
- Publication Date
- 2026-03-19
AI Technical Summary
Existing language model-based automatic speech translation (AST) technologies do not fully utilize prompt engineering, leading to inaccuracies and hallucinations in translations.
A two-stage transcription-assisted translation (TAT) system that processes speech input through an ASR model to generate a text transcription in the same language, which is then used as context for a language model to improve translation accuracy, utilizing encoder-decoder or decoder-only models for more precise translations.
The TAT system achieves more accurate and reliable automatic speech translations with reduced misidentifications and hallucinations by leveraging the context provided by the text transcription.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA OF INVENTION
[0001] At least one embodiment relates to computing resources used to perform and facilitate various speech-to-text tasks performed using machine learning, including, but not limited to, conversational artificial intelligence (AI), automatic speech recognition, and automatic speech translation. For example, at least one embodiment relates to the use of language models to facilitate and improve the translation of spoken language from one language to another. GENERAL STATE OF THE ART
[0002] Speech recognition, also known as automatic speech recognition (ASR), is a hybrid of computer technology and linguistics focused on techniques for recognizing and translating spoken language into text that can be displayed on a screen, printed, stored, used as input to a conversational model, as an instruction, or in any other way. ASR systems often employ machine learning models (MLMs), such as trained neural networks, to recognize patterns of spoken language in a given language and to identify units of spoken language, such as phonemes, graphemes, words, part words, sentences, and the like.ASR systems are commonly used in user-facing applications such as virtual agents, live captioning, clinical note-taking, and the like, and can be trained using speech samples produced by multiple speakers to process different dialects and accents. Automatic speech translation (AST) is a technology that converts spoken words, phrases, and sentences in a first language into corresponding speech units in a second language. AST can be performed using ASR to transcribe the spoken words into text, followed by machine translation of the transcribed text into the second language, or by directly converting spoken words into the second language. ASR and AST are parts of the speech-to-text (S2T) technology group.Various S2T systems and technologies can be used alone – e.g., to generate transcriptions and / or other recordings of spoken language, e.g., to synthesize new spoken language – or in conjunction with various text-to-speech algorithms (T2S algorithms), e.g., to perform natural language conversations.Other automatic speech tasks facilitated by machine learning include: speaker identification, which involves associating spoken utterances with speakers whose speech samples are stored in a database of speakers (or detecting a new speaker not represented in the database); speaker verification, which involves determining whether two or more utterances are made by the same speaker or different speakers; speech diarization, which involves partitioning unstructured speech among different participants in a conversation or meeting; and / or other tasks. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a block representation of an exemplary computer system according to at least one embodiment, which is capable of training and deploying an automatic transcription-assisted translation (TAT) system that uses language models; Fig. 2 illustrates an exemplary computing device, according to at least one embodiment, which supports the training and use of an automatic TAT system that uses language models; Fig. Figure 3A illustrates exemplary operations of the use of an automatic TAT system according to at least one embodiment; Fig. Figure 3B illustrates exemplary operations of another automatic TAT system according to at least one embodiment; Fig. Figure 4 illustrates exemplary operations, according to at least one embodiment, of training an automatic TAT system that uses language models; Fig. Figure 5 is a flowchart of an exemplary procedure for the deployment of automatic TAT systems according to at least one embodiment that uses language models; Fig. Figure 6 is a flowchart of an exemplary procedure for training automatic TAT systems according to at least one embodiment that uses language models; Fig. 7A illustrates inference and / or training logic according to at least one embodiment; Fig. 7B illustrates inference and / or training logic according to at least one embodiment; Fig. Figure 8 illustrates the training and use of a neural network according to at least one embodiment; Fig. 9 is an exemplary data flow diagram for an advanced computing pipeline, according to at least one embodiment; Fig. Figure 10 is a system representation for an exemplary system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline according to at least one embodiment; Fig. 11A is a block representation of an exemplary generative language model system suitable for use in implementing at least some embodiments of the present disclosure; Fig. 11B is a block representation of an exemplary generative language model system that includes a transformer-encoder-decoder suitable for use in implementing at least some embodiments of the present disclosure; Fig. 11C is a block representation of an exemplary generative language model system that includes a decoder-transformer-only architecture suitable for use in implementing at least some embodiments of the present disclosure; Fig. Figure 12 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. Figure 13 is a block representation of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0003] Language models (LMs), including large language models (LLMs), small language models (SLMs), vision language models (VLMs), multimodal language models (MMLMs), etc., have achieved remarkable success in a variety of natural language processing tasks, including supporting natural language conversations, understanding speaker intent and emotions, explaining complex topics, generating new text / images / audio / etc. upon receiving appropriate prompts, writing and testing software code, providing advice to users on topics of interest, and / or performing other functions. LMs are typically subjected to self-supervised training on vast amounts of text data (and / or other data, such as audio, image, video, 2D or 3D graphics or design data, etc.).) and learn to predict the next and / or missing word in an expression / sentence, detect a human speaker's intention and / or sentiment, determine whether two sentences are related or unrelated, and / or perform other basic language tasks. Following initial training, language learners (LMs) often undergo instructor-led (prompt-based), supervised fine-tuning, which enables them to acquire a more advanced level of language and / or master more specialized tasks, such as learning about financial markets, solving mathematical problems, etc. Fine-tuning may be supervised, for example, with learning prompts (questions, hints, etc.) accompanied by example texts (e.g., answers, sample essays, etc.) used as ground truth for LMs to emulate.Later stages of fine-tuning can also include reinforcement learning, where a human grader assigns marks indicating the degree to which the generated text resembles human-produced text. Existing learning modules (LMs) demonstrate in-context learning capability from a small number of representative examples, even when similar examples were not seen by the LMs in the preceding training.
[0004] These learned skills of language learners (LMs) are also attractive for extension to other input modalities (different from typed text), including speech modalities (audio modalities). For example, an LM can be trained to process audio input as prompts, e.g., represented by coded speech input, and to generate transcriptions of the audio input (in the case of ASR processing) in the same language or translations of the audio input (in the case of AST processing) into a different language. The performance of LMs is known to depend significantly on the quality of the prompts, which leads to the field of prompt engineering. A well-designed prompt can provide an LM with important context to greatly improve the accuracy and relevance of LM output and reduce the occurrence of hallucinations, e.g.,This can be reduced by guiding the LM in the right direction, recommending relevant keywords, and so on. However, existing LM-based AST technology does not fully utilize prompt engineering.
[0005] Aspects and embodiments of the present disclosure address these and other technological challenges by providing transcription-assisted translation (TAT) of speech input using language models. In particular, a TAT system can be configured to perform translations from speech in a first language to text in a second language via two-stage processing. The first stage can include processing a trained ASR model of the input speech to generate a text transcription in the same first language as the speech input. The second stage can include processing, by a LM, the same speech input together with the text transcription generated by the first stage.The speech input can first be processed using an audio encoder model that represents the speech input in a format the language learning (LM) can understand. The text transcription in the first language, used as additional input, provides the LM with context for the speech input, thus facilitating more accurate translations. As such, the two-stage TAT system emulates an iterative approach that a human translator can follow by using the context of the speech utterance or episode derived from its transcription before translating the speech in light of this global context. In some embodiments, the LM can be an encoder-decoder model that encodes the entire speech utterance / episode along with the text inputs and predicts the translation output as a whole sequence.In other embodiments, the LM can be a decoder-only model that autoregressively predicts individual tokens of the translation output, with individual token predictions based on speech utterances / episodes and text inputs, as well as preceding output tokens (e.g., via next-token prediction). The first stage can employ any suitable ASR model, such as a decoder-only model. In some embodiments, the model used in the first stage can be the same LM model used in the second stage.
[0006] Training of the TAT system can be performed in stages. For example, the ASR model can be pre-trained. Similarly, the LM can be pre-trained to achieve a general understanding of the first and second languages and / or to demonstrate basic translation skills. Parameters (e.g., weights and biases) of a substantial part of the LM can then be fixed ("frozen"), while a relatively small (e.g., 10 percent or less) adapter portion of the LM can be modified during training. In some embodiments, the audio encoder can be trained along with the adapter portion of the LM. A training input can include a speech component, e.g., training speech in the first language encoded by the audio encoder, and a text component, e.g., a transcription of the training speech.The transcription can be obtained by the same ASR model that is subsequently used as part of the TAT system for inference and can serve as a hypothesis or clue to be used to generate a more accurate translation. Training outputs—translations generated by the LM model—can be compared with ground-truth translations. In some embodiments, ground-truth translations can be generated from known ground-truth transcriptions of training inputs using a trained translation model.
[0007] The benefits of the disclosed techniques include, but are not limited to, improved automatic speech translations that are more accurate and have a lower number of misidentified and / or mistranslated words, phrases, and hallucinations.
[0008] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of the claims are described herein.
[0009] Devices, systems, and techniques are disclosed that implement the training and deployment of automatic transcription-based translation systems using language models. The techniques include: processing, using a first speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language. The second input includes at least a representation of the spoken language and the transcription of the spoken language.
[0010] The revelation extends to any novel aspects or features described and / or illustrated herein.
[0011] Further features of the disclosure are characterized by the independent and dependent claims.
[0012] Any feature in one aspect of the disclosure can be applied to other aspects of the disclosure in any suitable combination. In particular, procedural aspects can be applied to device or system aspects and vice versa.
[0013] Furthermore, features implemented in hardware can also be implemented in software, and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0014] Each system or device feature as described herein can also be provided as a process feature, and vice versa. Functionally described system and / or device aspects (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0015] It should also be apparent that certain combinations of the various features described and defined in any aspect of the disclosure can be implemented and / or provided and / or used independently of one another.
[0016] This disclosure also provides computer programs and computer program products comprising software code adapted, when executed on a data processing device, to perform any of the procedures described herein and / or to embody any of the device and system features described herein, including any or all of the partial steps of a procedure.
[0017] The disclosure also provides a computer or computing system (including networked or distributed systems) that includes an operating system supporting a computer program for performing any of the procedures described herein and / or embodying any of the device or system features described herein.
[0018] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0019] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0020] The disclosure extends to processes and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0021] Aspects and embodiments of the disclosure will now be described exclusively by way of example with reference to the accompanying drawings. SYSTEM ARCHITECTURE
[0022] Fig. Figure 1 is a block representation of an exemplary computer system 100 according to at least one embodiment, which is capable of training and deploying a transcription-based automatic translation system that uses language models. As in Fig. As illustrated in Figure 1, the computer system 100 can include a speech processing server 102, a data storage system 150, and a training server 160, connected to a network 140. The network 140 can be a public network (e.g., the Internet), a private network (e.g., a local area network, LAN, or a wide area network, WAN)), a wireless network, a personal area network (PAN), a combination thereof, and / or another type of network.
[0023] The speech processing server 102 can include a desktop computer, laptop computer, smartphone, tablet computer, server, wearable device, VR / AR / MR headset or head-up display, digital avatar or chatbot kiosk, in-vehicle infotainment computer, and / or any suitable computing device capable of performing the techniques described herein. The speech processing server 102 can be configured to receive a speech input 101 associated with a speech episode involving one or more speakers. Speech episodes can be a public or private conversation, a business meeting, a public or private presentation, an artistic event, a debate, or an interaction between a digital agent (e.g., chatbot, digital avatar, etc.).Speech input 101 can include communication between a vehicle and one or more users, in-vehicle communication (e.g., between two or more occupants and a chatbot, avatar, or digital assistant of the vehicle), and / or the like. Speech input 101 can include a statement, query, question, query for explanation / instruction, expression of feeling, narrative (or part of a narrative), memorandum, report, part of a conversation, and / or any other type of speech that can be produced by a user, including, but not limited to, a human user. In some embodiments, speech input 101 can include speech generated by a computer, such as a text-to-speech (T2S) model, a chatbot, a trained language model, and / or the like.The speech input 101 can be recorded using one or more devices connected to the speech processing server 102 (e.g., a microphone), retrieved from a memory 104 of the speech processing server 102, and / or received from an external computing device via a local network connection (e.g., via the network 140). The speech input 101 can be in any suitable format, e.g., WAV, AIFF, MP3, AAC, WMA, or any other compressed or uncompressed audio format. In some embodiments, the speech input 101 can be stored in the data memory 150 (e.g., together with other data, such as metadata).Furthermore, the data store 150 can store training speech 152 for training one or more models capable of speech recognition, speech translation, speech identification, speech verification, and / or speech diarization, according to some embodiments disclosed herein. The data store 150 can be accessed directly (e.g., via a bus, an intermediate connection, and / or the like) or (as in ) by the speech processing server 102. Fig. (1 shown) can be accessed via network 140.
[0024] The data store 150 can include persistent storage capable of storing audio data as well as metadata for the stored audio files. The data store 150 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tapes or hard disks, network-attached storage (NAS), storage area network (SAN), and so on. Although illustrated as separate from the speech processing server 102, in at least some embodiments the data store 150 can be part of the speech processing server 102.In at least some embodiments, the data storage 150 can be a network-connected file server, while in other embodiments, the data storage 150 can be another type of persistent storage, such as an object-oriented database, a relational database, etc., which can be hosted by the speech processing server 102 or by one or more different machines coupled to the speech processing server 102 via the network 140.
[0025] The speech processing server 102 can include the following: a memory 104 (e.g., one or more memory devices or units) communicatively coupled to one or more processing devices, such as one or more graphics processing units (GPUs) 110, one or more central processing units (CPUs) 130, one or more data processing units (DPUs), one or more parallel processing units (PPUs), and one or more other processing units (e.g., field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and / or the like). The memory 104 can include read-only memory (ROM), flash memory, dynamic random-access memory (DRAM), such as synchronous DRAM (SDRAM), static memory, such as static random-access memory (SRAM), and / or other memory capable of storing digital data.The memory 104 can store a transcription-based translation system (TAT system) 120, which implements various techniques of the present disclosure, e.g., automatic speech translation functions. In some embodiments, the TAT system can include a language model (LM) 122, e.g., a large language model, a multimodal language model (audio and text language model), and / or the like. The LM 122 can process audio representative of a speech input 101 in a first language together with text inputs that may include a transcription(s) of the speech. The transcriptions may be in the same first language and may be generated by an automatic speech recognition (ASR) model 124.Before being input to the LM 122, audio inputs can be processed by an audio encoder 126, which can be configured and trained to process audio data from a speech input 101 and convert the audio data into digital features (audio embeddings), thereby capturing the audio content of a speech input 101 and contextual relationships between different parts of a speech input 101. The memory 104 can also store a prompt generator 128, which generates a prompt for the LM 122 by combining the audio inputs and transcriptions, adding a query describing a task for the LM 122 to perform, and providing the prompt to the LM 122 for processing.
[0026] Speech inputs 101 can be generated by a user or any suitable computer application, e.g., an audio application, a video application, a gaming application, a digital assistant application, a communication application, a translation application, and / or any other application that receives, generates (e.g., synthesizes), processes, transmits, and / or otherwise uses speech. In some embodiments, a speech input 101 can be accompanied by a non-speech input 103 to the TAT system 120. In some embodiments, a non-speech input 103 can be generated by the same user (e.g.,typed) and / or the same application that generated the speech input 101, and may include a description of a task to be performed by the TAT system 120 and / or various keywords, expressions, contexts, explanations and / or the like associated with the speech input 101.
[0027] In one exemplary embodiment, a speech input 101 and a non-speech input 103 can be received via a suitable user interface (UI) 106, which may include one or more devices of different modalities. For example, a speech input 101 can be received via an audio device, e.g., a microphone, and a non-speech input 103 can be received using a keyboard, touchscreen, touchpad, writing pad, graphical interface, mouse, stylus, and / or using any pointing device capable of selecting words / phrases displayed, e.g., on a screen, and / or another suitable device. In some embodiments, the speech and non-speech input devices can be separate, inseparable devices, e.g.,A microphone of a digital camera to receive speech input 101, and a computer keyboard to receive non-speech input 103. In some embodiments, speech and non-speech input devices can be integrated together (e.g., in a smartphone, tablet computer, and / or the like).
[0028] In at least one embodiment, various models employed and / or used by the TAT system 120, e.g., LM 122, ASR model 124, audio encoder 126, and / or other employed models and components, can be implemented as deep learning neural networks exhibiting multiple layers of linear and nonlinear operations. For example, any or some of the employed models can include convolutional neural networks, recurrent neural networks, fully connected neural networks, long-short-term memory (LSTM) neural networks, attentional neural networks (e.g., neural transformer networks), a combination of a convolutional network and one or more transformers (a conformer network), and / or neural networks of other types.In at least one embodiment, no arbitrary, some, or all of the models employed may include multiple neurons, wherein an individual neuron receives its input from other neurons and / or from an external source and produces an output by applying an activation function to the sum of weighted (using trainable weights) inputs and possibly a bias value. In at least one embodiment, one or more of the models employed may include multiple neurons arranged in layers, including an input layer, one or more hidden layers, and / or an output layer. Neurons of adjacent layers may be connected by weighted edges.In some embodiments, the training server 160 can train a number of different models, which may be models that differ in the number of neurons, number of neuron layers, specific neural architecture and / or the like.
[0029] The training server 160 can use training speech 152 to train one or more models, for example, to identify parameters (neural weights, biases, activation function parameters, etc.) of the models in a way that maximizes the accuracy of various S2T tasks performed by the TAT system 120, such as conversational AI tasks, ASR tasks, AST tasks, and / or other similar tasks. In at least one embodiment, the training server 160 and the audio processing server 102 can be implemented on a single computing device.The Training Server 160 and / or the Speech Processing Server 102 can be hosted by a rackmount server, router, desktop computer, laptop, smartphone, tablet, server, media center, and / or any suitable computing device or combination thereof capable of performing (and / or including) the techniques described herein. In some embodiments, the Training Server 160 can be implemented using one or more libraries and / or frameworks for training and deploying machine learning models, such as PyTorch libraries, TensorFlow libraries, NVIDIA® TensorRT™ software development kits, NVIDIA® NIM microservices, NVIDIA® NeMo conversational AI cloud toolkits, NVIDIA® Riva multilingual speech and translation microservices, and / or other systems, toolkits, and / or the like.In some implementations, any, some, or all machine learning tools can reside in a cloud.
[0030] In some embodiments, the training of the TAT system 120 can be supervised, for example, using various annotations of training speech 152. Such annotations can include ground-truth transcriptions 154 of the training speech 152, which may be in the same language as the corresponding training speech 152. Such annotations can further include ground-truth translations 156 of the training speech 152 (or translations of the ground-truth transcriptions 154) and / or the like. The training speech 152 can be used for supervised training, unsupervised training, semi-supervised training, training that includes reinforcement learning, and / or other types of training.
[0031] The training speech language 152 can be used by the training engine 162 or the training server 160 as training input 165 to train one or more models (networks) deployed by the TAT system 120 to recognize spoken words of the training speech language 152 and translate them into a different language. For example, during the training of the ASR model 124, the training input 165 can include a training speech language 152 in a first language, and training outputs generated by the ASR model 124 can include training transcriptions in the first language. The training transcriptions in the first language can be compared with target outputs 167, such as ground truth transcriptions 154.Similarly, during training of the LM 122, the training input 165 can include a training speech language 152 and ground-truth transcriptions 154, while training outputs generated by the LM 122 can include training translations in a second language. The training translations can be compared with ground-truth translations 156. During co-training of the LM 122 and the ASR model 124, training inputs to the ASR model 124 can include training speech language 152 in a first language, while training inputs to the LM 122 can include both training speech language 152 and training translations generated by the ASR model 124. The training output generated by the LM 122 during such co-training can include training translations in the second language, which are subsequently compared with ground-truth translations 156.
[0032] During the training of the TAT system 120, the training engine 162 can also generate mapping data 166 (e.g., metadata) that associates training inputs 165 with correct target outputs 167 (ground truth). During training, the training engine 162 can identify patterns in training inputs 165 based on desired target outputs 167 and train various models and / or networks of the TAT system 120 to perform ASR, AST, and other suitable speech tasks.
[0033] The training speech language 152 can be stored in a data storage device 150 in a raw audio format, e.g., in the form of spectrograms, or in any other suitable representation of the speech language. For example, a spectrogram of the training speech language 152 can be obtained by recording air pressure caused by the speech language as a function of time and calculating a short-time Fourier transform for overlapping time intervals (frames) of a fixed duration. This maps the audio signal from the time domain to the frequency domain and generates a spectrogram that characterizes the spectral content of the training speech language 152. The amplitude of the audio signal can be represented on a logarithmic scale (decibel scale). In some embodiments, the obtained spectrograms can be further converted into Mel spectrograms by converting frequency f into a non-linear Mel domain. f→m=a ln(1+fb), is transformed to account for the human ear's ability to distinguish between evenly spaced frequencies (tones) at the lower end of the audible spectrum better than at its higher end. In an example, a = 1607 Hz and b = 700 Hz. In this entire disclosure, the term "speech-speech spectrogram" may be understood to include Fourier spectrograms, mel spectrograms, wavelet spectrograms, time-domain filter banks, custom filter banks, learnable filter banks, custom audio embeddings, and / or other forms of analog or digital audio intermediate representation, or a combination thereof.
[0034] Initially, parameters (e.g., edge weights and biases) of various models and / or networks being trained can be assigned some initial values (e.g., random values). For different training inputs 165, the training engine 162 can cause the TAT system 120 to generate one or more training outputs. The training engine 162 can then compare the training output(s) with the desired target output(s) 167. The resulting error or discrepancy, e.g., the difference between the target output(s) 167 and the training output(s) of the neural networks, can be analyzed using various models and / or networks. B. LM 122, ASR model 124, audio encoder 126, can be propagated back and the weights and biases in the neural networks can be adjusted to bring the training outputs closer to the target (ground-truth) outputs 167.This adjustment can be repeated until the output error for a given training input 165 meets a predetermined condition (e.g., falls below a predetermined value). Subsequently, a different training input 165 can be selected, a new output 166 generated, and a new series of adjustments implemented until the respective neural networks are trained to a target accuracy level or until the respective neural network(s) converge to a limit of its accuracy.
[0035] In some embodiments, the LM 122 can be trained by the training engine 162. In some embodiments, the LM 122 can be a model trained and deployed by an external entity (relative to the speech processing server 102), such as a speech model service 170, which can be a cloud service, a subscription service, and / or a combination thereof. In some embodiments, the LM 122 (and / or other deployed speech models) can be or include a large language model (LLM). The LM 122 can be trained to understand the syntax and semantics of human speech, for example, by predicting the next, preceding, and / or missing word in a sequence of words (e.g., one or more sentences of human speech or human text).The LM 122 can be trained using training data containing a large number of texts, such as human dialogues, newspaper articles, magazine articles, book texts, web-based texts, and / or other types of text. The trained LM 122 can then be capable of engaging in a (textual) conversation with a user (a human user or a computer) in natural language in a manner closely resembling a dialogue with a human speaker, including understanding the user's intent and responding in a way the user would expect from a conversation partner. The LM 122 can be implemented using neural networks with a large number (e.g., billions) of artificial neurons, including, but not limited to, deep learning neural networks equipped with a self-awareness mechanism (such as neural transformer networks).
[0036] The translation capability acquired by the TAT system 120 during training can subsequently be evaluated (validated or tested) using additional evaluation or validation inputs. The trained TAT system 120 can then be used during the inference stage to process new (previously unobtained) speech inputs 101.
[0037] Fig. Figure 2 illustrates an exemplary computing device 200, according to at least one embodiment, which supports the training and use of the automatic TAT system, which uses speech models. In at least one embodiment, the computing device 200 can be part of the speech processing server 102. In at least one embodiment, the computing device 200 can be part of the training server 160. In at least one embodiment, the computing device 200 supports the TAT system 120, which includes (but is not limited to) the LM 122, the ASR model 124, the audio encoder 126, the prompt generator 128, and / or other components. The TAT system 120 can be capable of processing an input 201 and generating a text output 202. The input 201 can be a speech input (e.g., speech input 101) received via any audio device (e.g.,a microphone) is received in real time, or it may include a previously recorded audio input that includes speech. The audio device may be part of a UI 106, which may further include non-audio devices, such as a keyboard, touchscreen, writing pad, and / or the like, to receive a non-speech portion of the input 201. The non-speech portion may include instructions to the TAT system 120 specifying: the type of text output 202 to produce, and the context associated with the speech input, such as one or more typed (or otherwise selected) keywords, phrases, acronyms, and / or the like.
[0038] The input 201 in a first language can be processed by the ASR model 124, which generates a transcription in a first language, which is combined by the prompt generator 128 with audio embeddings generated by the audio encoder 126 (representative of the speech language of the input 201) to form a prompt, which is processed by the LM 122 to generate a text output 202 (e.g. translation of the speech language into the second language).
[0039] Operations of the TAT system 120 can be performed using one or more GPUs 210, one or more CPUs 230, one or more parallel processing units (PPUs) or accelerators, such as a deep learning accelerator, data processing units (DPUs), and / or the like. In at least one embodiment, a GPU 210 includes multiple cores 211, each core being capable of executing multiple threads 212. Each core can execute multiple threads 212 concurrently (e.g., in parallel). In at least one embodiment, threads 212 can access registers 213. The registers 213 can be thread-specific registers with access to a register restricted to a particular thread. In addition, shared registers 214 can be accessed by one or more (e.g., all) threads of the core.In at least one embodiment, each core 211 can include a scheduler 215 to distribute computational tasks and processes among different threads 212 of the respective core 211. A processing unit 216 can implement scheduled tasks on appropriate threads using correct private registers 213 and shared registers 214. The computing device 200 can include one or more input / output components 234 to facilitate the exchange of information with one or more users or developers.
[0040] In at least one embodiment, the GPU 210 can have a (high-speed) cache 218, wherein multiple cores 211 can share access to it. Furthermore, the computing device 200 can include GPU memory 219, wherein the GPU 210 can store intermediate and / or final results (outputs) of various calculations performed by the GPU 210. Upon completion of a respective task, the GPU 210 (or the CPU 230) can move the output to (main) memory 204. In at least one embodiment, the CPU 230 can execute processes that involve serial computing tasks, wherein the GPU 210 can perform tasks (such as multiplying inputs of a neural node by weights and adding biases) that are accessible for parallel processing.In at least one embodiment, the TAT system 120 can determine which processes are to be executed on the GPU 210 and which processes are to be executed on the CPU 230. In other embodiments, the CPU 230 can determine which processes are to be executed on the GPU 210 and which processes are to be executed on the CPU 230. SPEECH PROCESSING SYSTEM
[0041] Fig. Figure 3A illustrates exemplary operations 300 of the use of an automatic TAT system according to at least one embodiment. Exemplary operations 300 can be performed by the TAT system 120 of Fig. 1 and / or Fig. 2 in some embodiments. Text outputs generated by the TAT system 120 can include translations of speech inputs 101 from a first language into a second language, which can include any languages from the same language family (e.g., Spanish-to-Italian translations) or from different language families (e.g., English-to-Mandarin-Chinese translations). In at least one embodiment, the TAT system 120 can be supported by a speech processing server 102, which can be located on a single computing device or distributed across multiple computing devices. Various in Fig. 3A with the same reference symbols as the blocks of Fig. 1 and / or Fig. The two marked blocks implement the same (or similar) functionality.
[0042] As in Fig. As illustrated in Figure 3A, the TAT system 120 can receive speech input 101, which is captured using one or more audio sensors, such as microphones. Microphones can include dynamic microphones, condenser microphones, ribbon microphones, unidirectional microphones, omnidirectional microphones, and / or any other type of microphone. In some embodiments, a microphone can be combined with other devices, such as computers, telephones, loudspeakers, TV screens, smart kiosks, smart speakers / displays, in-vehicle or in-cabin infotainment or computing devices, and / or the like. The speech input 101 collected by the audio sensors can be generated by any number of speakers, such as spoken speech, and can include a single speech episode or multiple speech episodes.The audio sensors can detect not only a spoken language signal, but also background noise, interference signals, e.g. emitted by TV devices, radio devices, alarm devices and / or any other devices, or naturally occurring sounds (e.g. sound of wind, water, birds, etc.).
[0043] The speech input 101 can be subjected to audio preprocessing 302. For example, the audio preprocessing 302 can include filtering, denoising, amplification, reverberation reduction, segmentation, and / or any other suitable audio signal enhancement. The audio preprocessing 302 can further include the removal of parts of the speech input 101 that do not contain speech content. For example, the audio preprocessing 302 can evaluate energy e(t) associated with the audio data as a function of time and identify regions that have energy less than a certain threshold (e.g., an empirically determined noise source). Such identified regions can be removed (trimmed) from the speech input 101 during the audio preprocessing 302. Segmentation can involve segmenting the speech input 101 into intervals of a predetermined size (duration), τ, e.g., 0.05–5 seconds., include. Such intervals need not correspond to a complete logical unit of speech and may contain one or more sentences, one or more words, part of a word, one or more exclamations, filler words, pauses, and / or the like. In some embodiments, the intervals may partially overlap.
[0044] Individual frames can be represented by one or more frames, e.g., T-frames, over a predetermined time interval. Frames can have a duration of 15 ms, 20 ms, 30 ms, 80 ms, and / or another duration. Frames can be subjected to a suitable frame-to-spectrogram transformation to generate spectrograms 310. Spectrograms 310 can include any suitable representation of audio data, including (but not limited to) Fourier spectrograms, mel spectrograms, wavelet spectrograms, time-domain filter banks, custom filter banks, learnable filter banks, custom audio embeddings, and / or other forms of analog or digital audio intermediate representation, or a combination thereof.For example, spectrogram(s) 310 of a frame can be obtained or generated by performing discrete Fourier transformations of acoustic energy(t) or air pressure p(t) associated with a specific utterance. The resulting spectrograms e(f. j ) can be used for a number of bands f1, f2... f C , for example, for C = 80 bands or C = 128 bands or any other number of bands. In some embodiments, the bands can be mel-bands and the spectrograms can be mel-spectrograms. Separate spectrograms 310 can be obtained for separate audio frames.
[0045] Spectrograms 310 can be processed by the ASR model 124, which can be any machine learning model trained to recognize spoken words, part words, phrases, and / or any other units of spoken utterances, and which can represent these utterances in a transcribed form, e.g., using letters, numbers, capitalization, punctuation, and / or other elements of written language. In some embodiments, the ASR model 124 can be a monolingual model trained to recognize spoken language in a specific language (e.g., English or Chinese). In other embodiments, the ASR model 124 can be a multilingual model trained to recognize any one language from a set of languages (e.g., either English or Spanish, or any combination thereof). The ASR model 124 can have any suitable architecture, e.g.,A decoder-only ASR model, encoder-decoder ASR model, encoder-only ASR model, attention-based model, transformer-based model, conformer-based model, and / or the like. In some embodiments, the ASR model 124 can be a Connectionist-Temporal-Classification (CTC) model, a Recurrent Neural-Network Transducer (RNNT) model, and / or the like. The ASR model 124 can generate a transcription 320 of a speech input 101 in the same first language as the speech input 101.
[0046] Spectrograms 310 of the speech input 101 can also be processed by an audio encoder 126, which represents the speech input 101 via audio embeddings (features) 330 in a format that can be understood by the LM 122. Audio embeddings 330 generated by the audio encoder 126 can capture temporal and frequency correlations of the speech input 101. An embedding should be understood as a suitable digital representation of a unit (e.g., frame, part of a frame, multiple frames, etc.) of a speech input 101, e.g., as a vector (string) of any number D of components, which may have integer or floating-point values. Embeddings can be viewed as vectors or points in a D-dimensional embedding space. The dimensionality D of the embedding space may be smaller than the size of the speech input 101 (or spectrograms 310).The audio encoder 126 can be trained to associate similar sets of the training audio spectrograms with similar embeddings represented by points that are closely adjacent in the embedding space, and furthermore to learn to associate dissimilar sets of the training audio spectrograms with points that are further apart in the embedding space. A given audio embedding 330 can encode (represent) one or more words or part (e.g., one or more syllables of phonemes) of a word.
[0047] In some embodiments, the audio encoder 126 can be of a conformer type. The conformer architecture combines elements of transformer networks, such as self-attention layers, with elements of convolutional networks, such as layers of kernels (filters) that narrow or widen a perceptual field. For example, a conformer network can include a stack of alternating multi-headed attention layers, depth-separable convolutional layers, and / or fully connected layers. Some of the layers of a conformer network can be connected via residual (skipped) connections. In some embodiments, a conformer network can include a downsampling module that can be deployed at the start of the conformer to modify a frame rate of audio embeddings 330, for example, from the 20 ms interval per frame to the 80 ms interval in an illustrative example.In some embodiments, a fast conformer may be employed. A fast conformer may differ from a conventional conformer in its use of a larger initial downsampling (e.g., 8x downsampling) to reduce the computational cost of subsequent attention layers, replacing some of the convolutional sub-sampling layers with depth-separable convolutions, reducing the number of convolutional filters in the downsampling block(s) (e.g., to 256), and further reducing the size of one or more convolutional kernels (e.g., to 9). In some embodiments, the audio encoder 126 may include a NeMo-type model having 100M or more learnable parameters, a canary model, and / or any other suitable model.
[0048] Audio embeds 330 can be used as input to the LM 122. The transcription 320 in the first language can be used as additional input to the LM 122 to provide the LM 122 with a context of the speech input 101 and to facilitate more accurate translation. The prompt generator 128 can combine audio embeds 330 with the transcription 320, for example, by concatenating them, to obtain the prompt 345. In some embodiments, the prompt generator 128 can also add a task query 340 to the prompt 345. The task query 340 can include a natural language description of a task to be performed by the LM 122. In one example, the prompt 345 can include the following task query: Transcribe the spoken German content into English text using the following German text as a guide: <<Verschandeln Sie die Stätte nicht durch Anbringen oder Einkratzen von Graffiti.> >
[0049] In some embodiments, the transcription 320 and the task query 340 can be performed by a suitable LM tokenizer (in Fig. (3A omitted for brevity) can be converted into text tokens with a format understood by the LM 122. For example, the LM 122 can work in conjunction with a known set of tokens that can include any suitable representation of units of spoken language (e.g., syllables, words, etc.) as numbers. In one example of GPT-4 tokens, the word "the" can be represented by the token "280", the word "import" by the token "476", the word "description" by the token "4097", and so on. In other embodiments, individual words can be represented using any number of tokens, or word transitions (e.g., end of one word, beginning of the next) can be represented using a single token. As such, tokenization can be performed in any way suitable for input into the network.
[0050] The LM can process the prompt 345 and output a translation 360 in the second language using the transcription 320 in the first language as a hint, hypothesis, etc., which provides a context for the speech input 101. In some embodiments, the LM 122 can be an encoder-decoder model that encodes any, some, or all of a speech input 101, audio embeds 330, and / or a task query 340 via one or more hidden states and then generates a translation 360 as a whole sequence. In encoder-decoder models, attention can be calculated between all units, e.g., tokens, audio embeds, etc., of the prompt 345 and any tokens of the translation 360. In other embodiments, the LM 122 can be a decoder-only model.In decoder-only models, translation tokens 360 can be computed autoregressively, with attention scores being calculated in a causal way, e.g., between a given translation token 360 and previous translation tokens (and other inputs, e.g., speech input 101, audio embeddings 330, task query 340, etc.).
[0051] In some embodiments, the LM 122 can be a frozen model, e.g., a model whose parameters are pre-trained (e.g., by a language modeling service 170 of Fig. 1 is performed) are fixed and during the training of the TAT system 120 (e.g. in conjunction with Fig. (4 disclosed in more detail below) may not be changed. In some embodiments, the TAT System 120 may include a trained LM Adapter 350 to facilitate learning and performing language tasks. The LM Adapter 350 may be a lightweight model that has a smaller (in some embodiments, much smaller) number of trainable parameters compared to the LM 122. The smaller number of parameters of the LM Adapter 350 makes training the TAT System 120 significantly faster and less costly; for example, it requires less training data and fewer training epochs.
[0052] In some embodiments, the LM adapter 350 can have a low-rank architecture. In particular, operations on a given linear layer of the LM 122 can be performed on a (frozen) h × d matrix of weights W. h × d The LM adapter 350 (for the same layer) can handle multiple matrices, e.g., two. h × r(with dimension h × r) and B r × d (with dimension r × d), where dimension r is much smaller than h or d (or both r « h, d). Learned (during supervised training) elements of matrices A h × r and B r × d can be used during inference to determine weights W h × d to extend the LM 122, e.g. according to: Wh×d→Wh×d+Ah×r⋅Br×d. Accordingly, an input to the layer of the LM 122 is provided by two parallel branches, e.g. frozen weights W. h × d of LM 122 and low-rank matrix product A h × r · B r × d of the LM adapter 350, processed. A similar extension can be carried out for other layers of the LM 122.
[0053] In some embodiments, e.g., when the TAT system 120 is used to perform offline translations, audio embeds 330 can represent an entire spoken utterance of a specific spoken episode. In some embodiments, audio embeds 330 can represent a portion of a spoken utterance (e.g., several minutes of a spoken episode), with different parts of the spoken utterance being processed independently, e.g., sequentially. In some embodiments, the TAT system 120 can be used to perform streaming translations. In some embodiments, audio embeds 330 can represent a certain sliding window of an empirically determined duration π, e.g., several seconds or more. Successive intervals τ1, τ2, ... τ nThe intervals Δτ can be non-overlapping or overlapping for a certain period of time. In such embodiments, individual translations of 360° are possible for successive intervals τ1, τ2, ... τ. n generated and then aggregated and displayed on a UI, stored / transmitted over a network and / or the like.
[0054] Fig. Figure 3B illustrates exemplary operations 301 of another automatic TAT system according to at least one embodiment. In the Fig. 3B illustrated TAT system can demonstrate the functionality of the ASR model (e.g., the ASR model 124 in Fig. 3A) are implemented by the LM 122 (which may or may not include an LM adapter 350). Although, for clarity, two instances of the LM 122 (and LM adapter 350) are shown in Fig. As illustrated in Figure 3B, the same instance of LM 122 can be used for processing audio embeds 330 during the first stage of processing and prompt 345 during the second stage of processing. Other blocks of Fig. 3B can perform operations that are the same (or essentially the same) as the operations of Fig. 3A are (audio preprocessing 302 is in Fig. 3B not shown for the sake of conciseness).
[0055] Fig. Figure 4 illustrates exemplary operations 400, according to at least one embodiment, of training an automatic TAT system that uses language models. The system, whose training with Fig. As illustrated in section 4, the TAT system can be used by 120 people. Fig. 1 and / or Fig. 2, whose operational deployments are linked to Fig. 3A and / or Fig. 3B. In at least one embodiment, the training of the TAT system can be carried out by the training engine 162 of the training server 160 and subsequently transferred to the speech processing server 102 (with reference to Fig. 1) be uploaded. Various in Fig. 5 with the same reference symbols as the respective blocks of Fig. 3A and / or Fig. Blocks marked 3B implement the same (or similar) functionality.
[0056] The training speech input 401 in a first language can be acquired by one or more audio sensors and can be subjected to audio preprocessing 302 to be represented by one or more spectrograms 410. Spectrograms 410 can be processed by the ASR model 124. The ASR model 124 can generate a transcription 420 in the first language, which is used by the prompt generator 128 to form a training prompt 445. The training prompt 445 can further include audio embeddings 430 generated by processing spectrograms 410 using the audio encoder 126, and can also include a training task query 440 with a natural language description of a training task to be performed by the LM 122. The LM 122 can process the training prompt 445 and generate a translation 460 in a second language.The translation 460 can be compared with a ground-truth translation 465 using an output evaluation function 470. The difference (the discrepancy) between the LM-generated translation 460 and the ground-truth translation 465 can be used to modify parameters for various models of the TAT system, e.g., LM adapter 350, ASR model 124, audio encoder 126, and / or the like. For example, the modification of the parameters can be performed using various algorithms of backpropagation, gradient descent, and / or other training techniques.
[0057] In some embodiments, the LM 122 can be pretrained and then frozen during subsequent training of the TAT system. The LM 122 can be a large language model with billions of parameters and can be pretrained (using self-supervised training) on a set of tokens, using web crawl data, messages, conversations, books, scientific texts, and / or the like. The LM 122 can be trained using both English and non-English texts. Subsequently, the LM 122 can be fine-tuned using any suitable public instruction datasets.
[0058] In some embodiments, the audio encoder 126 and / or the LM adapter 350 can be trained together. In some embodiments, the ASR model 124 can be trained before training the audio encoder 126 and / or the LM adapter 350. In some embodiments, the ASR model 124 can be pre-trained first, but then subjected to additional training (tuning) together with the audio encoder 126 and / or the LM adapter 350. In some embodiments, training speech and ground-truth transcripts of training speech can be obtained from publicly available audio / text pairs.
[0059] Fig. 5 and Fig. Figure 6 shows flowcharts of methods 500 and 60 respectively, which facilitate the training and deployment of automatic TAT systems using speech models according to at least one embodiment. Methods 500 and / or 600 can be performed using one or more processing units (e.g., CPUs, GPUs, accelerators, PPUs, DPUs, etc.) that can include (or communicate with) one or more storage devices. In at least one embodiment, methods 500 and / or 600 can be performed using processing units 600 of the speech processing server 102 or training server 160. Fig. 1. In at least one embodiment, the processing units performing any one of the methods 500 and / or 600 can execute instructions stored on a non-transient, machine-readable storage medium. In at least one embodiment, any one of the methods 500 and / or 600 can be performed using multiple processing threads (e.g., CPU threads and / or GPU threads), with individual threads performing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, processing threads implementing one of the methods 500 and / or 600 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, processing threads implementing one of the methods 500 and / or 600 can be executed asynchronously with respect to each other.Various operations of any of the procedures 500 and / or 600 can be compared with those in . Fig. 5 and Fig. The sequence shown in Figure 6 may be performed in a different order. Some operations of one of methods 500 and / or 600 may be performed simultaneously with other operations. In at least one embodiment, one or more of the following may be performed: Fig. 5 and / or Fig. The 6 operations shown are not always performed.
[0060] Procedures 500 and / or 600 may be performed in the context of speech-to-text applications, chatbot applications, audio applications, video applications (including streaming audio and / or video applications), gaming applications, digital assistant applications, video conferencing applications, audio conversation applications, audio / text messaging applications, and / or any other suitable applications that use speech-to-text translation. Procedures 500 and / or 600 may involve speech utterances produced by humans in any possible context, such as a conversation, a public speech, a public event, a business meeting, a conference, a street encounter, an interaction in a game, an interaction with a chatbot or digital avatar, an interaction with an in-vehicle infotainment system, and / or the like.“Speech language”, as used in the context of Procedures 500 and / or 600, should be understood to include sounds produced by humans as well as robotic speech language, e.g., synthesized or computer-generated speech, and / or the like.
[0061] Fig. Figure 5 is a flowchart of an exemplary method 500 for the deployment of automatic TAT systems according to at least one embodiment that uses speech models. One or more operations of the method 500 can be performed by one or more processing units of the speech processing server 102. Fig. 1. Operations of procedure 500 can be performed to translate spoken language in a first language into a text in a second language.
[0062] In block 510, one or more processing units executing procedure 500 can: process, using an initial speech-to-text (S2T) model (e.g., the ASR model 124 in Fig. 3A), a first input that is a spoken language in the first language (e.g. the spoken language input 101 in Fig. 3A) includes a transcription of the spoken language (e.g., the transcription 320 in Fig. 3A) to generate. In some embodiments, the transcription may be in the first language. In some embodiments, the first S2T model may be (or include) a decoder-only language model.
[0063] In Block 520, Method 500 can: generate a second input (e.g., the prompt 345) that includes a representation of the spoken language (e.g., audio embeddings 330) and the transcription of the spoken language. In some embodiments, the second input can further include a natural language description of a task to be performed by the second S2T (e.g., the task query 340) to, for example, generate the translation of the spoken language into the second language. In some embodiments, the second input can be generated using operations associated with highlighting Fig. Figure 5 illustrates this. In particular, for block 522, procedure 500 may include: generating the representation of the speech using an audio encoder network (e.g., audio encoder 126). For block 524, procedure 500 may include: concatenating the representation of the speech and the transcription of the speech to obtain the second input.
[0064] At block 530, procedure 500 can continue with the following: Processing, using a second S2T model (e.g., LM 122, LM 122 with LM adapter 350, and / or the like), the second input to perform a translation of the spoken language into a second language (e.g., the translation 360 into Fig. 3A). In some embodiments, the second S2T model may include an encoder-decoder language model. In some embodiments, the second S2T model may be the same as the first S2T model (e.g., as in Fig. 3B illustrated) or (or include this).
[0065] Fig. Figure 6 is a flowchart of an exemplary method 600 for training automatic TAT systems according to at least one embodiment that uses language models. One or more operations of the method 600 can be performed by the processing units of the training server 160. Fig. 1. For block 610, procedure 600 may include: Receiving a training input (e.g., the training prompt 445 in Fig. 4), which includes the following: a first part with a representation (e.g., audio embeddings 430 in Fig. 4) of a training speech language (e.g. spectrograms 410 of the training speech language input 401 in Fig. 4) in a first language and a second part which is a transcription (e.g. the transcription 420 in Fig. 4) the training speech in the first language. In some embodiments, the training input may further include: a third part (e.g., the training task query 440) with a natural language description of a task to be performed by the S2T model, e.g., to generate the translation of the training speech into the second language.
[0066] In some embodiments, receiving the training input may include operations associated with the top highlighting of Fig. 6 will be illustrated. In particular, for block 612, procedure 600 may include the following: processing, using an audio encoder network (e.g., audio encoder 126 in Fig. 4), the training speech (e.g., spectrograms 410), and to generate the representation of the training speech. For block 614, procedure 600 may include: processing, using a trained ASR model (e.g., the ASR model 124 in Fig. 4) the training speech in the first language to generate the transcription of the training speech.
[0067] For block 620, procedure 600 may include the following: processing, using a speech-to-text model (S2T model) (e.g., LM 122, LM 122 with LM adapter 350 and / or the like, with reference to FI. 4), the training input to produce a translation of the training speech into a second language (e.g., translation 460 into Fig. 4) to generate. In some embodiments, the S2T model may be (or include) an encoder-decoder language model.
[0068] For block 630, procedure 600 may include: comparing the translation of the training speech language into the second language with a ground-truth translation (GT translation) of the training speech language into the second language (e.g., the GT translation 465 in Fig. 4) In some embodiments, as illustrated by highlighting block 632, the method 600 may include: Obtaining the GT translation of the training speech language by translating a GT transcription into the first language using a translation model.
[0069] At block 640, procedure 600 may continue with the following: Modifying, at least based on the comparison performed at block 630, one or more parameters of the S2T model. As illustrated by the highlighted block 642 below, modifying one or more parameters of the S2T model may include: Modifying one or more parameters of the adapter network (e.g., the LM adapter 350 in Fig. 4) In some embodiments, the method 600, as illustrated by block 650, may include: modifying one or more parameters of the audio encoder network (e.g., of the audio encoder 126 in Fig. 4) In some embodiments, the method 600, as illustrated by block 660, may also include: modifying one or more parameters of the ASR model (e.g., of the ASR model 124 in Fig. 4).
[0070] The systems and methods described herein can be used for a variety of purposes, including but not limited to machine control (e.g., control of robots, vehicles, construction equipment, warehouse vehicles / machines, autonomous, semi-autonomous, and / or other machine types), machine locomotion, machine propulsion, synthetic data generation, model training (e.g., using real, augmented, and / or synthetic data, such as synthetic data generated using a simulation platform or system, synthetic generation techniques such as, but not limited to, those described herein, etc.), perception, analytical operations, factory operations, generation and / or presentation of augmented reality (AR), virtual reality (VR), mixed reality (MR), etc., robotics operations, medical operations, security, and monitoring (e.g.,in a smart city implementation form), autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, generative AI operations, conversational AI operations, operations involving vision language models, large language models, small language models, multimodal language models, light transport simulations (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation for 3D assets (e.g., using Universal Scene Descriptor (USD) data, such as OpenUSD, and / or other types), cloud computing applications, generative artificial intelligence applications (e.g., using one or more diffusion models, transformer models, etc.), and / or any other suitable applications.
[0071] Disclosed embodiments may be included in a range of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), an in-vehicle infotainment system for an autonomous or semi-autonomous machine, systems implemented using a robot or robotic platform, antenna systems, media systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulation, robot simulation, intelligent city or surveillance simulation, etc.), systems for performing digital twinning operations (e.g.,in conjunction with a collaborative content creation platform or system, such as, without limitation, NVIDIA's OMNIVERSE and / or any other platform, system, or service that uses USD or OpenUSD data types, systems implemented using an edge device, systems involving one or more virtual machines (VMs), systems for performing synthetic data generation operations (e.g., using one or more neural rendering fields (NERFs), Gaussian splatting techniques, diffusion models, transformer models, etc.).), systems that are at least partially implemented in a data center, systems for performing conversational AI operations, systems that implement one or more language models – such as one or more large language models (LLMs), one or more small language models (SLMs), one or more vision language models (VLMs), one or more multimodal language models, etc., systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets (e.g., using Universal Scene Descriptor (USD) data such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data, and / or other data types), systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0072] In some examples, the machine learning model(s) described herein (e.g., deep neural networks, language models, SLMs, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field models (NERF models), etc.) may be packaged as a microservice—such as an inference microservice (e.g., NVIDIA NIMs)—which may include a container (e.g., an operating system (OS)-level virtualization package) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model “engine.” For example, the inference microservice can include the container itself and the model(s) (e.g., weights and biases). In some cases, such as when the machine learning model(s) 502 is small enough (e.g.,(Having a sufficiently small number of parameters), the model(s) can be enclosed within the container itself. In other examples—such as when the model(s) is / are large—the model(s) can be hosted / stored in the cloud (e.g., in a data center) and / or hosted / stored on-premises and / or at the edge (e.g., on a local server or local computing device, but outside the container). In such embodiments, the model(s) can be accessed via one or more APIs, such as REST APIs. As such, and in some embodiments, the machine learning model(s) described herein can be used as an inference microservice to accelerate the development of a model(s) on a cloud, data center, or edge computing system while ensuring data security.The inference microservice can include, for example: one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that deliver low latency and high throughput for production applications - such as NVIDIA's TensorRT) and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring).The machine learning model(s) described herein may be included as part of the microservice along with an accelerated infrastructure capable of being deployed and / or orchestrated with a single command and autoscaling with a container orchestration system on accelerated infrastructure (e.g., from a single device to data center scale). As such, the inference microservice may include the machine learning model(s) (optimized, for example, for high-performance inference), inference runtime software for executing the machine learning model(s) and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity verification, and / or other monitoring capabilities.In such embodiments, the inference microservice may include software for performing in-place replacement and / or updates of the machine learning model(s). During replacement or update, the software performing the replacement / update may maintain user configurations of the inference runtime software and enterprise management software.
[0073] In some embodiments, the systems and methods described herein can be used in a speech or smart kiosk application. For example, a kiosk, tablet, smart display, or other device may include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing the model, image database, etc.). In some embodiments, the kiosk / tablet / display may communicate (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) with one or more locally hosted servers / computing devices and / or with one or more remote servers / computing devices (e.g., in one or more data centers). In such examples, the kiosk may communicate with the machine learning model(s) (e.g.,the language model (LLM, VLM, MMLM, diffusion model, transformer model, NeRF, DNN, etc.) and / or the database hosted on local and / or remote servers using one or more APIs - such as, without limitation, REST APIs.
[0074] In one or more embodiments, the systems and methods described herein can be used in a gaming application. For example, a gaming console, PC, tablet, or other gaming device may include one or more onboard and / or remote processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing the game model, game assets, player data, etc.). These devices may use one or more machine learning models (e.g., diffusion models, transformer models, neural rendering field models (NERF models), language models (e.g., LLMs, SLMs, VLMs, MMLMs, etc.), DNNs, etc.) to enhance gameplay, generate dynamic, real-time content, and personalize user experiences based on in-game behavior or pre-stored player profiles. In some implementations, the systems can be used in a cloud gaming application (e.g.,NVIDIA's GeForce NOW) can be used. In such cases, a client device (e.g., a smart display, tablet, or gaming controller) can be used to interact with the game, while the machine learning model(s) and / or visual rendering may occur on one or more remote servers / computing devices (e.g., in one or more data centers). The language model, AI processing, and rendering described herein can run in the cloud, processing player input received from one or more end-user devices (e.g., based on controller, keyboard, mouse, joystick, AR / VR / MR / etc. input), generating appropriate in-game responses, and sending or transmitting the content to the end-user device(s).During the receipt and / or transmission of data to and from the end user or the edge device(s), one or more data processing units (DPUs) and / or network interface cards (NICs) may be used.
[0075] In some embodiments, the systems and methods described herein can be used in a videoconferencing application. For example, a videoconferencing device, such as a dedicated conference unit, computer, tablet, and / or smartphone, may include one or more onboard processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage (e.g., for storing video, audio, or other communication-related data). The system may use the machine learning model(s) (e.g., diffusion models, transformer models, neural rendering field models (NERF models), speech models (e.g., LLMs, SLMs, VLMs, MMLMs, etc.) to enhance videoconferencing functionality, including real-time or near-real-time transcription, diarization, speech translation, automatic speech recognition (ASR), and / or background noise reduction.In some implementations, the system can allow users to interact with the video conferencing platform using natural language input. For example, users can use voice commands to schedule, join, or leave meetings, or to manage participants and screen sharing. One or more data processing units (DPUs) and / or network interface cards (NICs) may be used while receiving and / or sending data to and from the end user or edge device(s).
[0076] In some embodiments, the systems and methods described herein can be used in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), hardware-based programmable vision accelerators (PVAs)) – which may include one or more vector processing units (VPUs), direct memory access (DMA) systems and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.) and memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models).The robotic system can use these processors to execute one or more machine learning models (e.g., language models) that allow it to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects or navigating environments using sensors such as cameras, LiDAR, radar, ultrasonic sensors, and more. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, radar, accelerometers) to create a comprehensive model of the robot's environment. This data can be processed locally on the robot or sent to remote servers for more computationally intensive tasks, such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g.,Sensor data, task status, or environmental conditions) are uploaded to the cloud, where centralized AI models can analyze optimized commands and distribute them to an entire fleet. In some embodiments, the machine learning model(s) described herein (e.g., language models, VLMs, SLMs, LLMs, MMLMs, diffusion models, NeRF models, DNNs, etc.) can be used to allow the robot to perceive and reflect on its environment and / or communicate with one or more other robots and / or one or more people in that environment. In some embodiments, the robot can communicate (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) with one or more locally hosted servers / computing devices and / or with one or more remote servers / computing devices (e.g.,communicate in one or more data centers.
[0077] In some embodiments, the systems and methods described herein can be used in an in-vehicle infotainment (IVI) system or in an in-cabin experience (IX) application. For example, the infotainment system in a vehicle (e.g., cars, trucks, drones, construction equipment, robots, semi-autonomous vehicles, or autonomous vehicles) may include one or more onboard processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), hardware-based programmable vision accelerators (PVAs)) – which may include one or more vector processing units (VPUs), direct memory access (DMA) systems, and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.The system may include a memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models) and a memory and / or storage (e.g., for storing entertainment content, navigation data, and user preferences). The system can use these processors to execute one or more machine learning models (e.g., language models) to enable features such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services via network connectivity. The in-vehicle infotainment system may also use natural language processing (NLP) models to enable voice-based interaction.The one or more machine learning models can be stored locally or accessed via one or more APIs connected to cloud services, enabling the system to process requests in real time or near real time.
[0078] In some embodiments, one or more transformer engines (TEs) can be implemented. The transformer engine can use micro-tensor scaling to optimize performance and accuracy—for example, to enable artificial intelligence processing with 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FP4) data. For instance, the transformer engine can combine 16-bit or 8-bit floating-point precision and an 8-bit or 4-bit floating-point data format with software algorithms to enhance AI performance and capabilities. By reducing mathematical operations to 8 bits or 4 bits, the TE allows for faster training of larger networks without compromising accuracy.For example, the TEs can include a library for accelerating transformer models on processing devices—such as GPUs—to deliver improved performance with reduced memory usage during both training and inference. When combined with other technologies, such as high-speed interlinking between nodes (e.g., using switches—such as NVLink switches) and Tensor Cores (enabling mixed-precision computing, such as microscale processing support), server clusters can be better equipped to train massive networks (e.g., billions of parameters) at high speeds. As such, Tensor Core processing of FP64, TF32, BF16, FP16, FP8, INT8, FP6, and FP4, as well as CUDA Core processing of FP64, FP32, FP16, and BF16, can be supported.
[0079] Although examples described herein may refer to the use of machine learning models, such as neural networks or language models, this is not intended to be restrictive. For example, and without limitation, various machine learning models and / or neural networks described herein may include any type of machine learning model, such as linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., ...).Neural autocoding networks, artificial neural networks (ANNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), perceptrons, long / short-term / memory (LSTM) networks, multi-layer perceptron (MLP) networks, deep-stacking networks (DSNs), generative pre-training (GPT) models or networks, feed-forward networks, radial basis function ANNs, self-organizing maps (SOMs), Kohonen maps, Hopfield networks, Boltzmann machine, neural deep-belief networks, deconvolutional neural networks, generative adversarial networks GANs), liquid-state machines, modular neural networks, liquid-state machines, sequence-to-sequence models, networks using transformer architectures, state space models (SSMs) (e.g.Networks using Mamba architectures (e.g., Mamba-1, Mamba 2, etc.), networks using selective state-space models, networks using structured state-space sequence models, etc.), diffusion models (e.g., probabilistic diffusion models, score-based generative models, etc.), Neural Radiation Field (NeRF) models, Gaussian splat models, Kolmogorov-Arnold networks (KANs), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), large action models (LAMs), vision-language-action (VLA) models, etc., are used / use, and / or include other types of machine learning models. INFERENCE AND TRAINING LOGIC
[0080] Fig. Figure 7A illustrates inference and / or training logic 715, which is used to perform inference and / or training operations associated with one or more embodiments.
[0081] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include a code and / or data storage device 701 to store forward and / or output weights and / or input / output data and / or other parameters for configuring neurons or layers of a neural network that is trained and / or used for inferencing in aspects by one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to a code and / or data storage device 701 to store graph code or other software for controlling the timing and / or sequencing into which weight and / or other parameter information for configuring the logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs) or simply circuits), is loaded.In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture from a neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 701 stores weight parameters and / or input / output data from each layer of a neural network that is trained or used in conjunction with one or more embodiments during the forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any part of the code and / or data storage 701 can be enclosed with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.
[0082] In at least one embodiment, any part of the code and / or data storage 701 can be located internally or externally of one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data storage 701 can be a cache memory, dynamic randomly addressable memory (DRAM), static randomly addressable memory (SRAM), non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether code and / or code and / or data storage 701 is, for example, internal or external to a processor or comprises a DRAM, SRAM, Flash or other storage type, may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.
[0083] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include a code and / or data storage 705 to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network that is trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 705 stores weight parameters and / or input / output data from each layer of a neural network that is trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, the training logic 715 can include or be coupled to a code and / or data storage 705 to store graph code or other software for controlling the timing and / or sequence, in which weight and / or other parameter information for configuring the logic, including integer and / or floating-point units (collectively arithmetic logic units (ALUs)), is loaded.
[0084] In at least one embodiment, code, such as graph code, causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which such code corresponds. In at least one embodiment, any part of the code and / or data storage 705 can be enclosed with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory. In at least one embodiment, any part of the code and / or data storage 705 can reside internally or externally to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 705 is, for example, internal or external to a processor or includes a DRAM, SRAM, Flash or other storage type, may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.
[0085] In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be separate storage structures. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be a combined storage structure. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be partially combined and partially separate. In at least one embodiment, any part of the code and / or data storage 701 and the code and / or data storage 705 can be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.
[0086] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include one or more arithmetic logic unit(s) (“ALU(s)”) 710, including integer and / or floating-point units, for performing logical and / or mathematical operations at least partly based on or specified by a training and / or inference code (e.g., graph code), wherein a result thereof may produce activations (e.g., output values of layers or neurons within a neural network), stored in an activation memory 720, which are functions of input / output and / or weight parameter data, stored in the code and / or data memory 701 and / or the code and / or data memory 705.In at least one embodiment, activations stored in the activation memory 720 are generated according to linear algebraic and / or matrix-based mathematics, performed by ALU(s) 710 in response to the execution of instructions or other code, wherein weight values stored in the code and / or data memory 705 and / or data memory 701 are used as operands together with other values, such as bias values, gradient information, momentum values or other parameters or hyperparameters, any or all of which may be stored in the code and / or data memory 705 or the code and / or data memory 701 or any other on- or off-chip memory.
[0087] In at least one embodiment, ALU(s) 710 are enclosed within one or more processors or other hardware logic devices or circuits, while in another embodiment, ALU(s) 710 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, ALU(s) 710 may be enclosed within execution units of a processor or otherwise within a series of ALUs accessible by execution units of a processor, either in the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, the code and / or data storage 701, the code and / or data storage 705, and the activation storage 720 can share a processor or other hardware logic device or circuit, while in another embodiment they can reside in different processors or other hardware logic devices or circuits, or a combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any part of the activation storage 720 can be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache or system memory.Furthermore, inference and / or training code can be stored with other code accessible to a processor or other hardware logic or circuitry and retrieved and / or processed using fetch, decode, scheduling, execution, sleep, and / or other logic circuitry of a processor.
[0088] In at least one embodiment, the activation memory 720 can be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or another type of storage. In at least one embodiment, the activation memory 720 can be located wholly or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation memory 720 is located, for example, internally or externally from a processor or comprises a DRAM, SRAM, flash, or other type of storage may depend on the available on-chip versus off-chip storage, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.
[0089] In at least one embodiment, the in Fig. 7A illustrated inference and / or training logic 715 in conjunction with an application-specific integrated circuit (“ASIC”), such as Google’s Tensorflow® Processing Unit, a Graphcore™ Inference Processing Unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake-Crest” processor). In at least one embodiment, the Fig. 7A illustrated inference and / or training logic 715 in conjunction with central processing unit hardware (“CPU” hardware), graphics processing unit hardware (“GPU” hardware) or other hardware, such as field programmable gate arrays (“FPGAs”).
[0090] Fig. Figure 7B illustrates inference and / or training logic 715 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 715 may, without limitation, include hardware logic in which computer resources are dedicated or otherwise used exclusively in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the Fig. 7B illustrated inference and / or training logic 715 in conjunction with an application-specific integrated circuit (“ASIC”), such as Google’s Tensorflow® Processing Unit, a Graphcore™ Inference Processing Unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake-Crest” processor). In at least one embodiment, the Fig. Figure 7B illustrates the use of inference and / or training logic 715 in conjunction with central processing unit hardware (“CPU” hardware), graphics processing unit hardware (“GPU” hardware), or other hardware, such as field programmable gate arrays (“FPGAs”). In at least one embodiment, the inference and / or training logic 715 includes, without limitation, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, Fig. In the embodiment illustrated in Figure 7B, the code and / or data storage 701 and the code and / or data storage 705 are each associated with a dedicated computer resource, such as computer hardware 702 and computer hardware 706, respectively. In at least one embodiment, the computer hardware 702 and the computer hardware 706 each comprise one or more ALUs that perform mathematical functions, such as linear algebraic functions, solely on information stored in the code and / or data storage 701 and the code and / or data storage 705, respectively, with the results being stored in the activation storage 720.
[0091] In at least one embodiment, the code and / or data storage 701 or 705 and the corresponding computer hardware 702 or 706 correspond to different layers of a neural network, such that a resulting activation from one storage / computer pair 701 / 702 provides the code and / or data storage 701 and computer hardware 702 as input for the next storage / computer pair 705 / 706, mirroring the conceptual organization of a neural network. In at least one embodiment, each of the storage / computer pairs 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computer pairs (not shown) can be included after or in parallel to the storage / computer pairs 701 / 702 and 705 / 706 in the inference and / or training logic 715. TRAINING AND USE OF A NEURAL NETWORK
[0092] Fig. Figure 8 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training dataset 802. In at least one embodiment, the training framework 804 is a PyTorch framework, while in other embodiments, the training framework 804 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 804 trains an untrained neural network 806 and enables it to be trained using processing resources described herein to generate a trained neural network 808. In at least one embodiment, weights can be selected arbitrarily or by pretraining using a deep belief network.In at least one embodiment, the training can be carried out in a supervised, partially supervised or unsupervised manner.
[0093] In at least one embodiment, an untrained neural network 806 is trained using supervised learning, wherein a training dataset 802 includes an input paired with a desired output, or wherein the training dataset 802 includes an input that has a known output, and an output of the neural network 806 is manually graded. In at least one embodiment, an untrained neural network 806 is trained in a supervised manner and processes inputs from a training dataset 802 and compares resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are subsequently propagated back through the untrained neural network 806. In at least one embodiment, the training framework 804 adjusts weights that control the untrained neural network 806.In at least one embodiment, the training framework 804 includes tools to monitor how well the untrained neural network 806 converges to a model, such as a trained neural network 808, capable of generating correct answers, such as result 814, based on input data, such as a new dataset 812. In at least one embodiment, the training framework 804 repeatedly trains an untrained neural network 806 while adjusting weights to refine an output of the untrained neural network 806 using a loss function and a fitting algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 804 trains an untrained neural network 806 until the untrained neural network 806 achieves a desired accuracy.In at least one embodiment, the trained neural network 808 can then be used to implement any number of machine learning operations.
[0094] In at least one embodiment, an untrained neural network 806 is trained using unsupervised learning, while the untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, the training dataset 802 for unsupervised learning includes input data without associated output data or "ground truth" data. In at least one embodiment, an untrained neural network 806 can learn groupings within the training dataset 802 and determine how individual inputs relate to the untrained dataset 802. In at least one embodiment, unsupervised training can be used to generate a self-organizing mapping in the trained neural network 808, enabling operations that are useful for reducing the dimensionality of a new dataset 812.In at least one embodiment, unattended training can also be used to perform irregularity detection, which allows the identification of data points in the new data set 812 that deviate from normal patterns of the new data set 812.
[0095] In at least one embodiment, semi-supervised learning can be used, which is a technique in which the training dataset 802 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 804 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables a trained neural network 808 to adapt to a new dataset 812 without forgetting knowledge that was introduced into the trained neural network 808 during the initial training.
[0096] With reference to Fig. 9 is Fig. Figure 9 shows an exemplary data flow diagram for a process 900 for generating and deploying a processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 900 can be used to perform game name recognition analysis and inference on user feedback data at one or more facilities 902, such as a data center.
[0097] In at least one embodiment, the process 900 can be executed within a training system 904 and / or a deployment system 906. In at least one embodiment, the training system 904 can be used to train, deploy, and execute machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in the deployment system 906. In at least one embodiment, the deployment system 906 can be configured to offload processing and computing resources to a distributed computing environment to reduce infrastructure requirements at the facility 902. In at least one embodiment, the deployment system 906 can provide a simplified platform for selecting, customizing, and implementing virtual tools for use with computing resources at the facility 902.In at least one embodiment, virtual instruments can include software-defined applications for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in a pipeline can use or call services (e.g., inference, visualization, computation, AI, etc.) of the deployment system 906 during application execution.
[0098] In at least one embodiment, some applications used in advanced processing and inference pipelines can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models can be trained on the facility 902 using feedback data 908 (such as imaging data) stored on the facility 902, or feedback data 908 from another facility or facilities, or a combination thereof. In at least one embodiment, the training system 904 can be used to provide applications, services, and / or other resources for generating functional, deployable machine learning models for the deployment system 906.
[0099] In at least one embodiment, a model register 924 can be backed up by object storage, which can support version and object metadata. In at least one embodiment, the object storage can be, for example, via cloud storage (e.g., a cloud 1026 of Fig. 10) be accessible in a cloud platform compatible with an application programming interface (API). In at least one embodiment, machine learning models within the model register 924 can be uploaded, listed, modified, or deleted by developers or partners of a system interacting with an API. In at least one embodiment, an API can provide access to procedures that allow users with appropriate permissions to associate models with applications, so that models can be executed as part of the execution of containerized instantiations of applications.
[0100] In at least one embodiment, a training pipeline 1004 ( Fig. 10) include a scenario in which an institution 902 trains its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, feedback data 908 can be received from various channels, such as forums, web forums, or the like. In at least one embodiment, after the feedback data 908 has been received, AI-assisted annotation 910 can be used to assist in generating annotations in accordance with the feedback data 908, which are to be used as ground-truth data for a machine learning model. In at least one embodiment, AI-assisted annotation 910 can include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations in accordance with certain types of feedback data 908 (e.g.,to generate AI-supported annotations 910 (of certain devices) and / or certain types of irregularities in feedback data 908. In at least one embodiment, AI-supported annotations 910 can then be used directly or adjusted or fine-tuned using an annotation tool to generate ground-truth data. In at least one embodiment, in some examples, labeled data 912 can be used as ground-truth data for training a machine learning model. In at least one embodiment, AI-supported annotations 910, labeled data 912, or a combination thereof can be used as ground-truth data for training a machine learning model, e.g., via model training 914 in . Fig. 9-10. In at least one embodiment, a trained machine learning model can be designated as the output model 916 and used by the deployment system 906 as described herein.
[0101] In at least one embodiment, the training pipeline 1004 ( Fig. 10) Include a scenario in which a facility 902 requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 906, but the facility 902 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, an existing machine learning model can be selected from a model register 924. In at least one embodiment, the model register 924 can include machine learning models that are trained to perform a range of different inference tasks on imaging data. In at least one embodiment, machine learning models in the model register 924 may be trained on imaging data from facilities that are different from the facility 902 (e.g.,remote facilities) have been trained. In at least one embodiment, machine learning models may be trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a specific location, which may be a form of feedback data 908, the training may take place at that location or at least be carried out in a manner that protects the confidentiality of imaging data or restricts the transfer of imaging data off the business premises (e.g., to comply with HIPAA regulations, data protection regulations, etc.). In at least one embodiment, after a model has been trained—or partially trained—at a location, a machine learning model may be added to the model register 924.In at least one embodiment, a machine learning model can then be retrained or updated on any number of other devices, and a retrained or updated model can be made available in the model register 924. In at least one embodiment, a machine learning model can then be selected from the model register 924—and designated as the output model 916—and can then be used in the deployment system 906 to perform one or more processing tasks for one or more applications of a deployment system.
[0102] In at least one embodiment, the training pipeline 1004 ( Fig. 10) in a scenario that includes a facility 902 which requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 906, wherein the facility 902 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, a machine learning model selected from the model register 924 may not be fine-tuned or optimized for feedback data 908 generated at the facility 902 due to differences in populations, genetic variations, robustness of training data used to train a machine learning model, diversity of irregularities in training data, and / or other problems with training data.In at least one embodiment, AI-assisted annotation 910 can be used to support the generation of annotations according to the feedback data 908, which are to be used as ground-truth data for retraining or updating a machine learning model. In at least one embodiment, labeled data 912 can be used as ground-truth data for training a machine learning model. In at least one embodiment, the retraining or updating of a machine learning model can be referred to as model training 914. In at least one embodiment, the model training 914—e.g., AI-assisted annotations 910, labeled data 912, or a combination thereof—can be used as ground-truth data for retraining or updating a machine learning model.
[0103] In at least one embodiment, the deployment system 906 can include software 918, services 920, hardware 922, and / or other components, features, and functionality. In at least one embodiment, the deployment system 906 can include a software "stack" such that software 918 can be built on top of services 920 and use services 920 to perform some or all of the processing tasks, and services 920 and software 918 can be built on top of hardware 922 and use hardware 922 to perform processing, storage, and / or other computing tasks of the deployment system 906.
[0104] In at least one embodiment, the software 918 can include any number of distinct containers, each container capable of executing an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks in an advanced processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, for each type of computing device, there can be any number of containers capable of performing a data processing task with respect to feedback data 908 (or other data types, such as those described herein).In at least one embodiment, an advanced processing and inference pipeline can be defined based on selections of different containers that are desired or required for processing feedback data 908, in addition to containers that receive and configure imaging data for use by each container and / or for use by a device 902 after processing by a pipeline (e.g., to convert outputs back into a usable data type for storage and display at the device 902). In at least one embodiment, a combination of containers within the software 918 (e.g., constituting a pipeline) can be referred to as a virtual instrument (as described in more detail herein), and a virtual instrument can utilize services 920 and hardware 922 to perform some or all of the processing tasks of applications instantiated in containers.
[0105] In at least one embodiment, data can undergo preprocessing as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, postprocessing can be performed on an output from one or more inference tasks or other processing tasks of a pipeline to prepare output data for a subsequent application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks can be performed by one or more machine learning models, such as trained or deployed neural networks, which may include output models 916 of the training system 904.
[0106] In at least one embodiment, tasks of the data processing pipeline can be encapsulated in one or more containers, each representing a discrete, fully functional instantiation of an application and virtualized computing environment capable of referencing machine learning models. In at least one embodiment, containers or applications can be published in a private area (e.g., a restricted access area) from a container register (described in more detail herein), and trained or deployed models can be stored in model register 924 and associated with one or more applications. In at least one embodiment, images of applications (e.g.,Container images) are available in a container registry, and after being selected by a user from a container registry for use in a pipeline, an image can be used to generate a container for instantiating an application for use by a user system.
[0107] In at least one embodiment, developers can develop, publish, and store applications (e.g., as containers) for performing processing and / or inference on supplied data. In at least one embodiment, the development, publication, and / or storage can be performed using a software development kit (SDK) associated with a system (e.g., to ensure that a developed application and / or container is consistent or compatible with a system). In at least one embodiment, an application being developed can be deployed locally (e.g., on a first facility, to data from a first facility) as a system (e.g., architecture 1000 of) using an SDK that can support at least some services 920. Fig. 10) be tested. In at least one embodiment, an application, after it has been validated by Architecture 1000 (e.g., with regard to accuracy, etc.), can be made available in a container register for selection and / or execution by a user (e.g., a hospital, clinic, laboratory, healthcare provider, etc.) to perform one or more processing tasks relating to data at a user's facility (e.g., a second facility).
[0108] In at least one embodiment, developers can then make applications or containers available over a network for access and use by users of a system (e.g., Architecture 1000 of Fig. 10) share. In at least one embodiment, completed and validated applications or containers can be stored in a container register, and associated machine learning models can be stored in model register 924. In at least one embodiment, a requesting entity providing an inference or image processing request can search a container register and / or model register 924 for an application, container, dataset, machine learning model, etc., select a desired combination of elements for inclusion in a data processing pipeline, and send a processing request. In at least one embodiment, a request can include input data necessary to execute a request and / or can include a selection of applications and / or machine learning models to be executed when processing a request.In at least one embodiment, a request can be forwarded to one or more components of the deployment system 906 (e.g., a cloud) to perform processing in the data processing pipeline. In at least one embodiment, the processing by the deployment system 906 can include referencing selected elements (e.g., applications, containers, models, etc.) from a container register and / or model register 924. In at least one embodiment, after results have been generated by a pipeline, results can be returned to a user for reference (e.g., for viewing in a viewing application suite running on a local on-premises workstation or end device).
[0109] In at least one embodiment, services 920 can be used to support the processing or execution of applications or containers in pipelines. In at least one embodiment, services 920 can include computing services, collaborative content creation services, simulation services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, services 920 can provide functionality common to one or more applications in software 918, so that functionality can be abstracted to a service that can be called or used by applications. In at least one embodiment, functionality provided by services 920 can run dynamically and more efficiently while also scaling well by allowing applications to process data in parallel, for example, using a parallel computing platform 1030. Fig. 10) In at least one embodiment, a Service 920 can be shared between and by different applications, instead of each application sharing the same functionality offered by a Service 920 requiring its own instance of the Service 920. In at least one embodiment, services can include an inference server or inference engine that can be used as non-restrictive examples for performing detection or segmentation tasks. In at least one embodiment, a model training service can be included that can provide machine learning model training and / or retraining capabilities.
[0110] In at least one embodiment, where a service 920 includes an AI service (e.g., an inference service), one or more machine learning models associated with an irregularity detection application (e.g., tumors, growth abnormalities, scarring, etc.) can be executed by calling an inference service (e.g., an inference server) (e.g., as an API call) to execute one or more machine learning models, or processing thereof, as part of application execution. In at least one embodiment, if another application includes one or more machine learning models for segmentation tasks, an application can call an inference service to execute machine learning models to perform one or more processing operations associated with segmentation tasks.In at least one embodiment, software 918, which implements an advanced processing and inference pipeline, can be simplified because each application can call the same inference service to perform one or more inference tasks.
[0111] In at least one embodiment, Hardware 922 can include GPUs, CPUs, graphics cards, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX™ supercomputer system), a cloud platform, or a combination thereof. In at least one embodiment, different types of Hardware 922 can be used to provide efficient, purpose-built support for Software 918 and Services 920 in the deployment system 906. In at least one embodiment, the use of GPU processing can be implemented for local processing (e.g., at the facility 902), within an AI / deep learning system, in a cloud system, and / or in other processing components of the deployment system 906 to improve the efficiency, accuracy, and effectiveness of game name recognition.
[0112] In at least one embodiment, software 918 and / or services 920 can be optimized for GPU processing with respect to deep learning, machine learning, and / or high-performance computing, simulation, and visual computing, as non-limiting examples. In at least one embodiment, at least a portion of the computing environment of the deployment system 906 and / or the training system 904 can be run in a data center or on one or more supercomputers or high-performance computing systems, with GPU-optimized software (e.g., a hardware and software combination of NVIDIA's DGX™ system). In at least one embodiment, hardware 922 can include any number of GPUs that can be called upon to perform the parallel processing of data as described herein.In at least one embodiment, a cloud platform can further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computational tasks. In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC™) can run as a hardware abstraction and scaling platform using one or more AI / deep learning supercomputers and / or GPU-optimized software (e.g., as deployed on NVIDIA's DGX™ systems). In at least one embodiment, the cloud platform can integrate an application container clustering or orchestration system (e.g., Kubernetes) across multiple GPUs to enable seamless scaling and load balancing.
[0113] Fig. Figure 10 is a system representation for an exemplary architecture 1000 for generating and deploying a deployment pipeline according to at least one embodiment. In at least one embodiment, the architecture 1000 can be used to perform the process 900 of Fig. 9 and / or other processes that include advanced processing and inference pipelines. In at least one embodiment, the architecture 1000 may include a training system 904 and a deployment system 906. In at least one embodiment, the training system 904 and the deployment system 906 may be implemented using the software 918, services 920, and / or hardware 922 as described herein.
[0114] In at least one embodiment, Architecture 1000 (e.g., Training System 904 and / or Deployment System 906) can be implemented in a cloud computing environment (e.g., using Cloud 1026). In at least one embodiment, Architecture 1000 can be implemented locally with respect to a facility or as a combination of cloud and local computing resources. In at least one embodiment, access to APIs in Cloud 1026 can be restricted to authorized users by means of established security measures or protocols. In at least one embodiment, a security protocol can include web tokens that can be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and can include corresponding authorization.In at least one embodiment, APIs of virtual instruments (described herein), or other instantiations of Architecture 1000, can be restricted to a set of public IPs (Internet Service Providers, ISPs) that have been audited or authorized for interaction.
[0115] In at least one embodiment, various components of Architecture 1000 can communicate between themselves and with each other using a range of different network types, including, but not limited to, local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between devices and components of Architecture 1000 (e.g., for transmitting inference requests, receiving results of inference requests, etc.) can be conducted via a data bus or data buses, wireless data protocols (WiFi), wired data protocols (e.g., Ethernet), etc.
[0116] In at least one embodiment, the training system 904 can include training pipelines 1004, similar to those described herein with respect to Fig. 9 described. In at least one embodiment, training pipelines 1004 can be used when one or more machine learning models in deployment pipelines 1010 are to be used by the deployment system 906 to train or retrain one or more (e.g., pre-trained) models and / or to implement one or more of the pre-trained models 1006 (e.g., without the need for retraining or updating). In at least one embodiment, one or more output models 916 can be generated as a result of the training pipelines 1004.In at least one embodiment, the training pipelines 1004 can include any number of processing steps, AI-assisted annotation 910, labeling or annotating feedback data 908 to generate labeled data 912, model selection from a model register, model training 914, training, retraining, or updating models, and / or other processing steps. In at least one embodiment, different training pipelines 1004 can be used for different machine learning models employed by the deployment system 906. In at least one embodiment, a training pipeline 1004, similar to a first one with respect to . Fig. The example described in 9, which can be used for a first machine learning model, can be a training pipeline 1004, similar to a second one in terms of Fig. The example described in 9 can be used for a second machine learning model, and a training pipeline 1004 can be similar to a third in terms of Fig. The example described in Section 9 can be used for a third machine learning model. In at least one embodiment, any combination of tasks within a training system 904 can be used, depending on what is required for each machine learning model. In at least one embodiment, one or more machine learning models can already be trained and ready for use, so that machine learning models may not require any processing by the training system 904 and can be implemented by the deployment system 906.
[0117] In at least one embodiment, the output model(s) 916 and / or pretrained model(s) 1006 can include any type of machine learning model, depending on the embodiment. In at least one embodiment and without limitation, machine learning models used by Architecture 1000 may include machine learning model(s) that employ linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., auto-encoders, convolutional, recurrent, perceptrons, long / short-term / memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep-belief, deconvolutional, generative adversarial, liquid-state machine, etc.), and / or other types of machine learning models.
[0118] In at least one embodiment, training pipelines can include AI-assisted annotation. In at least one embodiment, labeled data (e.g., traditional annotation) can be generated by a variety of techniques. In at least one embodiment, labels or other annotations can be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of program suitable for generating annotations or labels for ground truth, and / or, in some examples, drawn by hand. In at least one embodiment, ground truth data can be produced as follows: synthetically (e.g., generated from computer models or renderings), real (e.g., conceived and produced from real-world data), machine-automated (e.g.,using feature analysis and learning to extract features from data and then generate labels), annotated by a human (e.g., a labeler or annotation expert defines the location of the labels), and / or a combination thereof. In at least one embodiment, for each instance of feedback data 908 (or a type of data used by machine learning models), there can be corresponding ground-truth data generated by the training system 904. In at least one embodiment, AI-assisted annotation can be performed as part of deployment pipelines 1010; either in addition to or instead of AI-assisted annotation included in training pipelines 1004. In at least one embodiment, the architecture 1000 can include a multi-layered platform that includes a software layer (e.g.,Software 918) for diagnostic applications (or other application types) that can perform one or more medical imaging and diagnostic functions.
[0119] In at least one embodiment, a software layer can be implemented as a secure, encrypted, and / or authenticated API through which applications or containers from external environments (e.g., facility 902) can be invoked (e.g., called). In at least one embodiment, applications can then call or execute one or more services 920 to perform computational, AI, or visualization tasks associated with the respective applications, and the software 918 and / or services 920 can utilize hardware 922 to perform processing tasks in an effective and efficient manner.
[0120] In at least one embodiment, the deployment system 906 can execute deployment pipelines 1010. In at least one embodiment, the deployment pipelines 1010 can include any number of applications that can be applied sequentially, non-sequentially, or otherwise to feedback data (and / or other data types), including AI-assisted annotation as described above. In at least one embodiment, a deployment pipeline 1010, as described herein, can be designated as a virtual instrument for an individual device. In at least one embodiment, there can be more than one deployment pipeline 1010 for a single device, depending on the information desired from data generated by the device.
[0121] In at least one embodiment, applications available for deployment pipelines 1010 can include any application that can be used to perform processing tasks on feedback data or other device data. In at least one embodiment, since different applications in some embodiments share common image operations, a data extension library (e.g., as one of services 920) can be used to accelerate these operations. In at least one embodiment, to avoid bottlenecks of conventional processing approaches that rely on CPU processing, a parallel computing platform 1030 can be used to GPU-accelerate these processing tasks.
[0122] In at least one embodiment, the deployment system 906 may include a user interface (UI) 1014 (e.g., a graphical user interface, a web interface, etc.) that can be used to select applications for inclusion in deployment pipelines 1010, to order applications, to modify or change applications or parameters or constructs thereof, to use and interact with deployment pipelines 1010 during setup and / or deployment, and / or to otherwise interact with the deployment system 906. In at least one embodiment, the UI 1014 (or another user interface), although not illustrated with respect to the training system 904, may be used to select models for use in the deployment system 906, to select models for training or retraining in the training system 904, and / or to otherwise interact with the training system 904.In at least one embodiment, the training system 904 and the deployment system 906 can include DICOM adapters 1002A and 1002B.
[0123] In at least one embodiment, a pipeline manager 1012 can be used, in addition to an application orchestration system 1028, to manage the interaction between applications or containers of deployment pipelines 1010 and services 920 and / or hardware 922. In at least one embodiment, the pipeline manager 1012 can be configured to allow application-to-application, application-to-service 920, and / or application-or-service-to-hardware 922 interactions. In at least one embodiment, the pipeline manager 1012, although illustrated as being included in the software 918 (which is not to be interpreted as limiting), can in some examples be included in services 920. In at least one embodiment, the application orchestration system 1028 (e.g., Kubernetes, Docker, etc.) can be configured to allow application-to-application, application-to-service 920, and / or application-to-service 922 interactions.) include a container orchestration system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, each application can be run in a self-contained environment (e.g., at the kernel level) by associating applications from deployment pipelines 1010 (e.g., a reconstruction application, a segmentation application, etc.) with individual containers to increase speed and efficiency.
[0124] In at least one embodiment, each application and / or container (or each image thereof) can be developed, modified, and deployed individually (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separately from the first user or developer). This allows the focus and attention to be directed to a task by a single application and / or container without being hindered by tasks performed by other applications or containers. In at least one embodiment, communication and cooperation between different containers or applications can be supported by the pipeline manager 1012 and the application orchestration system 1028.In at least one embodiment, the application orchestration system 1028 and / or the pipeline manager 1012 can enable communication among and between each container or application, and resource sharing among and between each of the applications or containers, as long as an expected input and / or output from each container or application is known to the system (e.g., based on constructs of applications or containers). In at least one embodiment, since one or more applications or containers in deployment pipelines 1010 can share the same services and resources, the application orchestration system 1028 can orchestrate, balance, and determine the load for sharing services or resources between and among different applications or containers.In at least one embodiment, a scheduler can be used to track resource requests from applications or containers, the current or planned use of these resources, and resource availability. In at least one embodiment, the scheduler can therefore allocate resources to different applications and distribute resources between and among applications with respect to the requirements and availability of a system. In some examples, the scheduler (and / or another component of the application orchestration system 1028) can determine resource availability and distribution based on constraints imposed on a system (e.g., user constraints), such as Quality of Service (QoS), urgency of data output needs (e.g., to determine whether to perform real-time or delayed processing), etc.
[0125] In at least one embodiment, services 920, which are used and shared by applications or containers in the deployment system 906, can include compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, and / or other service types. In at least one embodiment, applications can call (e.g., execute) one or more services 920 to perform processing operations for an application. In at least one embodiment, compute services 1016 can be used by applications to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1016 can be used to perform parallel processing (e.g.,using a parallel computing platform 1030) to process data across one or more applications and / or to perform one or more tasks of a single application, essentially simultaneously. In at least one embodiment, the parallel computing platform 1030 (e.g., NVIDIA's CUDA) can enable general-purpose computing on GPUs (GPGPU) (e.g., GPUs 1022). In at least one embodiment, a software layer of the parallel computing platform 1030 can provide access to virtual instruction sets and parallel computing elements of GPUs for executing computing kernels. In at least one embodiment, the parallel computing platform 1030 can include memory, and in some embodiments, memory can be shared between and among multiple containers and / or between and among different processing tasks within a single container.In at least one embodiment, inter-process communication (IPC) calls can be generated for multiple containers and / or for multiple processes within a container to use the same data from a shared segment of the memory of the Parallel Computing Platform 1030 (e.g., when multiple different stages of an application or multiple applications process the same information). In at least one embodiment, instead of copying data and moving it to different locations in memory (e.g., a read / write operation), the same data can be used at the same memory location for any number of processing tasks (e.g., at the same time, at different times, etc.).In at least one embodiment, when data is used to generate new data as a result of processing, this information can be stored in a new location and shared between different applications. In at least one embodiment, the location of data and a location of updated or modified data can be part of a definition of how a payload is understood within containers.
[0126] In at least one embodiment, AI services 1018 can be used to perform inference services for executing machine learning models associated with applications (e.g., those tasked with performing one or more processing tasks of an application). In at least one embodiment, AI services 1018 can utilize the AI system 1024 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, applications of deployment pipelines 1010 can use one or more output models 916 from the training system 904 and / or other application models to perform inference on imaging data (e.g., DICOM data, RIS data, CIS data, REST-compliant data, RPC data, raw data, etc.).In at least one embodiment, two or more examples of inference using the application orchestration system 1028 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can achieve higher quality-of-service agreements, such as performing inference on urgent requests in emergencies or for a radiologist during diagnosis. In at least one embodiment, a second category may include a standard-priority path that can be used for requests that may not be urgent or when the analysis can be performed at a later time. In at least one embodiment, the application orchestration system 1028 may distribute resources (e.g., services 920 and / or hardware 922) based on priority paths for different inference tasks of AI services 1018.
[0127] In at least one embodiment, shared storage can be connected to AI services 1018 within architecture 1000. In at least one embodiment, shared storage can be operated as a cache (or other storage device type) and used to process inference requests from applications. In at least one embodiment, when an inference request is sent, a request can be received by a set of API instances of the deployment system 906, and one or more instances can be selected (e.g., for best fit, load balancing, etc.) to process a request.In at least one embodiment, to process a request, a request can be entered into a database, a machine learning model can be found from a model register 924 if it is not already in a cache, a validation step can ensure that the appropriate machine learning model is loaded into a cache (e.g., shared storage), and / or a copy of a model can be stored in a cache. In at least one embodiment, the scheduler (e.g., a pipeline manager 1012) can be used to start an application referenced in a request if no application is already running or if there are not enough instances of an application. In at least one embodiment, an inference server can be started if no inference server has already been started to execute a model. In at least one embodiment, a set of inference servers can be started for each model.In at least one embodiment, models can be cached in a pull model where inference servers are clustered, if load balancing is advantageous. In at least one embodiment, inference servers can be statically loaded on corresponding distributed servers.
[0128] In at least one embodiment, inference can be performed using an inference server running in a container. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally a plurality of versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, a model can be submitted to an inference server when it is started, so that the same container can be used to serve different models, as long as the inference server runs as a different instance.
[0129] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., hosting an instance of an inference server) can be loaded (if not already loaded), and a start operation can be invoked. In at least one embodiment, preprocessing logic within a container can load, decode, and / or perform other additional preprocessing on incoming data (e.g., using a CPU(s) and / or GPU(s)). In at least one embodiment, after data has been prepared for inference, a container can perform inference on data as needed. In at least one embodiment, this can involve a single inference call on a single image (e.g., a hand X-ray) or it can require inference on hundreds of images (e.g., a breast CT scan).In at least one embodiment, an application can summarize results before completion, which can include, without limitation, generating a single confidence score, pixel-level segmentation, voxel-level segmentation, generating a visualization, or generating text to summarize findings. In at least one embodiment, different models or applications can be assigned different priorities. For example, some models can have a real-time priority (turnaround time less than one minute), while others can have a lower priority (e.g., turnaround time less than 10 minutes). In at least one embodiment, model execution times can be measured by a requesting institution or entity and can include partner network traversal time as well as execution on an inference service.
[0130] In at least one embodiment, the transfer of requests between Services 920 and inference applications can be hidden behind a software development kit (SDK), and robust transport through a queue can be provided. In at least one embodiment, a request is placed in a queue via an API for a single application / tenant ID combination, and an SDK retrieves a request from the queue and passes it to an application. In at least one embodiment, a queue name can be provided in an environment from which an SDK retrieves the request. In at least one embodiment, asynchronous communication through a queue can be beneficial, as it allows each instance of an application to pick up work as it becomes available.In at least one embodiment, results can be transferred back via a queue to ensure that no data is lost. In at least one embodiment, queues can also provide the ability to segment work, with highest-priority work going to a queue with the most associated instances of an application, while lowest-priority work going to a queue with a single associated instance that processes tasks in the order they were received. In at least one embodiment, an application can run on a GPU-accelerated instance generated in Cloud 1026, and an inference service can perform inference on a GPU.
[0131] In at least one embodiment, visualization services 1020 can be used to generate visualizations for viewing application outputs and / or deployment pipelines 1010. In at least one embodiment, GPUs 1022 can be used by visualization services 1020 to generate visualizations. In at least one embodiment, rendering effects, such as ray tracing or other light transport simulation techniques, can be implemented by visualization services 1020 to generate higher-quality visualizations. In at least one embodiment, visualizations can include, without limitation, 2D image renderings, 3D volume renderings, 3D volume reconstruction, 2D tomography slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, virtualized environments can be used to provide a virtual interactive display or environment (e.g.,to generate a virtual environment for interaction by users of a system (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, visualization services 1020 may include an internal visualizer, cinematography, and / or other rendering or image processing capabilities or functionality (e.g., ray tracing, rasterization, internal optics, etc.).
[0132] In at least one embodiment, Hardware 922 can include GPUs 1022, AI System 1024, Cloud 1026, and / or other hardware used to run the Training System 904 and / or Deployment System 906. In at least one embodiment, GPUs 1022 (e.g., NVIDIA's TESLA) can include GPUs 1022 (e.g., NVIDIA's TESLA) ® and / or QUADRO ®GPUs) include any number of GPUs used to perform processing tasks of compute services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, other services, and / or any features or functionality of software 918. For example, with respect to AI services 1018, GPUs 1022 may be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on outputs of machine learning models, and / or to perform inference (e.g., to run machine learning models). In at least one embodiment, the cloud 1026, the AI system 1024, and / or other components of the architecture 1000 may use GPUs 1022. In at least one embodiment, the Cloud 1026 can include a GPU-optimized platform for deep learning tasks.In at least one embodiment, the AI system can use 1024 GPUs, and the cloud 1026—or at least a part intended for deep learning or inference—can be run using one or more AI systems 1024. As such, although Hardware 922 is illustrated as discrete components, this is not to be interpreted as restrictive, and any components of Hardware 922 can be combined with or used by other components of Hardware 922.
[0133] In at least one embodiment, the AI system 1024 can include a dedicated computing system (e.g., a supercomputer or an HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, the AI system 1024 (e.g., NVIDIA's DGX™) can include GPU-optimized software (e.g., a software stack) that can be run using a variety of GPUs 1022, in addition to CPUs, RAM, storage, and / or other components, features, or functionality. In at least one embodiment, one or more AI systems 1024 can be deployed in the cloud 1026 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the architecture 1000.
[0134] In at least one embodiment, the Cloud 1026 can include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC™) that can provide a GPU-optimized platform for performing processing tasks of architecture 1000. In at least one embodiment, the Cloud 1026 can include an AI system (or systems) 1024 for performing one or more AI-based tasks of architecture 1000 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the Cloud 1026 can be integrated into an application orchestration system 1028 utilizing multiple GPUs to enable seamless scaling and load balancing between and among applications and services 920. In at least one embodiment, the Cloud 1026 can be provided for running at least some of the services 920 of the architecture 1000, including computing services 1016, AI services 1018 and / or visualization services 1020, as described herein.In at least one embodiment, the Cloud 1026 can perform small and large batch inference (e.g., running NVIDIA's TensorRT™), an accelerated parallel computing API and platform 1030 (e.g., NVIDIA's CUDA). ® ) provide, run an application orchestration system 1028 (e.g. KUBERNETES), provide a graphics rendering API and platform (e.g. for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher quality cinematography), and / or provide other functionality for architecture 1000.
[0135] In at least one embodiment, the Cloud 1026, in an effort to maintain patient confidentiality (e.g., when patient data or records are to be used externally), can include a register, such as a deep learning container register. In at least one embodiment, a register can store containers for instantiations of applications that can perform preprocessing, postprocessing, or other processing tasks on patient data. In at least one embodiment, the Cloud 1026 can receive data that includes both patient data and sensor data in containers, perform requested processing only on sensor data in these containers, and then provide a resulting output and / or visualizations to appropriate parties and / or devices (e.g.,(Medical on-premises devices used for visualization or diagnosis) are transmitted, all without needing to extract, store, or otherwise access patient data. In at least one embodiment, the confidentiality of patient data is maintained in accordance with HIPAA and / or other data regulations. EXEMPLARY LANGUAGE MODELS
[0136] In at least some embodiments, language models, such as large language models (LLMs) and / or other types of generative artificial intelligence (AI), can be implemented. These models may be capable of understanding, summarizing, translating, and / or otherwise generating text (e.g., natural language text, code, etc.), images, video, computer-aided design (CAD) assets, omniverse and / or metaverse file information (e.g., in USD format), and / or the like, based on the context provided in input prompts or queries. These language models may be considered "large" in embodiments based on the fact that the models are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—such as millions or billions of parameters. The LLMs / VLMs / etc.These LLMs can be implemented for summarizing textual data, analyzing and extracting insights from data (e.g., text, image, video, etc.), and generating new text / image / video / etc. in user-specified styles, tones, or formats. The LLMs of this disclosure can be used exclusively for text processing in some embodiments, while in other embodiments, multimodal LLMs can be implemented to accept, understand, and / or generate text along with other types of content, such as images, audio, and / or video. For example, visual language models (VLMs), or more generally, multimodal language models, can be implemented to accept image, video, audio, textual, 3D design (e.g., CAD), and / or other input data types, and / or to generate or output image, video, audio, textual, 3D design, and / or other output data types.
[0137] Different types of LLM / VLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented that employ different techniques for understanding and generating outputs such as text, audio, video, images, etc. In some embodiments, LLM architectures such as recurrent neural networks (RNNs) or long short-term memory networks (LSTMs) can be used, while in other embodiments, transformer architectures—such as those that depend on self-attention mechanisms—are used to understand and recognize relationships between words or tokens. The language models of this disclosure can include one or more encoder and / or decoder blocks.For example, discriminative or encoder-only LLMs like BERT (Bidirectional Encoder Representations from Transformers) can be implemented for tasks involving language comprehension, such as classification, sentiment analysis, question answering, and named entity recognition. As another example, generative or decoder-only LLMs like GPT (Generative Pretrained Transformer) can be implemented for tasks involving both language and content generation, such as text completion, story generation, and dialogue generation. LLMs that include both encoder and decoder components, such as T5 (Text-to-Text Transformer), can be implemented to understand and generate content, such as for translation and summarization.These examples are not intended to be limiting and any type of architecture – including, but not limited to, those described herein – may be implemented depending on the specific embodiment and the task(s) being performed using the model(s).
[0138] In various embodiments, LLMs / VLMs / etc. can be trained using unsupervised learning, where an LLM learns patterns from large amounts of unlabeled text / audio / video / image / etc. data. Due to the comprehensive training in these embodiments, the models may not require task-specific or domain-specific training. LLMs that have undergone extensive pretraining with vast amounts of unlabeled text data can be referred to as foundational models and may be capable of a range of tasks, such as answering questions, summarizing, filling in missing information, and translating. Some LLMs can be tailored to a specific use case using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), and adding adapters (e.g.,custom neural networks and / or neural network layers that tune or adapt prompts or tokens to direct the language model to a specific task or domain), and / or using other fine-tuning or tailoring techniques that optimize the models for use in specific tasks and / or within specific domains.
[0139] In some embodiments, the LLMs / VLMs / etc. of the present disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify impermissible or undesired inputs (e.g., prompts) and / or outputs of the model. In some non-limiting embodiments, the implemented guardrails can be similar to those described in U.S. Patent Application No. 18,304,341, filed on April 20, 2023, the contents of which are incorporated herein by reference in their entirety. In some embodiments, one or more additional models—or layers thereof—can be identified to identify problems with the inputs and / or outputs of the models.For example, these "safeguard" models can be trained to identify inputs and / or outputs that are "safe" or otherwise acceptable or desired, and / or that are "unsafe" or otherwise undesirable for the given application / execution mode. As a result, the LLMs / VLMs / etc. of this disclosure will be less likely to output speech / text / audio / etc. that is offensive, vulgar, impermissible, unsafe, out-of-domain, and / or otherwise undesirable for the given application / execution mode.
[0140] In some embodiments, the LLMs / VLMs / etc. may be configured or capable of accessing and using one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations for which the model is not ideally suited, instructions may be provided (e.g., as a result of training and / or based on instructions in a given prompt) to access one or more plug-ins (e.g., third-party plug-ins) to assist in processing the current input. In such an example, where at least part of a prompt relates to restaurants or weather, the model may access one or more restaurant or weather plug-ins (e.g., via one or more APIs) to retrieve the relevant information.As another example where at least part of a response requires a mathematical calculation, the model can access one or more math plugins or APIs to help solve the problem(s) and then use the response from the plugin and / or API in the model's output. This process can be repeated—e.g., recursively—for any number of iterations and using any number of plugins and / or APIs until a response to the input prompt can be generated that addresses each request / question / requirement / process / operation / etc. As such, the model(s) may depend not only on its own knowledge gained from training on a large dataset, but also on the know-how or optimized nature of one or more external resources—such as APIs, plugins, and / or the like.
[0141] Fig. Figure 11A is a block representation of an exemplary generative language model system 1100, which is suitable for use in implementing at least some embodiments of the present disclosure. In the Fig. In the illustrated example 11A, the generative language model system 1100 includes a request-extended generation component (RAG component) 1192, an input processor 1105, a tokenizer 1110, an embedding component 1120, plug-ins / APIs 1195 and a generative language model (LM) 1130 (which may include an LLM, an SLM, a VLM, a multimodal LM, etc.).
[0142] At a high level, the input processor 1105 can receive an input 1101 that includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Descriptor (USD) data – such as OpenUSD, etc.), depending on the architecture of the generative LM 1130 (e.g., LLM / SLM / VLM / MMLM / etc.). In some embodiments, the input 1101 can include simple text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 1101 can include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML).In some embodiments where the generative LM 1130 is capable of processing multimodal inputs, the input 1101 can combine text with image data, audio data, video data, design data, USD data, and / or other types of input data, such as, but not limited to, those described herein (or it can omit text). Taking raw input text as an example, the input processor 1105 can prepare raw input text in various ways. For example, the input processor 1105 can perform various types of text filtering to remove noise (e.g., special characters, punctuation, HTML tags, stop words, parts of an image (or images), parts of audio, etc.) from relevant textual content.In an example involving stop words (generic words with little semantic meaning), the Input Processor 1105 can remove stop words to reduce noise and allow the generative LM 1130 to focus on more meaningful content. The Input Processor 1105 can also apply text normalization, such as converting all characters to lowercase, removing accents, and / or handling special cases like abbreviations or suffixes to ensure consistency. These are just a few examples, and other types of input processing can be applied.
[0143] In some embodiments, a RAG component 1192 (which may include one or more RAG models and / or be implemented using the generative LM 1130 itself) can be used to retrieve additional information to be used as part of the input 1101 or the prompt. RAG can be used to enhance the input for the LLM / SLM / VLM / MMLM / etc. with external knowledge, making answers to specific questions, queries, or requirements more relevant—for example, in a case where specific knowledge is required. The RAG component 1192 can retrieve this additional information (e.g., basic information such as basic text / image / video / audio / USD / CAD / etc.) from one or more external sources, which can then be fed to the LLM / SLM / VLM / MMLM / etc. along with the prompt to improve the accuracy of the model's responses or outputs.
[0144] For example, in some embodiments, the input 1101 can be generated using the query or input to the model (e.g., a question, a request, etc.) in addition to data retrieved using the RAG component 1192. In some embodiments, the input processor 1105 can analyze the input 1101 and communicate with the RAG component 1192 (or the RAG component 1192 can, in embodiments, be part of the input processor 1105) to identify relevant text and / or other data in order to provide this to the generative LM 1130 as additional context or additional sources of information in order to identify the reaction, response, or output 1190, in general.For example, if the input indicates that the user is interested in a desired tire pressure for a specific make and model of vehicle, the RAG component 1192—for example, using a RAG model that performs a vector search in an embedding space—can retrieve the tire pressure information or the corresponding text from a digital (embedded) version of the user manual for that specific vehicle make and model. Similarly, if a user revisits a chatbot regarding a specific product offering or service, the RAG component 1192 can retrieve a previously saved conversation history—or at least a summary of it—and include the previous conversation history, along with the current task / request, as part of the input 1101 in the generative LM 1130.
[0145] The RAG component 1192 can employ various RAG techniques. For example, naive RAG can be used when documents are indexed, bundled, and applied to an embedding model to generate embeddings that correspond to the bundles. A user query can also be applied to the embedding model and / or another embedding model of the RAG component 1192, and the bundle embeddings, along with the query embeddings, can be compared to identify the embeddings most similar to / corresponding to the query. These embeddings can then be provided to the generative LM 1130 to generate output.
[0146] In some implementations, more advanced RAG techniques can be used. For example, before bundles are passed to the embedding model, the bundles can be subjected to pre-fetch processes (e.g., routing, rewriting, metadata analysis, expansion, etc.). Furthermore, before generating the final embeddings, post-fetch processes (e.g., reordering, prompt compression, etc.) can be performed on the outputs of the embedding model, before the final embeddings are used for comparison with an input query.
[0147] Another example is modular RAG techniques, such as those similar to naive and / or advanced RAG, but which also include features such as hybrid search, recursive retrieval and query engines, step-back approaches, subqueries and hypothetical document embedding.
[0148] As another example, a graph RAG can use knowledge graphs as a source of contextual or factual information. The graph RAG can be implemented using a graph database as a source of contextual information sent to the LLM / SLM / VLM / MMLM / etc. Instead of providing the model with bundles of data extracted from larger documents—which can result in a lack of context, factual accuracy, linguistic precision, etc.—the graph RAG can provide structured entity information to the LLM / SLM / VLM / MMLM / etc. by combining the structured textual entity description with its many properties and relationships, allowing for a deeper understanding on the part of the model.When the Graph-RAG is implemented, the systems and procedures described herein use a graph as a content store, extracting relevant bundles of documents and requesting the LLM / SLM / VLM / MMLM / etc. to respond using these. In such embodiments, the knowledge graph can contain relevant textual content and metadata about the knowledge graph and can be integrated into a vector database. In some embodiments, the Graph-RAG can use a graph as a domain expert, extracting descriptions of concepts and entities relevant to a query / prompt and passing them to the model as semantic context. These descriptions can include relationships between the concepts.In other examples, the graph can be used as a database, where part of a query / prompt can be mapped to a graph query, the graph query can be executed, and the LLM / SLM / VLM / MMLM / etc. can summarize the results. In such an example, the graph can store relevant factual information, and a query (natural language query) can be sent to the graph query tool (NL-to-Graph-query tool), along with entity linking. In some implementations, graph RAG (e.g., using a graph database) can be combined with a standard (e.g., vector database) RAG and / or other RAG types to benefit from multiple approaches.
[0149] In any embodiment, the RAG component 1192 can implement a plugin, API, user interface, and / or other functionality to perform RAG. For example, a graph RAG plugin can be used by the LLM / SLM / VLM / MMLM / etc. to query the knowledge graph to extract relevant information to feed into the model, and a standard or vector RAG plugin can be used to query a vector database. The graph database can, for example, interact with a plugin REST interface, thus decoupling the graph database from the vector database and / or the embedding models.
[0150] The Tokenizer 1110 can segment (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, the tokens can represent individual words, partial words, characters, or parts of audio, video, images, etc. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Partial word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 1130 to understand morphological variations and more effectively comprehend out-of-vocabulary words. Character-based tokenization represents each character as a separate token, allowing the generative LM 1130 to process text at a fine-grained level.The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or characteristics of the training dataset. As such, the tokenizer 1110 can convert the (e.g., processed) text into a structured format according to the tokenization scheme implemented in the specific embodiment.
[0151] The embedding component 1120 can use any known embedding technique to convert discrete tokens into (e.g., dense, continuous vector) representations of semantic meaning. For example, the embedding component 1120 can use pretrained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term-frequency-inverse-document-frequency (TF-IDF) coding, one or more embedding layers of a neural network, and / or other techniques.
[0152] In some embodiments where the input 1101 includes image data / video data / etc., the input processor 1101 can resize the data to a standard size to be compatible with the format of a corresponding input channel, and / or normalize pixel values to a common range (e.g., 0 to 1) to ensure a consistent representation, and the embedding component 1120 can encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features).In some embodiments where the input 1101 includes audio data, the input processor 1101 can resample an audio file to a consistent sampling rate for uniform processing, and the embedding component 1120 can use any known technique to extract and encode audio features—such as in the form of a spectrogram, which is said to include: Fourier spectrograms, mel spectrograms, wavelet spectrograms, time-domain filter banks, custom filter banks, learnable filter banks, custom audio embeddings, and / or other forms of analog or digital audio intermediate representation, or a combination thereof.In some embodiments where the input 1101 includes video data, the input processor 1101 can extract frames or apply resizing to extracted frames, and the embedding component 1120 can extract features such as optical flow embeddings or video embeddings and / or encode temporal information or sequences of frames. In some embodiments where the input 1101 includes multimodal data, the embedding component 1120 can fuse representations of the different types of data (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.
[0153] The generative LM 1130 and / or other components of the generative LM system 1100 can employ different types of neural network architectures, depending on the implementation. For example, transformer-based architectures, such as those used in models like GPT, can be implemented and include self-attention mechanisms that weight the meaning of different words or tokens in the input sequence and / or feedforward networks that process the output of the self-attention layers, thereby applying nonlinear transformations to the input representations and extracting higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, intermodal embedding models that learn shared embedding spaces, graphical neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative, adversarial networks or GANs, or adversarial autocoders (AAEs) for joint distributional learning, and others. As such, depending on the implementation and architecture, the embedding component 1120 can apply a coded representation of the input 1101 to the generative LM 1130, and the generative LM 1130 can process the coded representation of the input 1101 to generate an output 1190, which may include a response text and / or other types of data.
[0154] As described herein, in some embodiments, the generative LM 1130 may be configured to access or use plug-ins / APIs 1195 (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations for which the generative LM 1130 is not ideally suited, the model may have instructions (e.g., as a result of training and / or based on instructions in a given prompt, such as those retrieved using the RAG component 1192) to access one or more plug-ins / APIs 1195 (e.g., third-party plug-ins) to assist in processing the current input.In such an example, where at least part of a prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least part of the prompt related to the specific plugin / API 1195 to the plugin / API 1195, where the plugin / API 1195 can process the information and send a response back to the generative LM 1130, and the generative LM 1130 can use the response to generate the output 1190. This process can be repeated—e.g., recursively—for any number of iterations and using any number of plugins / APIs 1195 until an output 1190 can be generated that addresses every request / question / request / process / operation / etc. from the input 1101.As such, the model(s) may depend not only on its own knowledge gained from training with a large dataset (large datasets) and / or on data retrieved using the RAG component 1192, but also on the know-how or optimized nature of one or more external resources - such as the plug-ins / APIs 1195.
[0155] Fig. Figure 11B is a block representation of an exemplary generative LM 1130, which includes a transformer-encoder-decoder suitable for use in implementing at least some embodiments of the present disclosure; for example, input text such as "Who discovered gravity" can be tokenized into tokens, such as words (e.g., by the tokenizer 1110 of [reference missing]). Fig. 11A) and each token is encoded into a corresponding embedding (e.g., of size 512) (e.g., by the embedding component 1120 of FIG. 911A). Since these token embeddings typically do not represent the position of the token in the input sequence, any known technique can be used to add positional encoding to each token embedding in order to encode the sequential relationships and context of the tokens in the input sequence. As such, the (e.g., resulting) embeddings can be applied to one or more encoders 1135 of the generative LM 1130.
[0156] In an exemplary embodiment, the encoders 1135 form an encoder stack, where each encoder includes a self-attention layer and a feedforward network. In an exemplary transformer architecture, each token (e.g., word) flows through a separate path. As such, each encoder can accept a sequence of vectors, with each vector passing through the self-attention layer, then the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used.For example, to calculate a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. A self-attention score for pairs of tokens can be calculated by taking the dot product of the query vector with the corresponding key vectors, normalizing the resulting scores, multiplying them by the corresponding value vectors, and summing the weighted value vectors. The coder can apply multi-headed attention, applying the attention mechanism multiple times in parallel with different learned weight matrices. Any number of coders can be cascaded to generate a context vector, which is used to code the input. An attention projection layer 1140 can convert the context vector into attention vectors (keys and values) for the decoder(s) 1145.
[0157] In an exemplary embodiment, the decoder(s) 1145 can form a decoder stack, in which each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses the attention vectors (keys and values) from the decoder to focus relevant parts of the input sequence, and a feedforward network. As with the encoder(s) 1135, in an exemplary transformer architecture, each token (e.g., word) flows through a separate path in the decoder(s) 1145. During a first pass, the decoder(s) 1145, a classifier 1150, and a generation mechanism 1155 can generate an initial token, and the generation mechanism 1155 can apply the generated token as input during a second pass. The process can repeat in a loop, with tokens (e.g.,Words) are generated and added to the output from the previous pass, and the token embeddings of the composite sequence with positional encodings are applied as input to decoder(s) 1145 during a subsequent pass, generating one token sequentially at a time (known as autoregression) until a symbol or token representing the end of the response is predicted. Within each decoder, the self-attention layer is typically restricted to being attentive only to previous positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In an exemplary embodiment, the encoder-decoder attention layer operates similarly to the (e.g.,Multi-Headed) self-awareness in the coder(s) 1135, except that it creates its queries from the layer below it and takes the keys and values (e.g. matrix) from the output of coder(s) 1135.
[0158] As such, the decoder(s) 1145 can output a decoded (e.g., vector) representation of the input applied during a given iteration. The classifier 1150 can include a multi-class classifier comprising one or more neural network layers that project the decoded (e.g., vector) representation into an appropriate dimensionality (e.g., one dimension for each supported word or token in the output vocabulary), and a softmax operation that converts logits into probabilities. As such, the generation mechanism 1155 can select or sample a word or token based on an appropriate predicted probability (e.g., selecting the word with the highest predicted probability) and append it to the output of a previous iteration, with each word or token being generated sequentially.The generation mechanism 1155 can repeat the process, triggering successive decoder inputs and corresponding predictions until a symbol or token is selected or sampled that represents the end of the response, after which the generation mechanism 1155 can output the generated response.
[0159] Fig. Figure 11C is a block diagram of an exemplary embodiment in which the generative LM 1130 includes a decoder-transformer-only architecture suitable for use in implementing at least some embodiments of the present disclosure. For example, the decoder(s) 1160 of Fig. 11C similar to decoder(s) 1145 from Fig. 11B work, with the exception that the decoder(s) 1160 of Fig. 11C omits the encoder-decoder self-attention layer (since there is no encoder in this embodiment). As such, the decoder(s) 1160 can form a decoder stack, with each decoder including a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or token representing the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., corresponding embeddings with positional encodings) can be applied to the decoder(s) 1160. As with the decoder(s) 1145 of Fig. In 11B, each token (e.g., word) can flow through a separate path in the decoder(s) 1160, and the decoder(s), a classifier 1165, and a generation mechanism 1170 can use autoregression to sequentially generate one token at a time until a symbol or token representing the end of the response is predicted. The classifier 1165 and the generation mechanism 1170 can function similarly to the classifier 1150 and the generation mechanism 1155 of Fig. 11B operates wherein the generation mechanism 1170 selects and samples each successive output token based on a corresponding predicted probability and appends it to the output of a previous iteration, generating each token sequentially until a symbol or token is selected or sampled that represents the end of the response. This and other architectures described herein are merely examples, and other suitable architectures may be implemented within the scope of this disclosure. EXAMPLE CALCULATION DEVICE
[0160] Fig. Figure 12 is a block representation of an exemplary computing device(s) 1200, which is / are suitable for use in implementing some embodiments of the present disclosure. The computing device 1200 can include a connection system 1202 that directly or indirectly couples the following devices: memory 1204, one or more central processing units (CPUs) 1206, one or more graphics processing units (GPUs) 1208, a communication interface 1210, input / output (I / O) ports 1212, input / output components 1214, a power supply 1216, one or more presentation components 1218 (e.g., display(s)), and one or more logic units 1220. In at least one embodiment, the computing device(s) 1200 can include one or more virtual machines (VMs) and / or any of its components can include virtual components (e.g., virtual hardware components).For non-restrictive examples, one or more of the GPUs 1208 may comprise one or more vGPUs, one or more of the CPUs 1206 may comprise one or more vCPUs, and / or one or more of the logic units 1220 may comprise one or more virtual logic units. As such, a computing device(s) 1200 may include discrete components (e.g., a complete GPU dedicated to the computing device 1200), virtual components (e.g., a portion of a GPU dedicated to the computing device 1200), or a combination thereof.
[0161] Although the various blocks of Fig. Where components 12 are shown to be connected via lines through the connection system 1202, this is not intended to be restrictive and serves only for clarity. For example, in some embodiments, a presentation component 1218, such as a display device, can be considered an I / O component 1214 (e.g., if the display is a touchscreen). As another example, the CPUs 1206 and / or GPUs 1208 can include memory (e.g., the memory 1204 can be representative of a storage device, in addition to the memory of the GPUs 1208, the CPUs 1206, and / or other components). As such, the computing device of Fig. 12 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all fall under the category of computing device. Fig. 12 are considered.
[0162] The 1202 interconnect system can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 1202 interconnect system can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standard Association (VESA) bus, a Peripheral Component Connection (PCI) bus, a Peripheral Component Connection Express (PCIe) bus, and / or other bus or link types. In some embodiments, there are direct connections between components. For example, the CPU 1206 can be directly connected to the memory 1204. Similarly, the CPU 1206 can be directly connected to the GPU 1208. In the case of a direct or point-to-point connection between components, the 1202 interconnect system can include a PCIe link to establish the connection.In these examples, the computing device 1200 does not need to include a PCI bus.
[0163] Memory 1204 can include any of a range of computer-readable media. The computer-readable media can be any available media accessible to the computing device 1200. The computer-readable media can include both volatile and non-volatile media, and both removable and non-removable media. As an example, and not as a limitation, the computer-readable media can include computer storage media and communication media.
[0164] Computer storage media can include both volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1204 can store computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1200. As used herein, computer storage media do not, per se, include signals.
[0165] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include information delivery media. The term "modulated data signal" can refer to a signal for which one or more of its characteristics are set or modified in a way that encodes information in the signal. By way of example, and not limited to this, computer storage media can include wired media, such as a wired network or a directly wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of any of the foregoing should also be included in the scope of computer-readable media.
[0166] The CPU(s) 1206 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the procedures and / or processes described herein. The CPU(s) 1206 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a plurality of software threads concurrently. The CPU(s) 1206 can include any type of processor and can include different types of processors depending on the type of computing device 1200 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).For example, depending on the type of computing device 1200, the processor can be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 1200 can include one or more CPUs 1206 in addition to one or more microprocessors or supplementary coprocessors, such as math coprocessors.
[0167] In addition to or as an alternative to the CPU(s) 1206, the GPU(s) 1208 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the procedures and / or processes described herein. One or more of the GPU(s) 1208 may be an integrated GPU (e.g., in one or more of the CPU(s) 1206) and / or one or more of the GPU(s) 1208 may be a discrete GPU. In embodiments, one or more of the GPU(s) 1208 may be a coprocessor of one or more of the CPU(s) 1206. The GPU(s) 1208 may be used by the computing device 1200 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPU(s) 1208 can be used for general-purpose computing on GPUs (GPGPU).The GPU(s) 1208 can include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. The GPU(s) 1208 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 1206 received via a host interface). The GPU(s) 1208 can include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory can be included as part of the memory 1204. The GPU(s) 1208 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using an NVLINK) or connect them via a switch (e.g., using an NVSwitch).When combined, each GPU can generate 1208 pixel data or GPGPU data for different dividers, or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or share memory with other GPUs.
[0168] In addition to or as an alternative to the CPU(s) 1206 and / or the GPU(s) 1208, the logic unit(s) 1220 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 1206, the GPU(s) 1208, and / or the logic unit(s) 1220 may discretely or jointly perform any combination of the methods, processes, and / or parts thereof. One or more of the logic units 1220 can be part of and / or integrated into one or more of the CPU(s) 1206 and / or the GPU(s) 1208 and / or one or more of the logic units 1220 can be discrete components or otherwise external to the CPU(s) 1206 and / or the GPU(s) 1208.In embodiments, one or more of the logic units 1220 can be a coprocessor of one or more of the CPU(s) 1206 and / or one of the GPU(s) 1208.
[0169] Examples of the Logic Unit(s) 1220 include: one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Programmable Vision Accelerators (PVAs) – which may include one or more Direct Memory Access (DMA) systems, one or more Vision or Vector Processing Units (VPUs), one or more Pixel Processing Engines (Pixel Processing Engines, PPEs), one or more decoupled accelerators (e.g.decoupled lookup table (DLUT) accelerators, etc., vision processing units (VPUs), optical flow accelerators (OFAs), field programmable gate arrays (FPGAs), neuromorphic chips, quantum processing units (QPUs), associative process units (APUs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) devices, peripheral interconnect (PCI) or peripheral interconnect express (PCIe) devices, and / or the like.
[0170] The communication interface 1210 can include one or more receivers, transmitters, and / or transceivers that allow the computing device 1200 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1210 can include components and functionality to allow communication over any number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 1220 and / or communication interface 1210 may include one or more data processing units (DPUs) to transfer data received via a network and / or the connection system 1202 directly to (e.g., a memory of) one or more GPU(s) 1208.
[0171] The I / O ports 1212 allow the computing device 1200 to be logically coupled with other devices, including the I / O components 1214, the presentation component(s) 1218, and / or other components, some of which may be built into (e.g., integrated with) the computing device 1200. Illustrative I / O components 1214 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite table, scanner, printer, wireless device, etc. The I / O components 1214 can provide a natural user interface (NUI) that processes air gestures, speech, or other physical input generated by a user. In some cases, input can be transmitted to an appropriate network element for further processing.A NUI can implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, gesture recognition both on-screen and adjacent to the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below), associated with a display of the Computing Device 1200. The Computing Device 1200 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. Furthermore, the Computing Device 1200 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit, IMU) that allow motion detection.In some examples, the output of the accelerometers or gyroscopes can be used by the Computing Device 1200 to render immersive augmented reality or virtual reality.
[0172] The power supply 1216 can include a hardwired power supply, a battery power supply, or a combination of both. The power supply 1216 can provide power to the computing device 1200 to allow the components of the computing device 1200 to operate.
[0173] The presentation component(s) 1218 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 1218 can receive data from other components (e.g., the GPU(s) 1208, the CPU(s) 1206, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0174] Fig. Figure 13 illustrates an exemplary data center 1300 that can be used in at least one embodiment of the present disclosure. The data center 1300 can include a data center infrastructure layer 1310, a framework layer 1320, a software layer 1330, and / or an application layer 1340.
[0175] As in Fig. As shown in Figure 13, the data center infrastructure layer 1310 can include a resource orchestrator 1312, clustered compute resources 1314, and node compute resources (“node RR”) 1316(1)-1316(N), where “N” represents an integer, positive number. In at least one embodiment, Node-RR 1316(1)-1316(N) can include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc. In some embodiments, one or more Node-RRs can be controlled by the Node-RR1316(1)-1316(N) correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node RR 1316(1)-1316(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or may correspond to one or more of the node RR 1316(1)-1316(N) of a virtual machine (VM).
[0176] In at least one embodiment, grouped compute resources 1314 can include separate groupings of node RR 1316 located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of node RR 1316 within grouped compute resources 1314 can include grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RR 1316, including CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.
[0177] The resource orchestrator 1312 can configure or otherwise control one or more node RR 1316(1)-1316(N) and / or grouped compute resources 1314. In at least one embodiment, the resource orchestrator 1312 can include a software design infrastructure (SDI) management entity for the data center 1300. The resource orchestrator 1312 can include hardware, software, or a combination thereof.
[0178] In at least one embodiment, as in Fig. As shown in Figure 13, a framework layer 1320 can include a job scheduler 1328, a configuration manager 1334, a resource manager 1336, and / or a distributed file system 1338. The framework layer 1320 can include a framework to support software 1332 of software layer 1330 and / or one or more application(s) 1342 of application layer 1340. The software 1332 or application(s) 1342 can each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1320 can be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use the distributed file system 1338 for large-volume data processing (e.g. "Big Data").In at least one embodiment, the job scheduler 1328 can include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1300. The configuration manager 1334 can be capable of configuring various layers, such as the software layer 1330 and the framework layer 1320, including Spark and the distributed file system 1338, to support high-volume data processing. The resource manager 1336 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 1338 and the job scheduler 1328. In at least one embodiment, the clustered or grouped compute resources can include a grouped compute resource 1314 at the data center infrastructure layer 1310.The resource manager 1336 can coordinate with the resource orchestrator 1312 to manage these allocated or assigned computing resources.
[0179] In at least one embodiment, software 1332, which is enclosed in software layer 1330, can include software used by at least parts of the node RR 1316(1)-1316(N), grouped computing resources 1314, and / or distributed file system 1338 of framework layer 1320. One or more types of software can include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0180] In at least one embodiment, one or more application(s) 1342 enclosed in the application layer 1340 may include one or more types of applications used by at least parts of the node RR 1316(1)-1316(N), grouped compute resources 1314, and / or distributed file system 1338 of the framework layer 1320. One or more types of applications may include, but are not limited to, any number of a genomics application, a cognitive computation application, and a machine learning application, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0181] In at least one embodiment, any configuration manager 1334, resource manager 1336, and resource orchestrator 1312 can implement any number and any type of self-modifying operations based on any set and any type of data captured in any technically feasible manner. Self-modifying operations can relieve a data center operator of data center 1300 of the burden of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing parts of a data center.
[0182] The Data Center 1300 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using software and / or computing resources as described above in relation to the Data Center 1300.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer and predict information regarding Data Center 1300 using the resources described above, by using weight parameters calculated via one or more training techniques, such as, but not limited to, those described herein.
[0183] In at least one embodiment, the data center can use 1300 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0184] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be connected to one or more instances of the computing device(s) 1200. Fig. 12. Each device can include similar components, features, and / or functionality to the computing device(s) 1200. Furthermore, if backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can be included as part of a data center 1300; an example of this is described in more detail below with reference to… Fig. 13 described.
[0185] Components in a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the internet, and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0186] Compatible network environments can include one or more peer-to-peer network environments—in which case no server may be included in a network environment—and one or more client-server network environments—in which case one or more servers may be included in a network environment. In peer-to-peer network environments, functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0187] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more application(s) of an application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be, but is not limited to, a type of free and open-source software web application framework that can use a distributed file system for large-scale data processing (e.g., "big data").
[0188] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination of computing and / or data storage functions as described herein (or one or more parts thereof). Any of these various functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers that may be distributed across a state, region, country, the Earth, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server(s), a core server may designate at least some of the functionality for that edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0189] The client device(s) may include at least some of the components, features, and functionality described herein with reference to Fig.The 12 exemplary computing devices described include 1200. By way of example, and not as a limitation, a client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or global positioning device, video player, video camera, surveillance device or system, vehicle, boat, flying vehicle, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0190] The disclosure can generally be described in the context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or a handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure can be exercised in various system configurations, including handheld devices, consumer electronics, general-purpose computers, other specialized computing devices, etc. The disclosure can also be exercised in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0191] Other variations are consistent with the nature of the present disclosure. Therefore, while the disclosed techniques are receptive to various modifications and alternative constructions, certain illustrated embodiments are shown in the drawings and have been described in detail above. It should be understood, however, that there is no intention to limit the disclosure to any specific form or forms that are disclosed; on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents that are consistent with the nature and scope of the disclosure as defined in the appended claims.
[0192] The use of the terms "a," "an," "the," "a," and "a," and reference terms in the context of describing the disclosed embodiments (particularly in the context of the claims below) is to be interpreted as covering both singular and plural forms, unless otherwise specified herein or clearly contrary to the context, and not as defining a term. The terms "comprising," "having," "including," and "containing" are to be interpreted as open terms (meaning "including but not limited to") unless otherwise noted. "Connected," unless modified and referring to physical connections, is to be interpreted as partially or wholly contained within, attached to, or joined with, even if something is in between.The mention of values herein is intended as a quick method of individually referring to each separate value falling within the scope, unless otherwise specified herein, and each separate value is included in the patent specification as if it were individually named herein. In at least one embodiment, the use of the term "sentence" (e.g., "a set of things") or "substance," unless otherwise noted or contradicted by context, is to be interpreted as a non-empty collection comprising one or more elements. Furthermore, unless otherwise noted or contradicted by context, the term "substance" of a corresponding sentence does not necessarily denote a proper substance of the corresponding sentence; however, the substance and the corresponding sentence may be the same.
[0193] A conjunctive expression, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C," unless specifically stated otherwise or otherwise clearly contradicted by the context, is otherwise to be understood, in accordance with the context, as it is generally used to present that a thing, concept, etc., can be either A, B, or C, or a non-empty clause of the sentence of A, B, and C. For example, in an illustrative example of a sentence containing three elements, the conjunctive expressions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any one of the following: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Such a conjunctive expression therefore generally does not intend to imply that certain embodiments require that at least one of A, at least one of B and at least one of C be present.Furthermore, unless otherwise noted or contradicted by the context, the term "plurality" indicates that a state is plural (e.g., "a plurality of things" indicates multiple things). In at least one embodiment, the number of things in a plurality is at least two, but may be more if either explicitly stated or indicated by the context. Furthermore, unless otherwise stated or clear from the context, the phrase "based on" means "at least partly based on" and not "exclusively based on".
[0194] The operations of processes described herein may be performed in any suitable order unless otherwise specified herein or otherwise clearly contradicted by the context. In at least one embodiment, a process such as those described herein (or variants and / or combinations thereof) is carried out under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications), executed together on one or more processors, by hardware, or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transient computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electrical or electromagnetic transmission) but includes a non-transient data storage circuit (e.g., buffers, cache, and queues) within transceivers of transitory signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transient computer-readable storage media that have executable instructions stored thereon (or on other memory to store executable instructions) which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein.In at least one embodiment, a set of non-transient machine-readable storage media comprises several non-transient machine-readable storage media, and one or more individual non-transient storage media of several non-transient machine-readable storage media lack all the code, while several non-transient machine-readable storage media together store all the code. In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors—for example, a non-transient machine-readable storage medium stores instructions, and a central processing unit (“CPU”) executes some instructions, while a graphics processing unit (“GPU”) executes other instructions.In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of instructions.
[0195] Accordingly, in at least one embodiment, computer systems are configured to implement one or more services that, alone or together, perform operations of the processes described herein, and such computer systems are configured with applicable hardware and / or software that enables the performance of operations. Furthermore, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, it is a distributed computer system comprising several devices that operate differently, such that distributed computer systems perform the operations described herein and such that a single device does not perform all operations.
[0196] The use of an example and all examples or exemplary language (e.g., "such as"), as provided herein, is intended only to make embodiments of the disclosure more readily understandable and does not constitute a limitation of the scope of the disclosure, unless otherwise claimed. No expression in the patent specification shall be construed as implying that an unclaimed element is essential for exercising the disclosure.
[0197] All references, including publications, patent applications, and patents mentioned herein, are hereby incorporated by reference to the same extent as if each reference had been individually and specifically incorporated by reference and set forth herein in its entirety.
[0198] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It is understood that these terms may not be intended as synonyms. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other but nevertheless work together or interact.
[0199] Unless specifically stated otherwise, it may be apparent throughout the patent specification that terms such as "processing", "calculating", "calculating", "determining" or the like refer to operations and / or processes of a computer or computing system, or a similar electronic computing device, which manipulate data represented as physical, such as electronic quantities within the registers and / or memory of the computing system and / or transform them into other data represented likewise as physical quantities within the memory, registers or other such information storage, transmission or display devices of the computing system.
[0200] Similarly, the term "processor" can refer to a device or part of a device that processes electronic data from registers and / or memories and transforms that electronic data into other electronic data that can be stored in registers and / or memories. As non-restrictive examples, the "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, "software" processes can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently.In at least one embodiment, the terms “system” and “method” are used synonymously herein insofar as a system may contain one or more methods and methods may be considered as a system.
[0201] This document may refer to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, a process of obtaining, acquiring, receiving, or inputting analog and digital data can be implemented in various ways, such as receiving data as parameters of a function call or a call to an application programming interface. In at least one embodiment, the processes of obtaining, acquiring, receiving, or inputting analog and digital data can be implemented by transferring data over a serial or parallel interface.In at least one embodiment, the processes of obtaining, acquiring, receiving, or inputting analog or digital data can be implemented by transferring data over a computer network from a providing entity to an acquiring entity. In at least one embodiment, references can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, processes for providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transferring data as input or output parameters of a function call, a parameter of an application programming interface, or an inter-process communication mechanism.
[0202] Although the descriptions herein present exemplary embodiments of the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. Furthermore, although specific distributions of responsibilities may be defined above for the purpose of description, various functions and responsibilities may be distributed and subdivided in different ways depending on the circumstances.
[0203] Furthermore, although the subject matter has been described in terms specific to structural features and / or methodological actions, it is understood that the subject matter claimed in the attached patent claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.
[0204] The disclosure of this application also includes the following numbered clauses: Clause 1: A procedure comprising: processing, using a first speech-to-text model (S2T model), a first input that includes a speech in a first language to generate a transcription of the speech; and processing, using a second S2T model, a second input to generate a translation of the speech into a second language, the second input including at least a representation of the speech and the transcription of the speech. Clause 2. Procedure according to Clause 1, wherein the transcription is in the first language. Clause 3. Procedure according to a preceding clause, wherein the first S2T model includes a decoder-only language model. Clause 4. Procedure according to a preceding clause, wherein the second S2T model includes an encoder-decoder language model. Clause 5. Procedure according to a preceding clause, wherein the second S2T model includes the first S2T model. Clause 6. Procedure according to a preceding clause, wherein the second input further includes a natural language description of a task to be performed by the second S2T, wherein the task includes generating the translation of the spoken language into the second language. Clause 7. Procedure according to a preceding clause, wherein the representation of the spoken language is generated using an audio encoder network. Clause 8. Procedure following a preceding clause, wherein the representation of the spoken language and the transcription of the spoken language are chained together to obtain the second input. Clause 9. A procedure according to a preceding clause, wherein at least one of the first S2T model or the second S2T model is trained using operations that include: receiving a training input that includes: a first part comprising a representation of a training speech language in a first language, and a second part comprising a transcription of the training speech language in the first language generated by the first S2T model; processing, using the second S2T model, the training input to generate a training translation of the training speech language into the second language;and modify, at least based on a comparison of the training translation of the training speech language into the second language with a ground-truth translation (GT translation) of the training speech language into the second language, of one or more parameters of at least one of the first S2T model or the second S2T model. Clause 10. Procedure comprising: receiving a training input comprising: a first part comprising a representation of a training speech language in a first language, and a second part comprising a transcription of the training speech language in the first language; processing, using a speech-to-text (S2T) model, the training input to generate a translation of the training speech language into a second language; and modifying, at least based on a comparison of the translation of the training speech language into the second language with a ground-truth (GT) translation of the training speech language into the second language, one or more parameters of the S2T model. Clause 11. Procedure according to Clause 10, wherein the transcription is obtained by processing the training speech in the first language using a trained automatic speech recognition (ASR) model. Clause 12. Procedure according to one of Clauses 10 or 11, wherein the GT translation of the training speech language is obtained by translating a GT transcription into the first language using a translation model. Clause 13. Method according to any of Clauses 10 to 12, wherein the S2T model comprises an adapter network, and wherein modifying one or more parameters of the S2T model comprises: modifying one or more parameters of the adapter network. Clause 14. Method according to any of Clauses 10 to 13, wherein obtaining the training input comprises: processing, using an audio encoder network, the training speech to generate the representation of the training speech; and wherein the method further comprises: modifying one or more parameters of the audio encoder network. Clause 15. Procedure according to any of Clauses 10 to 14, wherein the S2T model comprises an encoder-decoder model. Clause 16. Procedure according to any of Clauses 10 to 15, wherein the training input further includes a third part with a natural language description of a task to be performed by the S2T model, wherein the task comprises generating the translation of the training speech language into the second language. Clause 17. System comprising: one or more processors for: translating spoken language from a first language into a second language at least based on language model processing (i) a representation of the spoken language and (ii) a transcription of the spoken language into the first language. Clause 18. System according to Clause 17, wherein the one or more processors further serve to: generate the transcription of the spoken language in the first language using the language model or a second language model. Clause 19. System according to one of Clauses 17 or 18, wherein the language model is trained using operations that include: receiving a training input comprising: a first part comprising a representation of a training speech language in a first language, and a second part comprising a transcription of the training speech language in the first language; processing, using the language model, the training input to generate a training translation of the training speech language into the second language; and modifying, at least based on a comparison of the training translation of the training speech language into the second language with a ground-truth translation (GT translation) of the training speech language into the second language, one or more parameters of the language model. Clause 20. System according to any of Clauses 17 to 19, wherein the system includes at least one of the following: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing a light transport simulation; a system for performing one or more medical operations; a system for performing one or more factory operations; a system for performing one or more analytical operations; a system implementing one or more inference microservices; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device;a system for generating or presenting at least one type of virtual reality content, mixed reality content, or augmented reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implementing one or more large language models (LLMs); a system implementing one or more small language models (SLMs); a system implementing one or more vision language models (VLMs); a system implementing one or more multimodal language models; a system implementing one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system involving one or more virtual machines (VMs); a system that is at least partially implemented in a data center;or a system that is implemented at least partially using cloud computing resources.
[0205] It is understood that the aspects and embodiments described above are merely examples and that modifications in detail may be made within the scope of the claims.
[0206] Each device, method and feature disclosed in the description and (if applicable) in the claims and drawings can be provided independently or in any suitable combination.
[0207] Reference numerals appearing in the claims serve only for illustration and are not intended to have any limiting effect on the scope of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 18,304,341
[0139]
Claims
[1] Procedure comprising the following: Processing, using an initial speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and Processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language, wherein the second input includes at least a representation of the spoken language and the transcription of the spoken language. [2] Method according to claim 1, wherein the transcription is in the first language. [3] Method according to a preceding claim, wherein the first S2T model includes a decoder-only language model. [4] Method according to a preceding claim, wherein the second S2T model includes an encoder-decoder language model. [5] Method according to a preceding claim, wherein the second S2T model includes the first S2T model. [6] Method according to a preceding claim, wherein the second input further includes a natural language description of a task to be performed by the second S2T, wherein the task includes generating the translation of the spoken language into the second language. [7] Method according to a preceding claim, wherein the representation of the spoken language is generated using an audio encoder network. [8] Method according to a preceding claim, wherein the representation of the spoken language and the transcription of the spoken language are chained together to obtain the second input. [9] Method according to a preceding claim, wherein at least one of the first S2T model or the second S2T model is trained using operations comprising: Receiving a training input that includes the following: a first part, which includes a representation of a training speech language in a first language, and a second part, which includes a transcription of the training speech generated by the first S2T model into the first language; Processing, using the second S2T model, the training input to generate a translation of the training speech into a second language; and Modify, at least based on a comparison of the translation of the training speech language into the second language with a ground-truth translation (GT translation) of the training speech language into the second language, of one or more parameters of at least one of the first S2T model or the second S2T model. [10] Method comprising the following: Receiving a training input that includes the following: a first part, which includes a representation of a training speech language in a first language, and a second part, which includes a transcription of the training speech in the first language; Processing the training input using a speech-to-text (S2T) model to generate a translation of the training speech into a second language; and Modify, at least based on a comparison of the translation of the training speech language into the second language with a ground-truth translation (GT translation) of the training speech language into the second language, of one or more parameters of the S2T model. [11] Method according to claim 10, wherein the transcription is obtained by processing the training speech in the first language using a trained automatic speech recognition (ASR) model. [12] Method according to one of claims 10 or 11, wherein the GT translation of the training speech language is obtained by translating a GT transcription into the first language using a translation model. [13] Method according to any one of claims 10 to 12, wherein the S2T model comprises an adapter network, and wherein modifying one or more parameters of the S2T model comprises: Modifying one or more parameters of the adapter network. [14] Method according to any one of claims 10 to 13, wherein receiving the training input comprises: Processing, using an audio encoder network, the training speech to generate the representation of the training speech; and the procedure further includes the following: Modifying one or more parameters of the audio encoder network. [15] Method according to any one of claims 10 to 14, wherein the S2T model comprises an encoder-decoder model. [16] Method according to any one of claims 10 to 15, wherein the training input further includes a third part with a natural language description of a task to be performed by the S2T model, wherein the task comprises generating the translation of the training speech language into the second language. [17] System comprising the following: one or more processors for: translating spoken language from a first language into a second language, at least based on language model processing (i) a representation of the spoken language and (ii) a transcription of the spoken language into the first language. [18] System according to claim 17, wherein one or more processors further serve to: generate the transcription of the spoken language in the first language using the language model or a second language model. [19] System according to one of claims 17 or 18, wherein the language model is trained using operations comprising: Receiving a training input that includes the following: a first part, which includes a representation of a training speech language in a first language, and a second part, which includes a transcription of the training speech in the first language; Processing, using the language model, the training input to generate a translation of the training speech into a second language; and Modify, at least based on a comparison of the training translation of the training language into the second language with a ground-truth translation (GT translation) of the training language into the second language, one or more parameters of the language model. [20] System according to any one of claims 17 to 19, wherein the system comprises at least one of the following: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for conducting digital twinning operations; a system for performing a light transport simulation; a system for performing one or more medical operations; a system for performing one or more factory operations; a system for performing one or more analytical operations; a system that implements one or more inference microservices; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, mixed reality content, or augmented reality content; a system that is implemented using a robot; a system for performing one or more conversational AI operations; a system that implements one or more large language models (LLMs); a system that implements one or more small language models (SLMs); a system that implements one or more Visual Language Models (VLMs); a system that implements one or more multimodal language models; a system that implements one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
18,304,341
US18304341B2
Cited By
Data processing method and device and storage medium
CN116911683A