Spatially aware audio enhanced dialog proxy
By converting multi-channel audio data into a device-independent format and training a multimodal language model, the problem of the inability to process multi-channel audio data in existing technologies is solved, thereby improving the utilization of spatial information in dialogue agents.
Patent Information
- Application Number
- CN202511147258.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-16
- Filing Date
- 2025-08-15
- Publication Date
- 2026-03-03
AI Technical Summary
Existing language models cannot effectively process multi-channel audio data, resulting in an inability to utilize the spatial characteristics of audio data and affecting the realism of the dialogue agent's performance.
Multichannel audio data is converted to a device-independent audio format, such as B-format audio, and processed through tokenization and machine learning layers to train a multimodal language model to recognize and process the spatial characteristics of the audio data.
This improves the realism of the dialogue agent, enabling it to utilize spatial information from audio during interactions, thereby enhancing the agent's intelligence and accuracy.
Smart Images

Figure CN121600941A_ABST
Abstract
Description
Background Technology
[0001] Language models such as large language models (LLMs) can be used to process text data (e.g., in natural language) to implement dialogue agents. Multimodal language models include language models trained to additionally process information in other modalities, such as audio data, image data, and / or other types of data. However, existing solutions are generally not designed to process multichannel (e.g., stereo or binaural) audio data, but only mono audio data (e.g., with a single audio channel). Summary of the Invention
[0002] Multimodal language models for processing audio data operate as follows: First, the input audio data is encoded into a digital format using techniques such as tokenization. Then, machine learning layers, such as transformer layers, are used to process the tokenized audio. The conventional approach to processing audio data using multimodal models operates only on single-channel audio data. Single-channel audio data is represented as audio information from a single source and can only include time-varying audio intensity (e.g., volume), without providing spatial (e.g., direction) characteristics of the sound.
[0003] In contrast, multichannel audio data includes audio from multiple audio sources, making it possible to derive spatial information, such as the location and motion of sound sources, from the changes in audio intensity between each audio channel. A common type of multichannel audio is stereo or binaural audio, which includes two channels of audio; however, in multichannel audio, any number of channels can be implemented. This variability in audio makes encoding multichannel audio for use with multimodal language models challenging.
[0004] The embodiments described herein enable the combined use of multi-channel audio with multimodal language models by converting multi-channel audio from any number of audio sources into a device-agnostic audio format. For example, the device-agnostic format could be B-format audio, which may include a fixed number of component channels that collectively represent a full-sphere sound field. Because the number of component channels is fixed and independent of the number of audio sources / channels used to record the initial audio data, the device-agnostic format can be tokenized using a tokenizer trained for the multimodal language model. Using multi-channel audio when training / updating the multimodal language model allows the model to learn the spatial properties of the audio. Therefore, dialogue agents deployed using multimodal language models that process audio in this way (e.g., chatbots, non-player characters (NPCs), digital humans, avatars, digital assistants, etc.) may perform more realistically because they are able to utilize spatial awareness from audio (and / or other sources, such as images, videos, environmental simulations, etc.) when interacting with users.
[0005] At least one aspect involves one or more processors. The one or more processors may include one or more circuits. The one or more circuits can generate an encoded representation of multi-channel audio data corresponding to a machine learning model. The one or more circuits can use the encoded representation to generate a training dataset for the machine learning model. The training dataset may indicate spatial information of at least one audio source represented in the multi-channel audio data. The one or more circuits can use the training dataset to update one or more parameters of the machine learning model to generate an output corresponding to the input spatial audio.
[0006] In some implementations, the machine learning model includes at least one of the following: a large language model (LLM), a visual language model (VLM), or a multimodal language model (MMLM). In some implementations, spatial information includes text data. In some implementations, one or more circuits can update one or more parameters of the machine learning model to generate output text data relating to at least one audio source represented in the input spatial audio. In some implementations, the output text data identifies one or more of the following: distance from the audio source represented in the input spatial audio, the number of audio sources represented in the input spatial audio, or a transcription or diarization output of speech from a moving audio source represented in the input spatial audio.
[0007] In some implementations, one or more circuits can generate multi-channel audio data by applying spatial transformation operations to multiple audio sources. In some implementations, the spatial transformation operation generates multi-channel audio data as B-format audio. In some implementations, one or more circuits can update one or more parameters of a machine learning model to generate output spatial audio based on the input spatial audio.
[0008] In some implementations, one or more circuits can generate a training dataset to include an encoded representation of video data. In some implementations, one or more circuits can use the training dataset to update one or more parameters of a machine learning model to generate output spatial audio that tracks at least one audio source depicted in the video data. In some implementations, one or more circuits can update one or more parameters of a machine learning model to receive encoded representations of single-channel audio and video data, thereby generating output spatial audio.
[0009] At least one aspect relates to a system. The system may include one or more processors. The system can receive input audio from a client device for a language model trained to process multi-channel audio data. The system can use the input and the language model to generate output data indicating spatial information of at least one audio source represented in the input audio. The system can provide the output data indicating the spatial information to the client device.
[0010] In some implementations, the system can generate an encoded representation of multi-channel audio data corresponding to a machine learning model. In some implementations, the system can provide the encoded representation as input to a language model. In some implementations, the system can receive input text for the language model. In some implementations, the system can use the language model to generate output data indicating spatial information based on the input text and input audio.
[0011] In some implementations, the system may receive input video for a language model. In some implementations, the system may use the language model to generate output data indicating spatial information based on the input video and input audio. In some implementations, the output data includes the encoded output of the language model. In some implementations, the system may generate multi-channel output audio based on the encoded output of the language model. In some implementations, the output data includes at least one of the following: the number of sound sources represented in the input audio, the estimated distance of the sound sources represented in the input audio, or the estimated location of the sound sources represented in the input audio.
[0012] At least one aspect relates to a method. The method may include: generating an encoded representation of multi-channel audio data corresponding to a machine learning model using one or more processors. The method may include: generating a training dataset for the machine learning model using the encoded representation with one or more processors. The training dataset may indicate spatial information of at least one audio source used for representation in the multi-channel audio data. The method may include: updating one or more parameters of the machine learning model using one or more processors and the training dataset to generate an output corresponding to the input spatial audio.
[0013] In some implementations, spatial information includes text data. In some implementations, the method may include updating one or more parameters of a machine learning model using one or more processors to generate output text data relating to at least one audio source represented in the input spatial audio. In some implementations, the output text data identifies one or more of the following: distance from the audio source represented in the input spatial audio, the number of audio sources represented in the input spatial audio, or transcription of speech from a moving audio source represented in the input spatial audio.
[0014] The processors, systems, and / or methods described herein can be implemented by or included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing generative AI operations using a large language model; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a language model; a system for performing generative AI operations using a large language model; a system for performing generative AI operations using a visual language model; a system for performing generative AI operations using a multi-model language model; a system for generating synthetic data; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Attached Figure Description
[0015] The system and method for spatially aware audio-enhanced dialogue agents are described in detail below with reference to the accompanying drawings, wherein:
[0016] Figure 1 This is a block diagram illustrating an example system for implementing a spatially aware audio-enhanced dialogue agent according to some embodiments of the present disclosure.
[0017] Figure 2 A data flow diagram illustrating how a spatially aware audio-enhanced dialogue agent can process different types of data is depicted according to some embodiments of the present disclosure;
[0018] Figure 3 This is a flowchart of a method for training / updating a spatially aware audio-enhanced dialogue agent according to some embodiments of this disclosure;
[0019] Figure 4A This is a block diagram of an example generative language model system suitable for implementing at least some embodiments of the present disclosure;
[0020] Figure 4B It is a block diagram of an example generative language model including a transformer encoder-decoder, suitable for implementing at least some embodiments of this disclosure;
[0021] Figure 4C It is a block diagram of an example generative language model including a decoder-only transformer architecture suitable for implementing at least some embodiments of this disclosure;
[0022] Figure 5 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0023] Figure 6 This is a block diagram of an example data center suitable for implementing some embodiments of this disclosure. Detailed Implementation
[0024] This disclosure relates to systems and methods for implementing spatially aware and audio-enhanced dialogue agents. The dialogue agent can be implemented using machine learning models such as Large Language Models (LLMs), Visual Language Models (VLMs), and Multimodal Language Models (MMLMs). Generative artificial intelligence models such as LLMs / VMLMs / MMLMs can receive and process information representing various media modalities, including audio, video, images, and text. Machine learning models trained / updated to receive input data with different media modalities can be referred to as "multimodal models" or MMLMs.
[0025] Processing information for use in generative multimodal models involves encoding that information into a digital format using techniques such as tokenization. Typically, tokenization converts input data into a format compatible with the input layer of a machine learning model. Some machine learning models implement audio-based processing, where a stream of audio information is encoded, and this encoded audio stream is provided as input to the machine learning model. These encoding processes convert single-channel audio data into a digital format for processing.
[0026] Single-channel audio data refers to audio recorded or encoded using only one audio channel. Single-channel audio data consists only of mono audio data without any spatial encoding. Conventional audio-based machine learning models that only process single-channel audio data cannot handle the spatial information of the audio source represented in the input audio data. Therefore, the various contextual information associated with the dialogue agent cannot be obtained using conventional machine learning models.
[0027] The systems and methods described in this paper implement techniques for processing multi-channel audio data that encodes spatial information from an arbitrary number of audio sources. The techniques described can be applied to any number of audio channels. To this end, spatial transformations can be applied to multi-channel microphone array signals to generate device-independent audio formats. An example of a device-independent spatial audio format is ambisonics representation. The spatial transformation may correspond to a multi-channel microphone system recording audio data.
[0028] Once encoded into a spatial audio format, a multichannel audio encoder encodes the spatial audio (e.g., via tokenization) into a format suitable for machine learning models. Such machine learning models can also include multimodal models that receive input from a combination of audio, text, or other modalities such as images or video. For example, special tokens or inputs for LLM / VLM / MMLM / etc. can specify portions of the input context sequence corresponding to audio data and portions corresponding to text data.
[0029] Machine learning models trained / updated using these techniques can process both the content and spatial context of input audio data. Unlike conventional methods, the techniques described in this paper enable machine learning models trained / updated to identify, isolate, and process the content of individual audio sources or directional audio. Spatial information can be used to track or estimate the location of speech or audio sources, or to derive additional insights from spatial audio that conventional machine learning models cannot generate. Therefore, the systems and methods described in this paper improve audio processing methods by extending the capabilities of conventional machine learning models.
[0030] refer to Figure 1 ,Should Figure 1This is an example computing environment including a system for implementing a spatially aware audio-enhanced dialogue agent according to some embodiments of this disclosure. It should be understood that such and other arrangements described herein are merely examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used to supplement or replace the illustrated arrangements and elements, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, a processor executing instructions stored in memory may be used to perform these functions.
[0031] System 100 is shown as including a data processing system 102. Data processing system 102 may include one or more processors, circuits, memories, and / or computing devices / systems capable of performing the various techniques described herein. Data processing system 102 may be implemented, for example, in a cloud computing environment and / or at the edge, and may maintain, update, and / or execute one or more language models 120 (e.g., LLM / VLM / MMLM / etc.). Data processing system 102 may implement the various techniques described herein for training / updating language models 120 to learn the spatial characteristics of extracting, processing, or interpreting multi-channel audio data. To this end, data processing system 102 (or components thereof) may access storage device 106 to generate and / or retrieve training dataset 110, which may include encoded audio samples 112 (e.g., encoded multi-channel audio data 108) and corresponding spatial information 114.
[0032] As illustrated, in this example, data processing system 102 communicates with storage device 106. Storage device 106 may be an external server, a distributed storage / computing environment (e.g., a cloud storage system), or any other type of storage device or system that communicates with data processing system 102. Although shown external to data processing system 102, it should be understood that storage device 106 may be part of data processing system 102 or may be internal to data processing system 102. Storage device 106 may be accessed when storing (e.g., provided by one or more client devices 122) multichannel audio data 108, generating training dataset 110, or performing any other operations described herein.
[0033] As described herein, conventional multimodal or audio-specific language models are not configured to process multichannel audio data 108, but are only configured to process single-channel (e.g., mono audio) data and / or use single-channel processing techniques to process multichannel audio data (which results in a loss of spatial reasoning or understanding). To address these issues, data processing system 102 can train / update one or more language models 120 to process multichannel audio data 108. Multichannel audio data 108 can include any type of digital signal representing an audio recording with two or more channels. An audio channel refers to a single path or stream carrying digital information of an audio signal, such as sound waves captured by a microphone or sound waves generated electronically. In the context of multichannel audio data, each channel can typically represent an independent audio capture device or audio component throughout the audio recording.
[0034] One example of multichannel audio is stereo audio, which includes a left channel and a right channel. Another example includes "surround sound," such as 5.1 surround sound, which includes audio channels such as front left, center, front right, rear left, rear right, and bass audio channels. In other examples, audio data can be collected from any number of microphones and / or microphone arrays distributed in any configuration / or orientation within the environment. In one example, multichannel audio data may include multiple individual "tracks" that are combined together to create a mixed audio signal. Each audio channel in the multichannel audio data 108 may have its own unique characteristics, such as gain level, frequency response, and / or spatial location in three-dimensional space.
[0035] Multichannel audio data 108 may include any number of multichannel audio samples, each of which may be stored in a corresponding file or data structure. Any suitable format may be used to store the multichannel audio data 108, including but not limited to WAV files, Audio Exchange File Format (AIFF) files, MPEG Audio Layer (MP3) files, Broadcast Waveform Format (BWF) files, Advanced Audio Codec (AAC) files, or Free and Free Lossless Audio Codec (FLAC) files. Each sample of the multichannel audio data 108 may be stored associated with various metadata, including characteristics such as the number of audio sources, information about how the multichannel audio samples were generated, a text-based description of the multichannel audio samples, or other information related to the multichannel audio samples. In some implementations, data processing system 102 may receive multichannel audio data 108 from client device 122 and subsequently store it in storage device 106 for processing according to the techniques described herein.
[0036] Samples of the multi-channel audio data 108 can be captured using any type of suitable equipment. In one example, one or more microphone arrays can be used to capture samples of the multi-channel audio data 108. The microphone array can include a configuration of microphones or other devices capable of capturing sound, arranged in a specific configuration to capture audio from multiple directions. Each signal captured using the microphones in the microphone array can be stored as a corresponding audio channel in a sample of the multi-channel audio data 108. Example configurations of the microphone array include, but are not limited to, linear arrays, circular arrays, or spherical arrays.
[0037] In some implementations, the multichannel audio data 108 can be stored in a stereo surround sound format, sometimes referred to herein as "B-format" audio. B-format audio data is a type of multichannel audio represented as a three-dimensional sound field. B-format audio can include four channels: W (omnidirectional), X (front-to-back), Y (left-to-right), and Z (up-down). In some implementations, audio can be recorded using a multichannel microphone array and converted to B-format audio data, which is then stored as a sample of the multichannel audio data 108. In some implementations, a stereo surround sound microphone array can be used to directly capture B-format audio data, which can then be stored as part of the multichannel audio data 108.
[0038] Unlike other multichannel audio formats, B-format audio includes a fixed number of channels to represent global surface audio. In contrast, other types of multichannel audio, such as stereo or surround sound, can include any number of audio channels, making such formats incompatible with the fixed input of the multimodal language model 120. Other formats can represent audio data in a microphone array-specific channel arrangement, which poses challenges to compatibility with the language model 120 due to the variety of possible audio formats and microphone arrangements. By representing multichannel audio in a microphone array-independent format, using B-format (e.g., stereo surround sound) audio data solves these limitations. B-format audio uses a fixed number of channels to represent global surface three-dimensional audio in a manner independent of the device used to capture three-dimensional audio.
[0039] Using a fixed number of channels (e.g., W, X, Y, and Z channels in B format) enables direct encoding (e.g., tokenization) of the multichannel audio data 108 and its use as input to one or more language models 120, regardless of the type or arrangement of the device used to capture the multichannel audio data 108. The multichannel audio data 108 can be encoded by the device used to capture the audio data (e.g., client device 122) or by the data processing system 102. In some implementations, microphone array-specific transformation functions can be used to convert the audio signal captured using the microphone array into B format audio for inclusion in the multichannel audio data 108. The multichannel audio data 108 can be used to train / update the language model 120 according to the techniques described herein.
[0040] In some implementations, the multi-channel audio data 108 may include high-order stereo surround sound (sometimes referred to as high-order B-format) rather than typical four-channel B-format audio data. High-order B-format may include extensions to conventional B-format data that provide high-precision capture and reproduction of spatial information within the sound field. For example, microphones may be attached to capture high-order B-format to record high-order coefficients that describe characteristics of the recorded sound field beyond those captured by a standard B-format microphone array.
[0041] The number of audio channels in a higher-order B-format can depend on the specific order of the B-format audio. For example, the number of channels can increase quadratically with the order (N) of the B-format data. Traditional first-order B-format audio includes 4 channels, second-order B-format audio includes 9 channels, third-order B-format audio includes 16 channels, and so on. The language model 120 may include an input layer that receives encoded audio data (e.g., encoded audio sample 112, encoded input data 119) generated from any type of B-format audio described herein.
[0042] In some implementations, storage device 106 may store different sets of multichannel audio data 108, each set of multichannel audio data 108 including B-format data of a specific order (e.g., first-order B-format audio from one corpus, second-order B-format audio from another corpus, etc.). In some implementations, each language model 120 may be trained / updated using training dataset 110 constructed from multichannel audio data 108 having a specific order (e.g., first-order, second-order, etc.) of B-format. In some implementations, language model 120 may be trained / updated to process encoded audio generated from multi-order B-format data (e.g., both first-order and second-order B-format data, etc.). As described herein, training / updating of one or more language models 120 may be performed using one or more training datasets 110.
[0043] Data processing system 102 can use multi-channel audio data 108 to generate training datasets 110 for one or more language models 120. In some implementations, each training / update sample in the training dataset 110 may include at least one encoded audio sample 112 and corresponding spatial information 114. The encoded audio sample 112 may be an encoded form of the corresponding sample of the multi-channel audio data 108. Each encoded audio sample 112 may be encoded in a format compatible with one or more language models 120. In some implementations, data processing system 102 can use one or more tokenizers 118 to generate the encoded audio samples 112. The one or more tokenizers 118 used to generate the encoded audio samples 112 may correspond to language models 120 and may form part of language models 120 to be trained using the corresponding training dataset 110.
[0044] One or more tokenizers 118 may include one or more audio tokenizers that can encode microphone array-independent spatial audio (e.g., B-format / stereo surround sound audio) into a format compatible with one or more language models 120. In some implementations, one or more tokenizers 118 may be executed by data processing system 102 to encode each channel of B-format audio in samples of multi-channel audio data 108. For example, as described herein, first-order B-format audio data may include four channels: W, X, Y, and Z. To tokenize the first-order B-format data, each channel of samples of multi-channel audio data 108 may be encoded separately using a corresponding audio codec model.
[0045] In one example, each audio codec model may include a non-autoregressive convolutional encoder-quantizer-decoder model for audio codec extraction. Such audio codec models may include any number of convolutional layers, recursive layers, or other machine learning layers capable of encoding one or more windows of audio data. The audio codec models may be trained or utilized to process specific channels of first-order (or higher-order) B-format audio data. Encoding samples of multi-channel audio data 108 may include providing each channel of the sample as input to the corresponding audio codec model to generate a sequence of output tags, each output tag corresponding to a specific timestep in the audio data for a given channel. Continuing the example above, each timestep might produce four tags, with one tag corresponding to the W channel, a second tag to the X channel, a third tag to the Y channel, and a fourth tag to the Z channel. In some implementations, an audio codec for higher-order channels may be used to generate more tags within a given timestep.
[0046] The tags generated by the audio codec model can encode any information represented in the audio at a corresponding time step. A sequence of tags generated from the entire sample of multi-channel audio data 108, when provided sequentially, can represent the samples in an encoded format. Each tag can contain a digital representation of the audio information in a given channel within a given time step. In some implementations, multi-channel audio tags can be generated by concatenating or otherwise grouping the tags generated from the samples of multi-channel audio data within a given time step. The sequence of multi-channel audio tags representing the samples of multi-channel audio data 108 can be stored as encoded audio samples 112 as part of the training dataset 110.
[0047] Each encoded audio sample 112 of the training dataset 110 may be stored together with corresponding spatial information 114. In some implementations, each encoded audio sample 112 may also be stored in association with corresponding supplementary media information 115. Spatial information 114 may include textual information indicating various spatial characteristics of the audio stored as encoded audio samples 112 in the training dataset 110. For example, spatial information 114 may include information relating to one or more audio sources represented in the audio sample, including but not limited to the number of directional sources, the azimuth / elevation angle of each source, and / or the distance from each source. In another example, spatial information may include information relating to the speaker (e.g., individual speaker) in the audio sample, including but not limited to the number of individual speakers in the audio sample (e.g., the number of speech sources), the speaker's transcription regarding a given direction (e.g., transcription of any speaker in the left, right, up, or down direction, etc.).
[0048] Spatial information 114 may include characteristics of different sources of audio, including but not limited to the relative volume of different sources represented in the encoded audio sample 112, changes in pitch of different sources represented in the encoded audio sample 112, changes in position / direction of different sources represented in the encoded audio sample 112, general attributes of different sources represented in the encoded audio sample 112 (e.g., timbre, classification, duration, intonation, etc.) and / or any other possible spatial characteristics of the corresponding audio sample.
[0049] In some implementations, spatial information 114 may be represented as text data that can be used as the corresponding target output of a multimodal language model 120. For example, training dataset 110 may include task-specific training / update examples for one or more language models 120 used to train / update them to perform task-specific operations related to spatial audio. Continuing this example, spatial information 114 may include text data representing input prompts (e.g., text-based requests or prompts) and / or output prompts representing the expected response that the language model 120 will generate based on the input prompts. This can be any type of text data that completes the corresponding input prompt. In one example, the input prompt may include “count the number of audio sources in the audio sample,” and the output prompt may include “there are three audio sources in the audio sample.” In the foregoing example, spatial information 114 includes an indication that the encoded audio sample 112 includes three separate audio sources. For any type of task specific to spatial audio, similar prompts may be provided as part of spatial information 114.
[0050] For example, input / output cue pairs may be included in the training dataset 110 in association with corresponding encoded audio samples 112 for targeted transcription, such as “transcribe the speaker speaking from the left” or “transcribe the speaker speaking from the right”. Continuing this example, the corresponding output cue could be a transcript of the speaker most obviously heard from the left and right portions of the corresponding spatial audio encoded as encoded audio sample 112. Various other input / output cue pairs may be provided as part of the spatial information 114 of one or more training / updating examples, for example, modeling user requests paired with corresponding outputs for learning by one or more language models 120 according to the techniques described herein.
[0051] In some implementations, the encoded audio samples 112 of the training dataset 110 may include additional media information 115, which may include any type of output text, video, audio, image, or other information to be generated by the language model 120. As described herein, the language model 120 may be a multimodal model capable of ingesting and / or generating media of various modalities. Such modalities may include, but are not limited to, text data, audio data, image data, video data, or combinations thereof. In some implementations, additional media information 115 may be generated for training / updating examples of the training dataset 115 to include output audio data to be generated by the trained / updated language model 120.
[0052] In one example, the output audio data may include isolated audio of a speaker having a discernible spatial location represented in the encoded audio sample 112. Continuing with this example, the spatial information may include a text input / output cue requesting transcription and isolation of one of the many speakers in the encoded audio sample 112. Additional media information 115 may include an encoded representation of the audio (e.g., a tagged sequence) including the requested isolated speaker, and spatial information 114 may include a text output cue including the transcription of the isolated speaker.
[0053] In another example, the output audio data in the supplementary media information 115 of the training / update example may include spatial audio (e.g., tokenized B-format audio) generated as a response or reply to input text or input audio prompts. For example, in some implementations, the encoded audio sample 112 may be an instruction from a user, such an instruction describing a query or describing the output to be generated by the language model 120. The supplementary media information 115 for this type of training / update example may include encoded audio data representing the output that the language model 120 will generate based on the input query in the encoded audio sample 112. In some implementations, the input text prompt may also specify parameters for generating spatial audio. Since the encoded output in the training / update example is multi-channel audio data, the user can request any desired spatial characteristics of the generated output, including the position, movement, or orientation of one or more audio sources.
[0054] Data processing system 102 can generate training dataset 110, which can be used to train / update one or more language models 120 to process user requests in audio data (e.g., recorded using one or more microphones or capture devices). For example, training / update examples in training dataset 110 may include examples of one or more speakers moving relative to a corresponding sample of the recording multichannel audio data 108. Such encoded audio samples 112 may include background noise or other speakers to be distinguished from the moving speaker. In one example, spatial information may include input / output cues that request and provide transcribed and / or logged output from the moving speaker, allowing one or more language models 120 to learn to distinguish the moving speaker from other noise or speakers in the environment. Such training examples can be used to enable language models 120 to automatically resolve confusion when multiple speakers are present or when a speaker is moving relative to the multichannel recording device.
[0055] Various training / update examples may be included to train / update one or more language models 120 to track or otherwise focus on user requests in moving audio data (e.g., in the presence of a fixed audio source from the environment). Training / updating one or more language models 120 to accurately capture, transcribe, or otherwise respond to user requests in such audio data facilitates the use of one or more language models 120 in conjunction with audio captured in noisy or dynamic environments.
[0056] In some implementations, additional media information in the training / update examples of training dataset 110 may include video data corresponding to the encoded audio sample 112. In one example, the video data may be encoded / tagged video data depicting an object or event corresponding to the sound source of the encoded audio sample 112. Continuing with this example, encoded video data may be used as input data when training / updating one or more language models 120 to generate the corresponding encoded audio sample 112. Training dataset 110 may be generated to provide generative capabilities for spatial audio, such that one or more language models 120 are trained / updated to generate three-dimensional audio data corresponding to and synchronized with the video data. In some implementations, spatial information 114 for such audio samples may include indications of the location of one or more sound sources in the video data. In some implementations, spatial information 114 may not necessarily be included in such training / update examples, thereby enabling one or more language models 120 to learn to map changes in the location of objects / events in the video data to the location of sound sources in the output audio using the corresponding encoded audio samples 112 as ground truth output data.
[0057] In another example, the encoded audio data sample 112 may include mono audio (e.g., having a single source). Additional media information 115 and spatial information 114 may specify the relative positions of one or more sound sources represented in the mono audio to be represented in the output spatial audio. Continuing this example, such a training / update example may include encoded mono audio samples to be used as input for one or more language models 120 and corresponding encoded audio samples 112 representing the spatial audio to be generated by one or more language models 120.
[0058] The training / update examples in training dataset 110 may also include examples for improving audio quality using the generative capabilities of one or more language models 120. In one example, coded audio sample 112 may include single-channel or multi-channel audio data of relatively poor quality—including, but not limited to, noisy backgrounds, audio artifacts or distortions, and other interferences. Video / image / sensor (e.g., LiDAR, RADAR, sonar, ultrasound, etc.) data may also be included, for example, to visually represent the different sound sources of the input coded audio sample 112. Continuing with this example, additional media information 115 of the training / update examples may include spatially improved coded audio data. The improved audio data may be used during the update / training of one or more language models 120 as a comparison with the data to be generated by one or more language models 120. The improved coded audio data may represent spatial audio without distortion, artifacts, or noise. The improved coded audio data may improve the spatial mapping between sound sources, as indicated by any input video data. In some implementations, such training / update examples can be generated by comprehensively reducing the quality of spatial audio (which may correspond to video data).
[0059] Other training / update examples included in training dataset 110 may include training examples for virtual, augmented, and / or mixed reality systems. In one example, encoded audio samples 112 may be generated based on audio captured from one or more virtual reality or augmented reality headsets or devices. Such audio samples may include sounds from different environments, which may include different objects, speakers, or environmental hazards. Such training / update examples can be used to distinguish the speech of a user of a virtual reality or augmented reality system in noisy or dynamic soundscapes. In such implementations, such training / update examples of training dataset 110 may include similar spatial information 114 as described herein (e.g., text data, request information, output responses, etc.), or may include audio responses generated by one or more language models 120 as described herein.
[0060] Other training / update examples may include examples for training / updating one or more language models 120 to generate spatial audio for accessibility purposes (e.g., in an augmented reality or virtual reality context). For example, training / update examples may include input video data for one or more language models 120 paired with corresponding output encoded audio samples 112 that provide auditory warnings or indications in the environment. In another example, training / update examples in training dataset 110 may include examples where spatial audio from augmented reality devices, virtual reality devices, or accessibility devices (e.g., smart glasses, hearing aids, etc.) can be used to train / update one or more language models 120 to generate visual indications of oncoming obstacles, environmental hazards, or other objects in the environment. Truth data corresponding to these indications may be provided as part of the spatial information and may include directional notifications, warning signals, or alarms indicating potential hazards detected in the corresponding encoded audio samples 112 of the training / update examples.
[0061] The data processing system 102 can generate any number of training datasets 110 to achieve any of the training / update objectives described herein. To generate the training dataset 110, the data processing system 102 can receive or access individual samples of multi-channel audio data 108 in a storage device and can associate said samples with corresponding spatial information 114 and / or additional media information 115. In some implementations, spatial information 114 and / or additional media information 115 can be retrieved from one or more data sources corresponding to the multi-channel audio data 108. In some implementations, synthetic data generation techniques can be used to generate one or more of the spatial information 114 and / or additional media information 115. Such techniques may include the execution of generative artificial intelligence models, including large language models or visual language models.
[0062] In some implementations, while communicating with the data processing system 102, one or more samples of multichannel audio data 108 can be received from one or more external computing devices (such as client device 122). Spatial information 114 and / or additional media information 115 can also be provided by the external computing system, or can be generated using other techniques such as manual generation / annotation. Various combinations of techniques can be used to generate spatial information 114 and / or additional media information 115 for samples of multichannel audio data 108.
[0063] To generate the training dataset 110, the data processing system 102 may use one or more tokenizers 118 to generate encoded audio samples 112 for each sample of the multi-channel audio data 108 provided for the training dataset 110. This may include providing each channel of the multi-channel audio data 108 as input to a corresponding model, which is trained / updated to generate a sequence of output tokens, each output token corresponding to a specific time step in the audio data for a given channel, as described herein. The sequence of tokens for each channel may be combined into one or more data structures to correspond to the encoded audio samples. In some implementations, the data processing system 102 may execute a similar tokenizer to tokenize various other media to be provided as input to one or more language models 120 during training. Such tokenizers 118 may include video tokenizer models that are trained / updated to tokenize one or more frames of video data before initiating the input sequence; or text-based tokenizers that are trained / updated to segment and tokenize text data included in the additional media information 115 for each training / updated example in the spatial information 114 and / or training dataset 110.
[0064] Once the data has been tokenized, the data processing system 102 can store the encoded audio samples 112 in association with corresponding spatial information 114 and / or additional media information 115 in one or more data structures. Each group including the encoded audio samples 112, spatial information 114, and / or additional media information 115 can be stored in association with the identifier of its corresponding training / update sample. In some implementations, the training dataset 110 can be stored in association with the identifier of the training / update objective for training the dataset 110 (e.g., to train / update the language model 120 to learn spatial relationships, isolate the speaker represented in the audio data, generate spatial audio for video data, generate improved audio / spatial quality based on audio / text input, etc.). In some implementations, the data processing system 102 can generate multiple training datasets 110 for multiple training / update objectives. In some implementations, the data processing system 102 can generate training datasets 110 to include training / update examples for multiple training / update objectives.
[0065] As shown, the data processing system 102 can maintain, execute, and train / update one or more language models 120. The one or more language models 120 can include any type of multimodal language model capable of handling natural language text input, audio input, video input, or image input, as well as other media modalities. The one or more language models 120 can be or include transformer-based models (e.g., generative pre-trained transformer (GPT) models). In some implementations, the one or more language models 120 can be or include large language models (LLMs) or visual language models (VLMs). In some implementations, the one or more language models 120 can be one or more tokenizers 118 capable of converting media data into an encoded format (e.g., one or more tagged or “tagged” formats) compatible with the layers of the one or more language models 120.
[0066] In some implementations, the data processing system 102 can maintain, store, update, and / or deploy multiple language models 120. For example, different language models 120 may include different media processing capabilities (e.g., one language model 120 can process video data, while another language model 120 can process audio and text data, etc.). In some implementations, one or more different language models 120 can be trained / updated according to different training / update objectives by using one or more corresponding training datasets 110.
[0067] Data processing system 102 can use model updater 116 to train / update language model 120. In one example, language model 120 can be trained / updated in response to a corresponding request received from an external computing device or in response to input received from an operator of data processing system 102. Model updater 116 can include any software, hardware, or combination thereof to perform training / update operations on one or more language models 120 as described herein. A request to train / update language model 120 can instruct one or more training datasets 110 for use in training / updating one or more language models 120. In some implementations, training datasets 110 can be automatically identified or selected based on one or more training / update objectives specified in the request (e.g., by selecting training datasets 110 with training / update objectives that match those specified in the request, etc.).
[0068] To train / update the language model 120 using the training dataset 110, the model updater 116 can iterate over each training / update example in the training dataset 110 based on hyperparameters of the training / update process (e.g., number of epochs, batch size, etc.), which can be specified via a request to train / update the language model 120 or via a configuration environment. For each training / update example, the model updater 116 can generate a context for the language model 120 to be trained. Generating the context may include concatenating tokenized input data (e.g., encoded audio samples 112, any encoded supplementary media data 115, encoded text cue data, etc.) with encoded output data into a single sequence. Special encoded markers can be used in the input context to specify the start and end of media modalities, allowing the language model 120 to learn to depict different types of input data. In some implementations, positional encoding or other relevant embeddings can be added to the context to preserve the order of certain input / output data in the sequence and to distinguish between input and output segments of the context.
[0069] Then, the model updater 116 can apply an attention mask (e.g., cross-attention, self-attention, etc.) to the context, causing the language model 120 to focus only on the encoded input data (e.g., input tokens). The attention mask can include replacing the masked tokens with special tokens that indicate that the encoded data should not be focused on. The masked data in the context can be data that the language model is to predict during training. Such an attention mask can instruct the language model 120 to use the encoded input when predicting each token in the output sequence. The attention mask can be applied in any suitable masking mode. In some implementations, the attention mask can be applied to the encoded data in the generated context, which represents the output data that the language model 120 is to generate. For example, if the language model is being trained / updated for a generative task, the attention mask can be applied to the tokens in the context representing the output audio data (e.g., the encoded audio sample 112 used as ground truth data).
[0070] In training / update iterations, the model updater 116 can execute the language model 120 by passing a sequence of encoded context data to each layer of the language model 120 while performing mathematical / machine learning operations at each layer. The output of the language model 120 may include a distribution of candidate token outputs from which one or more output tokens are selected. In some implementations, the outputs can be predicted in an autoregressive manner, where the model updater 116 appends the predicted output tokens to the initial context to generate an extended context. The extended context is then provided as input to the language model 120 until all output tokens have been predicted. In some implementations, a "teacher forcing" technique can be used, where ground truth tokens (rather than the model's own predictions) from the output portion of the context sequence are appended to the initial input context for predicting the next token. In some implementations, the language model 120 can generate tokens in a non-autoregressive manner, where the language model 120 is executed to predict all output tokens simultaneously.
[0071] In some implementations, language model 120 may generate multiple output tags in an autoregressive manner. For example, language model 120 may include layers of tags predicted simultaneously for each channel of a single time step in B-format audio. For example, in first-order B-format audio, language model 120 may simultaneously generate output tags for each of the W, X, Y, and Z channels within a single time step. In some implementations, a larger number of tags may be generated simultaneously for higher-order B-format audio. In some implementations, language model 120 may generate only a single tag in an autoregressive manner per iteration.
[0072] Model updater 116 can use a loss function (such as cross-entropy loss) to compare the ground truth labels of the training / update examples with each output label predicted by language model 120 to quantify the difference between the predicted and actual labels. In one example where cross-entropy loss is used, model updater 116 can compare the predicted probability distribution (e.g., a softmax function) of the output of language model 120 with the one-hot encoded true distribution representing one or more actual next labels in the output sequence. Model updater 116 can compute the cross-entropy loss as the negative log probability of the ground truth label based on the predicted distribution of language model 120. Model updater 115 can compute the total loss used for the training / update sequence as the sum (or, in some implementations, the average) of the cross-entropy losses at all label positions in the output sequence predicted by language model 120. In some implementations, similar methods can be used to compute other types of loss functions.
[0073] Model updater 116 can use backpropagation to train / update the parameters of language model 120 using the computed loss. Backpropagation can be performed by computing the gradient of the loss with respect to each parameter and adjusting the parameters in the direction that minimizes the loss. Parameter tuning can be performed using a suitable optimization function, such as a gradient descent function or an Adam optimizer function. Model updater 116 can iteratively repeat this process using a number of training / update samples from one or more training datasets 110 until a training / update termination condition (such as an accuracy threshold being met) is met or when training / updating one or more language models 120 using a predetermined number of training / update samples.
[0074] As described herein, training / update examples can be provided for training / updating language model 120 to achieve various objectives. In some implementations, training / update examples can be provided to train / update language model 120 to generate output text data that indicates various spatial properties of one or more audio sources represented in the input multichannel audio. In some implementations, the output text can identify the distance to the audio sources represented in the input multichannel audio, the number of audio sources represented in the input multichannel audio, or the transcription of speech from moving audio sources represented in the input multichannel audio.
[0075] In some implementations, the language model 120 can be updated to process spatial audio for generative tasks. For example, a model updater 116 can use training / update examples from one or more training datasets 110 to update the language model 120 to generate output multichannel audio from input audio. Such objectives can include noise reduction, improving spatial quality, or converting single-channel audio to multichannel audio based on instructions or input video data, as described herein. In examples where video / image / sensor data is used, the training / update examples can include an encoded representation of the video / image / sensor data (e.g., as part of additional media information 115, etc.), and the model updater 116 can use the encoded representation of the video / image / sensor data as part of the input sequence in the generated context for the training / update examples. The model updater 116 can generate a context such that the ground truth output data includes encoded multichannel audio (e.g., encoded audio sample 112) that represents audio tracking of at least one audio source depicted in the video / image / sensor data. For example, video / image / sensor data can depict travel from left to right, and the corresponding encoded audio sample 112 can be an audio sample of train noise moving from left to right, synchronized with the video / image / sensor data.
[0076] In another example, training / updating examples may include an encoded representation of video / image / sensor data (e.g., as part of additional media information 115, etc.), and model updater 116 may use the encoded representation of video / image / sensor data as part of the input sequence in the generated context for training / updating examples. Additionally, the input context may include encoded single-channel audio data corresponding to and synchronized with the video / image / sensor data. To enable language model 120 to learn to convert single-channel audio for video into multi-channel audio, model updater 116 may generate a context such that ground truth output data includes encoded multi-channel audio (e.g., encoded audio sample 112), which represents a multi-channel version of the input single-channel audio synchronized with the video / image / sensor data. Similar methods can be used to train / update language models using any type of multimodal data related to multi-channel audio, according to any suitable objective.
[0077] In some implementations, once trained / updated, the language model 120 can be executed to generate model output 126 in response to receiving input prompts (e.g., input data 124) from one or more client devices 122. System 100 is shown to include client devices 122, which may include one or more input / output devices such as microphones, video / image / sensor data capture devices (e.g., integrated cameras, LiDAR, RADAR, ultrasound, sonar, etc.) and text input devices (e.g., touchscreens, keyboards, AR / VR / MR devices, gesture recognition systems, etc.). Client devices 122 may include any type of device capable of communicating with data processing system 102 (e.g., via one or more networks), including but not limited to smartphones, laptops or mobile computers, augmented and / or virtual reality devices, digital assistant devices, accessibility devices (e.g., hearing aids or appliances, etc.), personal computers, servers, cloud computing systems, in-vehicle or cockpit infotainment systems, or other types of computing systems capable of providing input data 124 to data processing system 102. In some implementations, the client device 122 may include one or more communication interfaces that can send input data 124 to one or more external computing systems, which may include a data processing system 102.
[0078] Input data 124 may include any type of data that can be provided as input to one or more language models 120. This type of data includes, but is not limited to, text data, multi-channel audio data, single-channel audio data, video data, image data, sensor data, 2D or 3D design or graphic data, etc. In some implementations, input data 124 may be captured via one or more input devices of client device 122. In some implementations, input data 124 may be stored in one or more input devices at client device 122. In some implementations, client device 122 may execute one or more applications that enable a user to provide text input, capture audio, or capture video as input to language model 120. In some implementations, such applications may include augmented reality or virtual reality applications. In some implementations, the application may include a front-end for a dialogue agent.
[0079] Input data 124 generated or retrieved by client device 122 can be sent to data processing system 102 for processing using trained / updated language model 120. In some implementations, input data 124 can be provided via input from an operator of data processing system 102. Upon receiving input data 124, data processing system 102 can execute one or more tokenizers 118 (e.g., each tokenizer corresponds to a corresponding channel of multi-channel audio data or to a different media modality, etc.) to generate encoded input data 119. Encoded input data 119 may include one or more sequences of tags that represent input data 124 in a numerical format compatible with one or more input layers of trained / updated language model 120.
[0080] For example, if input data 124 includes B-format audio data, data processing system 102 can execute tokenizer 118, which converts each channel of the B-format audio into a sequence of tokens that can be formatted for the input context of one or more language models 120, as described herein. Further, one or more tokenizers 118 can be executed to convert other types of media into encoded formats for inclusion in encoded input data 119. Generating encoded input data 119 can include: generating encoded video data using a video-specific tokenizer 118, generating encoded text data using a text-specific tokenizer 118, or generating encoded single-channel audio data using a tokenizer 118 specific to processing single-channel audio. In some implementations, data processing system 102 can update encoded input data 119 to include additional markers indicating the start and end of different sequences corresponding to different media modalities. For example, data processing system 102 can provide corresponding start / stop markers for sequences of encoded audio data, sequences of corresponding encoded text data, and / or sequences of encoded video data.
[0081] Once encoded input data 119 is generated, data processing system 102 can execute language model 120 by providing the encoded input data 119 to one or more input layers of language model 120. Data processing system 102 can perform mathematical operations on each layer of the language model, propagating the results of each layer to the next layer for processing, until (e.g., from the output softmax layer, etc.) one or more output distributions of token probabilities are generated. Data processing system 102 can use one or more configuration environments to select one or more tokens from one or more output distributions to include in the output response. Data processing system 102 can execute large language model 120 in an autoregressive manner to model a sequence of output tokens corresponding to one or more media modalities, including multi-channel (spatial) audio data, video data, and / or text data. For example, data processing system 102 can execute language model 120 to predict one or more next tokens in the output sequence, which can then be included in the input context for the next iteration, as described herein.
[0082] The data processing system 102 can iteratively execute the language model 120, merging previously generated tags into a context for generating subsequent tags, until a termination condition is met. One type of termination condition may be a context length limit or a configurable limit on the number of tags that the language model 120 can generate and / or process. In some implementations, the termination condition may be satisfied when the language model 120 generates a tag indicating the end of a response. In some implementations, the language model 120 may be trained / updated as a dialogue agent. For example, the language model 120 may generate realistic natural language in response to natural language input, which may be in the form of audio data representing natural human speech.
[0083] Once a termination condition for executing language model 120 has been detected, data processing system 102 can convert the encoded output generated by language model 120 into a decoded format for transmission to a client device. In some implementations, this may include performing the inverse operation of the tokenization process. For example, in some implementations, one or more tokenizers 118 may include one or more detokenizer models trained / updated to convert the numerical tokens generated by language model 120 into corresponding media modalities. In one example, data processing system 102 may execute a tokenizer model that converts sequence tokens representing channels of B-format audio into B-format audio samples.
[0084] Similar operations can be performed by data processing system 102 to generate decoded text data and / or video data for inclusion in model output 126. For example, text data can be generated by de-tagging text-specific tags generated using language model 120, and video data can be generated by de-tagging video-specific tags generated using language model 120. In some implementations, media-specific tags can be extracted based on media-type-specific start / stop sequence tags generated by language model 120. Output text, video, and / or audio generated by language model 120 can be provided as part of model output data 126.
[0085] Model output 126 may include text data generated using a large language model 120, which may include text-based responses generated from one or more input multichannel audio samples. For example, the text data may specify data indicating spatial information of the audio in the model, including but not limited to distances from one or more audio sources represented in the input audio, the number of audio sources represented in the input audio, or transcriptions of speech from moving audio sources represented in the input audio.
[0086] Model output 126 may include multi-channel audio data 108 generated by language model 120. In one example, in addition to audio data, data processing system 102 may also provide encoded video data as input to the language model (e.g., as part of encoded input data 119) and generate an output indicating spatial information based on the input audio and video. Continuing with this example, the input audio may include single-channel audio, and the data processing system may generate multi-channel audio as an output that converts the single-channel audio into multi-channel spatial audio. Spatial audio may include the content of a single-channel audio track that is correctly mapped in a three-dimensional sound space to at least one audio source depicted in the video data. For example, the video may depict movement from left to right, and the corresponding output audio may include audio samples of train noise (represented in single-channel audio) moving from left to right in sync with the video.
[0087] A similar approach can be used to generate model output 126, which includes audio samples with improved audio quality. In some implementations, the sound / spatial quality of spatial / multichannel audio captured using an end-user device (e.g., a smartphone, etc.) may be poor. The data processing system can use the poor-quality input audio as input to execute language model 120 to generate output multichannel audio samples with improved spatial quality (e.g., attenuating distortion and background noise, amplifying audio in the speaker's direction, etc.). In some implementations, model output 126 may include one or more alerts or messages indicating potential hazards detected in the input audio data for use with accessibility-based devices, such as visual accessibility devices or augmented / virtual / mixed reality devices for hearing-impaired individuals.
[0088] Model output 126 can be provided to client device 122 for presentation to a user. In one example, text data in model output 126 can be displayed in one or more applications (e.g., using the display of client device 122), which are executed on client device 122. Audio data included in model output 126 can be sent to client device 122 so that it can be played via one or more audio output devices (such as the integrated speaker of client device 122). In some implementations, video data sent as part of model output 126 can be presented via one or more display devices of client device 122 (in conjunction with any associated audio data). Any model output 126 sent to client device 122 can be stored in the memory of client device 122 for later access or processing. In some implementations, model output 126 can be stored in the memory of data processing system 102 in association with input data 124. In some implementations, for example, this information can be used to generate a future training dataset 110 using reinforcement learning techniques.
[0089] In some implementations, the client device 122 or the data processing system 102 may store / maintain records of input data 124 and corresponding model outputs 126 in sequence, allowing the data processing system 102 to provide a dialogue agent using one or more language models 120. In such implementations, the data processing system 102 may provide the client device 122 with one or more web-based interfaces for the dialogue agent, through which a user can provide input data 124 using the input / output devices of the client device 122. Using these techniques, the data processing system 102 can train / update and execute the language models 120 to process spatial / multichannel audio as a dialogue agent.
[0090] In various embodiments where one or more language models 120 are used to deploy the dialogue agent, the dialogue agent may be deployed as a non-player character (NPC) in a video game, such as, for example, a video game locally managed (e.g., using a computing device or game console) and / or remotely managed using a content streaming platform or service (e.g., NVIDIA's GeForce Now). In other embodiments, the dialogue agent may be deployed in a vehicle or other machine type, such as part of an in-cabin infotainment system and / or as a digital assistant within the vehicle or machine (e.g., to help control components and / or features of the vehicle or machine—such as windows, doors, audio / video playback, navigation, etc.). In some embodiments, the dialogue agent may be deployed alongside a digital avatar, digital human, or robot to allow the avatar, human, or robot to use spatial awareness to converse with the user in the environment, which is recognized using one or more language models 120. In some embodiments, the dialogue agent may be deployed on a stationary object (such as a screen of a conversation / smart kiosk) or on a moving object (such as a robot). In any example, renderings of the dialogue agent can be generated in a simulation platform and / or a collaborative content generation and sharing platform for digital assets, such as a platform that uses generic scene descriptor (USD) data (e.g., NVIDIA's OMNIVERSE). For example, a simulation platform can be used to generate renderings of digital humans, digital avatars, etc., and stream them for display on end-user devices.
[0091] refer to Figure 2 ,Should Figure 2 The illustration shows a data stream diagram 200 illustrating how a spatially aware audio-enhanced dialogue agent can process different types of data according to some embodiments of the present disclosure. As described herein, this can be achieved, for example, through... Figure 1 The data processing system 102 executes the process shown in the data flow diagram 200. As described herein, a multimodal language model 212 (e.g., one or more language models 120) can be trained / updated to process different types of spatial audio data and associated data, including but not limited to text input data 204 and video input data 214.
[0092] In this example, the multimodal language model 212 is trained / updated to process audio input data 202, video input data 214, and text input data 204. However, it should be understood that other configurations are possible, and the multimodal language model 212 can be trained / updated to process any type of input data and combinations thereof. Continuing with this example, the audio input data 202 is multichannel audio data in a format different from B-format audio (e.g., stereo surround sound audio). To generate B-format audio compatible with the multimodal language model 212, a spatial transformation function 206 can be used. The spatial transformation function 206 can be specific to the apparatus used to capture the microphone array. The number of audio channels included in the audio input data 202 can correspond to the number of microphones in the capture device used to capture the audio input 202. The spatial transformation function 206 can be a function of the relative distance between each microphone, the gain of each microphone, and other properties of the device used to capture the audio input data 202. In some implementations, the spatial transformation function 206 may include multiplying the transformation matrix by the audio channels in the audio input data 202 to generate a set of B-format audio output channels.
[0093] As shown, once generated by the spatial transformation function 206, B-format audio (which in some implementations may include higher-order B-format audio) can be provided as input to the multi-channel audio encoder 208. The multi-channel audio encoder 208 may include a set of tokenizers (e.g., a set of tokenizers 118) trained / updated to generate an encoded representation of the B-format audio output from the spatial transformation function 206. For example, each tokenizer in this set may correspond to a channel of the B-format audio and generate encoded data for that channel. Continuing the example, a first tokenizer may process the signal for the W channel of the B-format audio, a second tokenizer may process the signal for the X channel of the B-format audio, a third tokenizer may process the signal for the Y channel of the B-format audio, and a fourth tokenizer may process the signal for the Z channel of the B-format audio. As described herein, each token output by the tokenizer may correspond to a corresponding time step; and when provided sequentially, it may represent an encoded representation of the audio input data 202 compatible with the multi-channel audio encoder 208.
[0094] As shown, the output of the multi-channel audio encoder 208 can be provided as input to the multimodal language model 212, for example, as part of the input context for the multimodal language model 212. In some implementations, a sequence of markers representing the encoded audio input data 202 can be identified by special start / stop markers in the input context. As described herein, the multimodal language model 212 can be trained / updated to receive additional data from different media modalities. As shown, the multimodal language model 212 is trained / updated to receive video input data 214 and text input data 204.
[0095] To convert video input data 214 and text input data 204 into a format compatible with the multimodal language model 212, a video encoder 216 and a text encoder 210 can be used, respectively. The video encoder 216 may include one or more video tokenizer models (e.g., one tokenizer in tokenizer 118) trained / updated to generate sequences of tokens representing the video input data 214. The text encoder 210 may include one or more video tokenizer models (e.g., one tokenizer in tokenizer 118) trained / updated to generate sequences of tokens representing the text input data 204. As described herein, the outputs of the multi-channel audio encoder 208, video encoder 216, and text encoder 210 can be combined into a single input context for the multimodal language model 212.
[0096] As described herein, the multimodal language model 212 can be executed (e.g., in an autoregressive manner) to generate one or more sequences of output tokens. In some implementations, if the multimodal language model 212 is to generate audio data, the output sequence generated by the multimodal language model 212 may include special start / stop tokens indicating tokens corresponding to encoded audio data in the output sequence. In some implementations, if the multimodal language model 212 is to generate video data, the output sequence generated by the multimodal language model 212 may include special start / stop tokens indicating tokens corresponding to encoded video data in the output sequence. In some implementations, if the multimodal language model 212 is to generate text data, the output sequence generated by the multimodal language model 212 may include special start / stop tokens indicating tokens corresponding to encoded text data in the output sequence.
[0097] As shown, encoded audio data, video data, and / or text data can be extracted from the output sequence generated by the multimodal language model 212 based on their corresponding start / stop markers, and used to generate audio output data 211, video output data 218, and text output data 213. Generating audio output data 211, video output data 218, and text output data 213 may include providing the encoded audio data, video data, and text data as input to one or more decoders. As described herein, each decoder can perform the inverse operation of its corresponding encoder (e.g., multi-channel audio encoder 208, video encoder 216, text encoder 210). In some implementations, as described herein, one or more decoder models for multi-channel audio can decode the sequence markers for each channel of B-format audio. Similar methods can be used to generate video output data 218 and text output data 213. Audio output data 211, video output data 218 (which may additionally or alternatively include image, sensor and / or other data type outputs), and text output data 213 may be provided as output, for example, as part of a conversational agent application.
[0098] Now, for reference Figure 3 Each block of the method 300 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a standalone application, service, or managed service (standalone or in combination with another managed service), or a plug-in to another product (to name just a few). Additionally, by way of example, regarding... Figure 1 The system described herein describes method 300. However, the method may be performed additionally or alternatively by any system or any combination of systems, including but not limited to the systems described herein.
[0099] Figure 3This is a flowchart of a method for training / updating a spatially aware audio-enhanced dialogue agent according to some embodiments of this disclosure. At block B302, the method 300 includes: identifying multichannel audio data (e.g., multichannel audio data 108). The multichannel audio data may be received from a client device (e.g., client system 101) via a network or retrieved from a data source (e.g., storage device 106). The multichannel audio data may include, but is not limited to, first-order B-format audio or higher-order B-format audio. In some implementations, the multichannel audio data may be identified in response to a request (e.g., from a client device) for training / updating one or more language models (e.g., one or more language models 120). In some implementations, the request may be provided via input to a computing system (e.g., data processing system 102) performing method 300. As described herein, the multichannel audio data may be associated with corresponding spatial information (e.g., spatial information 114) and / or corresponding supplementary media data (e.g., supplementary media data 115).
[0100] At box B304, method 300 includes generating an encoded representation (e.g., encoded audio sample 112) of multi-channel audio data corresponding to a machine learning model (e.g., language model 120). The encoded representation can be generated, for example, by executing one or more tokenizer models (e.g., one or more tokenizers 118) associated with the machine learning model to be trained. As described herein, generating an encoded representation of the multi-channel audio data may include tokenizing the signal from each channel of the multi-channel audio data. In some implementations, tokens for each channel at the same time step may be combined (e.g., concatenated, etc.). In some implementations, separate tokenizer models may be used to generate an encoded representation for each channel of the multi-channel audio data. In some implementations, a single tokenizer model may be used to tokenize all channels of the multi-channel audio data.
[0101] At box B306, method 300 includes: generating a training dataset (e.g., training dataset 110) for a machine learning model using an encoded representation. As described herein, generating the training dataset may include: generating one or more training / update examples based on one or more training / update objectives. For example, a training / update example for enhancing spatial audio quality may include poor-quality encoded audio data as input and quality-enhanced encoded audio data as ground truth output. A training / update example for generating multichannel audio data may include text and / or video data as input to the model and output multichannel audio data as ground truth data. A training / update example for transcribing a moving speaker may include multichannel audio data as input to the model and output text data as ground truth data, the output text data including the transcription of the moving speaker represented in the multichannel audio data. Any number of training / update examples may be generated based on any of the training / update objectives described herein.
[0102] At box B308, method 300 includes: training / updating a machine learning model (e.g., one or more language models 120) using a training dataset to generate an output (e.g., model output 126) corresponding to the input spatial audio. Training / updating the language model described herein may include: iterating over each training / updating example of the training dataset and constructing an input context for the machine learning model. Constructing the input context may include: concatenating an encoded representation of the input data and ground truth data from the training / updating examples into a single data structure. An attention mask may be applied to mask the encoded representation of the ground truth output.
[0103] In training / update iterations, a machine learning model can be executed by passing a sequence of encoded context data to each layer of the model while performing mathematical / machine learning operations on each layer. The output of the machine learning model can include a distribution of candidate label outputs from which one or more output labels are selected. In some implementations, the outputs can be predicted in an autoregressive manner, such that the predicted output labels are appended to the initial context to generate an extended context. The extended context is then provided as input to the machine learning model to predict the next label, until all output labels have been predicted. In some implementations, a "teacher-forced" technique can be used, where ground truth labels (rather than the machine learning model's predictions) from the output portion of the context sequence are appended to the initial input context to predict the next label. In some implementations, the machine learning model can generate labels in a non-autoregressive manner, where the machine learning model is executed to predict all output labels simultaneously.
[0104] As described in this paper, the ground truth labels for training / updating examples can be compared with each output label predicted by the machine learning model using a loss function such as cross-entropy loss to quantify the difference between the predicted and actual labels. Backpropagation can then be used to train / update the parameters of the machine learning model using the computed loss. Backpropagation may involve computing the gradient of the loss with respect to each parameter and adjusting the parameters in the direction that minimizes the loss. Parameter tuning can be performed using a suitable optimization function, such as gradient descent or the Adam optimizer function. The training / updating process can be iteratively repeated using different training / updating examples from the training dataset until a training / updating termination condition has been met.
[0105] The systems and methods described herein can be used for a variety of purposes, such as, but not limited to, machine (e.g., robots, vehicles, construction machinery, warehouse vehicles / machines, autonomous, semi-autonomous and / or other machine types) control, machine motion, machine driving, synthetic data generation, model training (e.g., using real, augmented and / or synthetic data, such as synthetic data generated using simulation platforms or systems, synthetic data generation techniques (e.g., but not limited to the techniques described herein)), perception, augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security and supervision (e.g., in smart city implementations), autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, converter models, etc.) and / or any other suitable application.
[0106] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots or robotic platforms, aerial systems, medical systems, rowing systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulations, in robot simulations, in smart city or supervised simulations, etc.), systems for performing digital twin operations (e.g., in conjunction with a collaborative content creation platform or system, such as, but not limited to, NVIDIA's OMNIVERSE and / or another platform, system, or service using USD or OpenUSD data types), systems implemented using edge devices, systems incorporating one or more virtual machines (VMs), and systems for performing synthetic data generation operations (e.g., using one or more neural rendering fields (NERF), Gaussian splashing, etc.). Systems that are at least partially implemented in a data center (e.g., using splat technology, diffusion models, converter models, etc.), systems for performing conversational AI operations, systems that implement one or more language models (e.g., one or more large language models (LLMs), one or more visual language models (VLMs), one or more multimodal language models, etc.), systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data and / or other data types), systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0107] Example large language model
[0108] In at least some embodiments, language models such as Large Language Models (LLMs), Visual Language Models (VLMs), Multimodal Language Models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD formats such as OpenUSD), and / or the like based on context provided in input prompts or queries. In embodiments, these language models may be considered “large” because they are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—e.g., millions or billions of parameters. LLMs / VLMs / MMLMs / etc. can be implemented for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), extracting insights from data (e.g., text, images, videos, etc.), and generating new text / images / videos / etc. in a user-specified style, tone, and / or format. In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein may be specifically designed for text processing, while in others, a multimodal LLM may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio, 2D and / or 3D data (e.g., USD format) and / or video. For example, a Visual Language Model (VLM) or more specifically a Multimodal Language Model (MMLM) may be implemented to accept images, video, audio, text, 3D designs (e.g., CAD) and / or other input data types and / or generate or output images, video, audio, text, 3D designs and / or other output data types.
[0109] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented using different techniques to understand and generate outputs (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, converter architectures (e.g., architectures relying on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. may also include one or more diffusion blocks (e.g., noise reduction blocks). The LLM / VLM / MMLM / etc. of this disclosure may include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc., including encoder and decoder components (e.g., T5 (Text-to-Text Transformer)), can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting and any architecture type (including, but not limited to, those described herein) can be implemented depending on the specific implementation and the task performed using LLM / VLM / MMLM / etc.
[0110] In various embodiments, unsupervised learning can be used to train LLM / VLM / MMLM / etc., where LLM / VLM / MMLM / etc. learns patterns from a large amount of unlabeled text / audio / video / image / design / USD / etc. data. Due to extensive training, in these embodiments, the model may not require task-specific or domain-specific training. An LLM / VLM / MMLM / etc. extensively pre-trained on a large amount of unlabeled data can be referred to as a base model and can excel at various tasks, such as question answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as cue tuning, fine-tuning, retrieval augmentation generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers to tune or adjust cues or labels to bias the language model towards a specific task or domain), and / or using optimization models for specific tasks and / or other fine-tuning or customization techniques within a specific domain.
[0111] In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using LLM / VLM / MMLM / etc., and / or to prevent the output or presentation of information generated by LLM / VLM / MMLM / etc. (e.g., displays, audio outputs, etc.). In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with the model's inputs and / or outputs. For example, these "protective" models can be trained to identify "safe" or otherwise okay or desired inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a particular application / implementation. Therefore, the LLM / VLM / MMLM / etc. disclosed herein are unlikely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, insecure, out of scope, and / or unwanted for a particular application / implementation.
[0112] In some embodiments, an LLM / VLM / etc. can be configured or able to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations where the model is not ideally suited, the model may have instructions for accessing one or more plugins (e.g., third-party plugins) to help process the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such an example, when at least part of the prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least part of the response requires mathematical computation, the model can access one or more mathematical plugins or APIs to help solve the problem, and then the response from the plugins and / or APIs can be used in the model's output. This process can be repeated (e.g., recursively) an arbitrary number of iterations, using any number of plugins and / or APIs, until a response to each query / question / request / process / action / etc. can be generated in response to the input prompt. Therefore, models can rely not only on their own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (such as APIs, plugins, etc.).
[0113] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple hints provided to the same language model or instances of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, the same input query and hints (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents, for example, providing more than one hint to constrain, guide, or otherwise influence the style, content, or character of the provided output. In one or more exemplary non-limiting embodiments, the same language model can be required to provide output corresponding to different roles, perspectives, characters, or different knowledge bases, as defined by the provided hints.
[0114] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated proxies of at least one language model, and / or provided to two or more prompts for at least one language model can be further processed, such as aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or proxy) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be required to generate or otherwise obtain output about the input source material, wherein the output is associated with the input source material. This association may include, for example, generating captions or text portions embedded (e.g., as metadata) within the input source text or image. In one or more embodiments, the output of the language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, the language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, wherein the text or image is annotated to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in the curatorial dataset, for example, but not limited to this.
[0115] Figure 4A This is a block diagram of an example generative language model system 400 applicable to implementing at least some embodiments of this disclosure. Figure 4A In the example shown, the generative language model system 400 includes a retrieval augmentation generation (RAG) component 492, an input processor 405, a tokenizer 410, an embedding component 420, a plugin / API 495, and a generative language model (LM) 430 (which may include LLM, SLM, VLM, multimodal LM, etc.).
[0116] At a high level, the input processor 405 can receive input 401, which includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, generic scene descriptor (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 430 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, input 401 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 401 may include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., tabular format, JSON, or XML). In generative LM In some implementations of 430 capable of handling multimodal input, input 401 can combine text (or text that may be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking raw input text as an example, input processor 405 can prepare the raw input text in various ways. For example, input processor 405 can perform various types of text filtering to remove noise from relevant text content (e.g., special characters, punctuation marks, HTML tags, stop words, portions of images, portions of audio, etc.). In examples involving stop words (common words that often have little semantic meaning), input processor 405 can remove stop words to reduce noise and allow the generative LM 430 to focus on more meaningful content. Input processor 405 can apply text normalization, for example, by converting all characters to lowercase, removing accent marks, and / or handling special cases (such as abbreviations or abbreviations) to ensure consistency. These are just a few examples; other types of input processing can be applied.
[0117] In some embodiments, RAG component 492 (which may include one or more RAG models, and / or may be performed using generative LM 430 itself) may be used to retrieve additional information to be used as part of input 401 or a prompt. RAGs can be used to enhance input to LLM / VLM / MMLM / etc. with external knowledge to make the answer to a specific question or query or request more relevant, for example, where specific knowledge is required. RAG component 492 may obtain this additional information from one or more external sources (e.g., basic information such as basic text / images / videos / audio / USD / CAD / etc.), and then feed it along with the prompt to LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.
[0118] For example, in some embodiments, in addition to the data retrieved using RAG component 492, input 401 may also be generated using query or model input (e.g., questions, requests, etc.). In some embodiments, input processor 405 may analyze input 401 and communicate with RAG component 492 (or in embodiments, RAG component 492 may be part of input processor 405) to identify relevant text and / or other data to provide to generative LM 430 as additional context or information source, typically from which to identify responses, answers, or outputs 490. For example, when the input indicates that a user is interested in the required tire pressure for a particular brand and model of vehicle, RAG component 492 may use a RAG model, for example, to perform a vector search in the embedding space to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the owner's manual for that particular vehicle brand and model. Similarly, when a user revisits the chatbot related to a specific product sale or service, the RAG component 492 can retrieve previously stored conversation history (or at least its summary) and provide the previous conversation history, along with the current inquiry / request, as part of the generative LM 430 as input 401.
[0119] RAG component 492 can use various RAG techniques. For example, it can use naive RAG ( The document is indexed, chunked, and applied to an embedding model to generate embeddings corresponding to chunks. User queries can also be applied to this embedding model and / or another embedding model of the RAG component 492, and the embeddings of the chunks can be compared with the embeddings of the query to identify the most similar / relevant embeddings to the query. These most similar / relevant embeddings can be provided to the generative LM 430 to generate output.
[0120] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.
[0121] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.
[0122] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which may result in a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide structured entity information to LLM / VLM / MMLM / etc. by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, or it may be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can aggregate the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.
[0123] In any embodiment, RAG component 492 may implement plugins, APIs, user interfaces, and / or other functions to perform RAG. For example, LLM / VLM / MMLM / etc. may use graph RAG plugins to run queries on knowledge graphs to extract relevant information to feed into the model, and may use standard or vector RAG plugins to run queries on vector databases. For example, the graph database may interact with the plugin's REST interface, thus decoupling the graph database from the vector database and / or the embedded model.
[0124] The tokenizer 410 can segment (e.g., processed) text data into smaller units (tags) for subsequent analysis and processing. Depending on the implementation, the tags can represent individual words, subwords, characters, audio / video / images / etc., or partial or multi-channel audio data (e.g., B-format audio with one or more channels, etc.). Word-based tokenization divides the text into individual words, treating each word as a separate tag. Similar methods can be used to generate tags representing one or more samples of one or more channels of multi-channel audio data, as described herein. Subword tokenization breaks down words into smaller, meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 430 to understand morphological changes and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate tag, enabling the generative LM 430 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, the tokenizer 410 can convert (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.
[0125] Embedding component 420 can use any known embedding technique to transform discrete tokens into semantically meaningful (e.g., dense, continuous vector) representations. For example, embedding component 420 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.
[0126] In some implementations where input 401 includes image data / video data, etc., input processor 401 may resize the data to a standard size compatible with the format of the corresponding input channel and / or normalize pixel values to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 420 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where input 401 includes audio data, input processor 401 may resample the audio file to a consistent sampling rate for uniform processing, and embedding component 420 may use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where input 401 includes video data, input processor 501 may extract frames or apply resizing to extracted frames, and embedding component 420 may extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where input 401 includes multimodal data, the embedded component 420 may use techniques such as early fusion (concatenation), late fusion (sequential processing), and attention-based fusion (e.g., self-attention, cross-attention) to fuse representations of different types of data (e.g., text, images, audio, data, video, design, etc.).
[0127] Other components of the generative LM 430 and / or generative LM system 400 may use different types of neural network architectures depending on the implementation scheme. For example, a transducer-based architecture (e.g., the architecture used in models such as GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feedforward network that processes the output of the self-attention layer, applying a nonlinear transformation to the input representation and extracting higher-level features. Some non-limiting example architectures include transducers (e.g., encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning), etc. Therefore, depending on the implementation scheme and architecture, the embedded component 420 can apply the encoded representation of the input 401 to the generative LM 430, and the generative LM 430 can process the encoded representation of the input 401 to generate an output 490, which may include response text and / or other types of data.
[0128] As described herein, in some embodiments, the generative LM 430 may be configured to access or use (or be able to access or use) plugins / APIs 495 (which may include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations where the generative LM 430 is not ideally suited, the model may have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 492) to access one or more plugins / APIs 495 (e.g., third-party plugins) to help process the current input. In such an example, when at least part of the prompt is related to a restaurant or weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs), sending at least part of the prompt related to a particular plugin / API 495 to the plugin / API 495, which may process the information and return an answer to the generative LM 430, which may then use the response to generate output 490. This process can be repeated (e.g., recursively) an arbitrary number of iterations and repeated using any number of plugins / APIs 495 until an output 490 that resolves each query / question / request / process / action / etc. from input 401 is generated. Therefore, the model can rely not only on its own knowledge acquired from training on a large dataset and / or from data retrieved using the RAG component 492, but also on the expertise or optimized properties of one or more external resources (e.g., plugins / APIs 495).
[0129] Figure 4B This is a block diagram of an example implementation scheme, where the generative LM 430 includes a converter encoder-decoder. For example, suppose the input text (e.g., “Who discovered gravity”) is tokenized (e.g., by...) Figure 4A The tokenizer 410) is used for tokens such as words, and each token is encoded (e.g., by...). Figure 4A The embedding component 420 is a corresponding embedding (e.g., of size 512). Since these token embeddings do not typically represent the position of the tokens in the input sequence, positional encoding can be added to each token embedding using any known technique to encode the order relation and context of the tokens in the input sequence. Thus, (e.g., the resulting) embeddings can be applied to one or more encoders 435 of the generative LM 430.
[0130] In the example implementation, encoder 435 forms an encoder stack, where each encoder includes a self-attention layer and a feedforward network. In the example converter architecture, each token (e.g., a word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, pass each vector through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used. For example, to compute a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. The self-attention score for a token pair can be computed by taking the dot product of the query vector and the corresponding key vector, normalizing the resulting score, multiplying by the corresponding value vector, and summing the weighted value vectors. The encoder can apply multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate context vectors that encode the input. Attention projection layer 440 can transform the context vectors into attention vectors (keys and values) for decoder 445.
[0131] In the example implementation, decoder 445 forms a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network. Similar to encoder 435, in the example converter architecture, each token (e.g., a word) flows through a separate path in decoder 445. During the first pass, decoder 445, classifier 450, and generation mechanism 455 can generate a first token, and generation mechanism 455 can apply the generated token as input during a second pass. This process can be repeated cyclically, generating tokens (e.g., words) and adding them to the output of the previous pass, and in subsequent passes applying token embeddings of positionally encoded composite sequences as input to decoder 445, generating one token at a time (called autoregression) until a symbol or token indicating the end of the response is predicted. In each decoder, the self-attention layer is typically restricted to focusing only on preceding positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In the example implementation, the encoder-decoder attention layer operates similarly to the (e.g., multi-head) self-attention operation in encoder 435, except that it creates its query from the layer below it and obtains keys and values (e.g., matrices) from the output of encoder 435.
[0132] Therefore, decoder 445 can output some decoded (e.g., vector) representation of the input applied during a particular pass. Classifier 450 can include a multi-class classifier comprising one or more neural network layers and a softmax operation that transforms logit probabilities into probabilities, the neural network layers projecting the decoded (e.g., vector) representation onto corresponding dimensions (e.g., one dimension for each supported word or token in the output vocabulary). Thus, generation mechanism 455 can select or sample words or tokens based on corresponding predicted probabilities (e.g., selecting the word with the highest predicted probability) and append it to the output of the previous pass, thereby generating each word or token sequentially. Generation mechanism 455 can repeat this process, triggering successive decoder inputs and corresponding predictions until a symbol or token representing the end of the response is selected or sampled, at which point generation mechanism 455 can output the generated response.
[0133] Figure 4C This is a block diagram of an example implementation where the generative LM 430 includes a decoder-only converter architecture. For example, Figure 4C The decoder 460 can be used with Figure 4B The decoder 445 operates similarly, except... Figure 4C Each decoder 460 omits the encoder-decoder self-attention layer (because there is no encoder in this implementation). Therefore, decoders 460 can form a decoder stack, where each decoder includes a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or tag indicating the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., a corresponding embedding with positional encoding) can be applied to decoder 460. Figure 4B Similar to decoder 445, each tag (e.g., a word) can flow through a separate path in decoder 460, and decoder 460, classifier 465, and generation mechanism 470 can use autoregression to generate one tag at a time sequentially until a symbol or tag indicating the end of the response is predicted. Classifier 465 and generation mechanism 470 can be combined with... Figure 4B The classifier 450 and generation mechanism 455 operate similarly, wherein generation mechanism 470 selects or samples each consecutive output label based on the corresponding predicted probability and appends it to the output of the previous iteration, generating each label sequentially until a symbol or label representing the end of the response is selected or sampled. The architectures described herein, and others, are merely examples, and other suitable architectures may be implemented within the scope of this disclosure.
[0134] Example computing device
[0135] Figure 5This is a block diagram of an example computing device 500 suitable for implementing some embodiments of the present disclosure. The computing device 500 may include an interconnect system 502 directly or indirectly coupled to: a memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., one or more displays), and one or more logic units 520. In at least one embodiment, one or more computing devices 500 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 508 may include one or more vGPUs, one or more CPUs 506 may include one or more vCPUs, and / or one or more logic units 520 may include one or more virtual logic units. Accordingly, one or more computing devices 500 may include discrete components (e.g., a full GPU dedicated to computing device 500), virtual components (e.g., a portion of the GPU dedicated to computing device 500), or a combination thereof.
[0136] although Figure 5 The various blocks are shown as being connected to lines via interconnect system 502, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 518 (such as a display device) may be considered I / O component 514 (e.g., if the display is a touchscreen). As another example, CPU 506 and / or GPU 508 may include memory (e.g., memory 504 may represent a storage device in addition to the memory of GPU 508, CPU 506, and / or other components). Therefore, Figure 5 The computing devices described are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop computer," "desktop computer," "tablet computer," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all are conceived in… Figure 5 Within the scope of computing devices.
[0137] Interconnect system 502 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 502 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Fast Peripheral Component Interconnect (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 506 may be directly connected to memory 504. Further, CPU 506 may be directly connected to GPU 508. In cases where there is a direct connection or point-to-point connection between components, interconnect system 502 may include a PCIe link to perform that connection. In these examples, a PCI bus is not required in computing device 500.
[0138] The memory 504 may include any of a variety of computer-readable media. The computer-readable media may be any available medium accessible by the computing device 500. The computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, the computer-readable media may include computer storage media and communication media.
[0139] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 504 may store computer-readable instructions (e.g., representing programs and / or program elements, such as operating systems). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Universal Disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 500. As used herein, computer storage media does not include the signal itself.
[0140] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals (such as carrier waves or other transmission mechanisms) and include any information transmission medium. The term "modulated data signal" can refer to a signal whose characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the foregoing should also be included within the scope of computer-readable media.
[0141] CPU 506 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. Each CPU 506 may include one or more cores (e.g., 1, 2, 4, 8, 28, 72, etc.) capable of processing multiple software threads simultaneously. CPU 506 may include any type of processor and may include different types of processors depending on the type of computing device 500 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 500, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplemental coprocessors such as math coprocessors, computing device 500 may also include one or more CPUs 506.
[0142] In addition to or replacing CPU 506, one or more GPUs 508 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. One or more GPUs 508 may be integrated GPUs (e.g., having one or more CPUs 506) and / or one or more GPUs 508 may be discrete GPUs. In embodiments, one or more GPUs 508 may be a coprocessor of one or more CPUs 506. GPUs 508 may be used by computing device 500 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPUs 508 may be used for general-purpose computing on a GPU (GPGPU). GPUs 508 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPUs 508 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received from CPU 506 via a host interface). GPU 508 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 504. GPU 508 may include two or more GPUs operating in parallel (e.g., via a link). The link may be directly connected to the GPUs (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 508 may generate pixel data or GPGPU data for different portions of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0143] In addition to or replacing CPU 506 and / or GPU 508, one or more logic units 520 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPUs 506, one or more GPUs 508, and / or one or more logic units 520 may perform any combination of methods, processes, and / or portions thereof, discretely or jointly. One or more logic units 520 may be part of and / or integrated into one or more of CPUs 506 and / or GPUs 508, and / or one or more logic units 520 may be discrete components or otherwise external to CPUs 506 and / or GPUs 508. In embodiments, one or more logic units 520 may be coprocessors of one or more CPUs 506 and / or one or more GPUs 508.
[0144] Examples of logic unit 520 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.
[0145] Communication interface 510 may include one or more receivers, transmitters, and / or transceivers that allow computing device 500 to communicate with other computing devices via electronic communication networks (including wired and / or wireless communications). Communication interface 510 may include components and functions that allow communication over any of a plurality of different networks (e.g., wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or wireless bandwidth), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet). In one or more embodiments, logic unit 520 and / or communication interface 510 may include one or more data processing units (DPUs) for directly transmitting data received via a network and / or via interconnect system 502 to one or more GPUs 508 (e.g., the memory of one or more GPUs 508).
[0146] I / O port 512 allows computing device 500 to be logically coupled to other devices including I / O component 514, presentation component 518, and / or other components, some of which may be built into (e.g., integrated into) computing device 500. Illustrative I / O component 514 includes a microphone, mouse, keyboard, joystick, gamepad, game controller, disc satellite dish, scanner, printer, wireless device, etc. I / O component 514 provides a natural user interface (NUI) that handles air gestures, voice, or other physiological input generated by the user. In some instances, the input may be passed to appropriate network elements for further processing. NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 500 (as described in more detail below). The computing device 500 may include a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, touchscreen technology, and combinations thereof, for attitude detection and recognition. Additionally, the computing device 500 may include an accelerometer or gyroscope that allows motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the computing device 500 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0147] Power supply 516 may include a hardwired power supply, a battery power supply, or a combination thereof. Power supply 516 may provide power to computing device 500 so that the components of computing device 500 can operate.
[0148] One or more presentation components 518 may include displays (e.g., monitors, touchscreens, television screens, head-up displays (HUDs), other display types, or combinations thereof), speakers, and / or other presentation components. Presentation component 518 may receive data from other components (e.g., GPU 508, CPU 506, DPU, etc.) and output data (e.g., as images, videos, sounds, etc.).
[0149] Example Data Center
[0150] Figure 6 An example data center 600 that can be used in at least one embodiment of this disclosure is shown. The data center 600 may include a data center infrastructure layer 610, a framework layer 620, a software layer 630, and / or an application layer 640.
[0151] like Figure 6As shown, the data center infrastructure layer 610 may include a resource coordinator 612, grouped computing resources 614, and node computing resources (“nodes CR”) 616(1)-616(N), where “N” represents any integer, a positive integer. In at least one embodiment, the nodes CR 616(1)-616(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW”) devices, and network network interfaces (“NW”). I / O devices, network switches, virtual machines ("VMs"), power modules and / or cooling modules, etc. In some embodiments, one or more nodes CR616(1)-616(N) may correspond to a server having one or more of the aforementioned computing resources. Furthermore, in some embodiments, nodes CR616(1)-616(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more nodes CR616(1)-616(N) may correspond to virtual machines (VMs).
[0152] In at least one embodiment, the grouped computing resources 614 may include separate groups of nodes CR616 housed within one or more racks (not shown) or within a plurality of racks in data centers (also not shown) located in different geographical locations. The separate groups of nodes CR616 within the grouped computing resources 614 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, a plurality of nodes CR616, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0153] Resource coordinator 612 may be configured or otherwise controlled to control one or more nodes CR616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 612 may include a Software Design Infrastructure (“SDI”) management entity for data center 600. Resource coordinator 612 may include hardware, software, or some combination thereof.
[0154] In at least one embodiment, such as Figure 6As shown, framework layer 620 may include a job scheduler 628, a configuration manager 634, a resource manager 636, and / or a distributed file system 638. Framework layer 620 may include a framework of software 632 supporting software layer 630 and / or one or more applications 642 of application layer 640. Software 632 or application 642 may respectively include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 620 may be, but is not limited to, a type of free and open-source software web application framework that can utilize the distributed file system 638 for large-scale data processing (e.g., "big data"), such as Apache Spark. TM (Hereinafter referred to as "Spark"). In at least one embodiment, job scheduler 628 may include Spark drivers to facilitate the scheduling of workloads supported by various layers of data center 600. Configuration manager 634 may be able to configure different layers, such as software layer 630 and framework layer 620 including Spark and distributed file system 638 for supporting large-scale data processing. Resource manager 636 may be able to manage cluster or group computing resources mapped to or allocated for supporting distributed file system 638 and job scheduler 628. In at least one embodiment, cluster or group computing resources may include group computing resources 614 at data center infrastructure layer 610. Resource manager 636 may coordinate with resource coordinator 612 to manage these mapped or allocated computing resources.
[0155] In at least one embodiment, the software 632 included in software layer 630 may include software used by at least a plurality of portions of nodes CR616(1)-616(N), grouped computing resources 614, and / or the distributed file system 638 of framework layer 620. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0156] In at least one embodiment, the application 642 included in the application layer 640 may include one or more types of applications used by at least a plurality of portions of nodes CR616(1)-616(N), grouped computing resources 614, and / or the distributed file system 638 of the framework layer 620. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications (including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.) and / or other machine learning applications used in combination with one or more embodiments).
[0157] In at least one embodiment, any of the configuration manager 634, resource manager 636, and resource coordinator 612 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can free the data center operator of data center 600 from making potentially undesirable configuration decisions and potentially avoid underutilized and / or poorly performing portions of the data center.
[0158] According to one or more embodiments described herein, data center 600 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 600. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 600 by using weight parameters calculated through one or more training techniques (such as, but not limited to, those described herein).
[0159] In at least one embodiment, the data center 600 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured to allow users to train or execute information inference services, such as image recognition, speech recognition, or other artificial intelligence services.
[0160] Example network environment
[0161] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network attached storage (NAS), other back-end devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 5 The implementation is carried out on one or more instances of computing device 500, for example, each device may include similar components, features and / or functions of computing device 500. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 600, examples of which are provided in this document. Figure 6 To describe in more detail.
[0162] Components of a network environment can communicate with each other via one or more networks, which may be wired, wireless, or both. A network can include multiple networks or a network of networks. For example, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0163] A compatible network environment may include one or more peer-to-peer network environments—in which case the network environment may not include a server—and one or more client-server network environments—in which case the network environment may include one or more servers. In a peer-to-peer network environment, the functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0164] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, job scheduler, resource manager, and distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for software supporting the software layer and / or one or more applications supporting the application layer. The software or application may respectively include web-based service software or applications. In embodiments, one or more client devices may use web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, free and open-source software web application frameworks, such as those used for large-scale data processing (e.g., "big data") using distributed file systems.
[0165] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions may be distributed across multiple locations from a central or core server (e.g., across one or more data centers distributed across states, regions, countries, globally, etc.). If the connection to the user (e.g., client device) is relatively close to the edge server, the core server may assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0166] One or more client devices may include the information described in this article. Figure 5 At least some of the components, features, and functions of one or more example computing devices 500 described. By way of example and not limitation, a client device may be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0167] This disclosure can be described in the general context of computer code or machine-usable instructions (including computer-executable instructions, such as program modules) that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This disclosure can also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked via a communication network.
[0168] As used herein, the phrase “and / or” relating to two or more elements should be interpreted as meaning only one element, or a combination of elements. For example, “element A, element B, and / or element C” can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0169] This document provides a detailed description of the subject matter of this disclosure to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have anticipated that the claimed subject matter may also be embodied in other ways in combination with other current or future techniques to include combinations of different steps or steps similar to those described in this document. Furthermore, although the terms “step” and / or “box” may be used herein to refer to different elements of the method employed, such terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and only if the order of individual steps is explicitly described.
Claims
1. One or more processors, comprising: One or more circuits are used for: Generate an encoded representation of multi-channel audio data corresponding to a machine learning model; The machine learning model is generated using the encoded representation, the training dataset indicating spatial information of at least one audio source represented in the multi-channel audio data; as well as The training dataset is used to update one or more parameters of the machine learning model to generate an output corresponding to the input spatial audio.
2. The processor of claim 1 or more, wherein the machine learning model comprises at least one of the following: a large language model (LLM), a visual language model (VLM), or a multimodal language model (MMLM).
3. The processor of claim 1, wherein the spatial information comprises text data, and wherein the one or more circuits are configured to update the one or more parameters of the machine learning model to generate output text data relating to at least one audio source represented in the input spatial audio.
4. The processor of claim 3 or more, wherein the output text data identifies one or more of the following: the distance from the audio source represented in the input spatial audio, the number of audio sources represented in the input spatial audio, or the transcription or logger output of speech from a moving audio source represented in the input spatial audio.
5. The processor of claim 1, wherein the one or more circuits are configured to generate the multi-channel audio data by applying spatial transformation operations to a plurality of audio sources.
6. The processor of claim 5 or more, wherein the spatial transformation operation generates the multi-channel audio data as B-format audio.
7. The processor of claim 1, wherein the one or more circuits are configured to update one or more parameters of the machine learning model to generate output spatial audio based on the input spatial audio.
8. One or more processors according to claim 1, wherein the one or more circuits are used for: The training dataset is generated to include an encoded representation of the video data; and The training dataset is used to update one or more parameters of the machine learning model to generate output spatial audio that tracks at least one audio source depicted in the video data.
9. The processor of claim 8, wherein the one or more circuits are configured to update the one or more parameters of the machine learning model to receive the encoded representation of the single-channel audio data and the video data, thereby generating the output spatial audio.
10. The processor of claim 1 or more, wherein the processor is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for executing optical transmission models; A system for performing collaborative content creation for 3D assets; A system used to perform deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing conversational AI operations; A system for performing generative AI operations using language models; Systems for performing generative AI operations using large language models (LLMs); A system for performing generative AI operations using a visual language model (VLM); A system for performing generative AI operations using multimodal language models; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
11. A system comprising: One or more processors are used for: The system receives input audio from a language model trained to process multi-channel audio data from a client device. The input audio and the language model are used to generate output data, the output data indicating spatial information of at least one audio source represented in the input audio; as well as The output data indicating the spatial information is provided to the client device.
12. The system of claim 11, wherein the one or more processors are configured to: The encoded representation of the input data used to generate the language model; and The encoded representation is provided as input to the language model.
13. The system of claim 11, wherein the one or more processors are configured to: Receive the input text of the language model; and The language model is used to generate output data indicating the spatial information based on the input text and the input audio.
14. The system of claim 11, wherein the one or more processors are configured to: Receive the input video of the language model; and The language model is used to generate output data indicating the spatial information based on the input video and the input audio.
15. The system of claim 11, wherein the output data comprises the encoded output of the language model, and the one or more processors are configured to: Multichannel audio is generated based on the encoded output of the language model.
16. The system of claim 11, wherein the output data includes one or more of the following: the number of sound sources represented in the input audio, the estimated distance of the sound sources represented in the input audio, or the estimated location of the sound sources represented in the input audio.
17. The system of claim 11, wherein the system is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system used to perform deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing conversational AI operations; A system for performing generative AI operations using language models; Systems for performing generative AI operations using large language models (LLMs); A system for performing generative AI operations using a visual language model (VLM); A system for performing generative AI operations using multimodal language models; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
18. A method comprising: Use one or more processors to generate an encoded representation of multi-channel audio data corresponding to a machine learning model; The training dataset for the machine learning model is generated using one or more processors, the training dataset indicating spatial information of at least one audio source represented in the multi-channel audio data; as well as The machine learning model is updated using the one or more processors and the training dataset to generate an output corresponding to the input spatial audio.
19. The method of claim 18, wherein the spatial information comprises text data, and wherein the method further comprises: The one or more processors are used to update the one or more parameters of the machine learning model to generate output text data relating to at least one audio source represented in the input spatial audio.
20. The method of claim 19, wherein the output text data identifies one or more of the following: the distance from the audio source represented in the input spatial audio, the number of audio sources represented in the input spatial audio, or the transcription of speech from a moving audio source represented in the input spatial audio.