Multilingual speech recognition models for speech processing systems and applications

CN122658293APending Publication Date: 2026-08-28NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610073672.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-01-20
Publication Date
2026-08-28

AI Technical Summary

Benefits of technology

[0003]In contrast to conventional systems, in some embodiments, the system of this disclosure uses an end-to-end machine learning model capable of performing both ASR processing and translation processing to generate text presented in various languages. More specifically, compared to conventional systems that use a specially trained ASR model to translate speech in one direction, the system of this disclosure is capable of using an end-to-end machine learning model to translate speech in multiple directions (e.g., translating into various languages). Additionally, compared to conventional systems that use different models to perform ASR processing and language translation, the system of this disclosure is capable of using a single machine learning model to perform both speech processing and language translation. In either instance, the system of this disclosure can reduce the amount of computational resources (e.g., models), storage (e.g., storing only one model, rather than storing a separate model for each language-to-language translation), latency, and/or training associated with speech processing. Furthermore, by training a single model for any number of languages, the model can learn from the nuances or styles of similar languages, allowing a single model to outperform a separate model trained specifically for one-to-one language translation in translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658293A_ABST
    Figure CN122658293A_ABST
Patent Text Reader

Abstract

The present disclosure relates to multilingual speech recognition models for speech processing systems and applications. In various examples, described herein are multilingual speech processing models for speech processing systems and applications. The systems and methods described herein can use an end-to-end model that is capable of performing both ASR processing and translation processing to generate text presented in various languages. For instance, a user can provide at least audio data representing speech and an indication of a target language for translating the speech. The model can then generate one or more audio representations associated with the speech and one or more language representations associated with the target language. Additionally, the model can combine the audio representations with the language representations (such as by performing stitching, masking, fusing, adding, etc.) to generate one or more combined representations. The model can then process the combined representations to generate text that corresponds to the speech and is presented in the target language.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Automatic speech recognition (ASR) models are used to process speech from users in order to convert speech into text. In some instances, ASR models are also capable of performing language translation in one direction (such as from the initial language of the speech to another language for which the ASR model was trained), or additional machine learning models can be used to translate text from the ASR model into the target language desired by the user. For example, if a user wants speech translated from English to French, the system can use an ASR model specifically trained to perform the processing of converting speech presented in English into corresponding text presented in French. Alternatively, the system can use an ASR model that first converts the speech into text presented in English, and then uses an additional model to translate the text into French. Thus, traditional systems performing ASR processing either require a specially trained model to translate in one direction or multiple models to translate different languages ​​(e.g., a separate model for each language-to-language conversion). Summary of the Invention

[0002] The various embodiments of this disclosure relate to multilingual speech processing models for speech processing systems and applications. The systems and methods described herein can utilize end-to-end machine learning models capable of performing both ASR processing and translation processing to generate text presented in various languages. For example, a user or system (e.g., automatically, based on known parameters of a task, operation, or system) can provide at least audio data representing speech and an indication of the target language for speech translation. The machine learning model can then generate one or more audio representations (e.g., one or more vectors, embeddings, tensors, kernels, features, etc.) associated with the speech and one or more language representations (e.g., one or more vectors, embeddings, tensors, kernels, features, etc.) associated with the target language. Additionally, the machine learning model can combine one or more audio representations with one or more language representations (e.g., by performing concatenation, masking, fusion, addition, etc.) to generate one or more combined representations. The machine learning model can then process one or more combined representations (e.g., by using one or more decoders) to generate text corresponding to the speech and presented in the target language.

[0003] In contrast to conventional systems, in some embodiments, the system of this disclosure uses an end-to-end machine learning model capable of performing both ASR processing and translation processing to generate text presented in various languages. More specifically, compared to conventional systems that use a specially trained ASR model to translate speech in one direction, the system of this disclosure is capable of using an end-to-end machine learning model to translate speech in multiple directions (e.g., translating into various languages). Additionally, compared to conventional systems that use different models to perform ASR processing and language translation, the system of this disclosure is capable of using a single machine learning model to perform both speech processing and language translation. In either instance, the system of this disclosure can reduce the amount of computational resources (e.g., models), storage (e.g., storing only one model, rather than storing a separate model for each language-to-language translation), latency, and / or training associated with speech processing. Furthermore, by training a single model for any number of languages, the model can learn from the nuances or styles of similar languages, allowing a single model to outperform a separate model trained specifically for one-to-one language translation in translation. Attached Figure Description

[0004] The following describes in detail the system and method for multilingual speech recognition models used in speech processing systems and applications, with reference to the accompanying drawings, wherein: Figure 1 The illustration shows an example data flow diagram of a process for translating speech into different text languages ​​using one or more multilingual machine learning models, according to some embodiments of the present disclosure. Figure 2 Examples of generating audio representations and speech representations according to some embodiments of the present disclosure are illustrated; Figure 3 The illustration depicts a first technique for combining audio representation and language representation using splicing or other vector / tensor combination techniques according to some embodiments of the present disclosure; Figure 4 The illustration depicts a second technique using a combination of audio and speech representations with sinusoidal transforms according to some embodiments of the present disclosure; Figure 5 The illustration shows examples of performing ASR processing and translation processing using one or more machine learning models according to some embodiments of the present disclosure; Figure 6 The illustration shows an example of a process for training one or more machine learning models to perform ASR processing and translation processing according to some embodiments of the present disclosure; Figure 7The illustration shows examples of one or more systems that can perform one or more of the processes described herein to translate speech, according to some embodiments of the present disclosure; Figure 8 The illustration shows a flowchart of a method for translating speech into text using a multilingual model, according to some embodiments of the present disclosure; Figure 9 The illustration shows a flowchart of a method for translating speech into text presented in a target language, according to some embodiments of the present disclosure; Figure 10A This is a block diagram of an example generative language model system applicable to implementing at least some embodiments of the present disclosure; Figure 10B It is a block diagram of an example generative language model including a transformer encoder-decoder suitable for implementing at least some embodiments of the present disclosure; Figure 10C It is a block diagram of an example generative language model that includes a decoder-only transformer architecture suitable for implementing at least some embodiments of the present disclosure; Figure 11 This is a block diagram of an exemplary computing device applicable to implementing some embodiments of the present disclosure; and Figure 12 This is a block diagram of an example data center applicable to implementing some embodiments of this disclosure. Detailed Implementation

[0005] Systems and methods for multilingual speech recognition models for speech processing systems and applications are disclosed. For example, one or more systems may receive audio data representing speech and input data indicating a target language for generating text corresponding to the speech. In some examples, the target language may include the same language as the speech, while in other examples, the target language may include a different language compared to the speech. Additionally, as described herein, the language may include, but is not limited to, English, Spanish, French, Arabic, German, Chinese, Japanese, Dutch, and / or any other language. One or more systems may then use an end-to-end machine learning model (“model”) configured to perform both ASR processing and translation processing to process the audio data and generate text corresponding to the speech and presented in the target language.

[0006] For example, the model can process audio data (such as by using one or more encoders (and / or any other type of processing component)) to generate one or more audio representations associated with audio information. As described herein, audio representations can include, but are not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation of the audio data. The model can also process input data (such as by using one or more encoders (and / or any other type of processing component)) to generate one or more language representations associated with target language information. As described herein, language representations can include, but are not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation. In some examples, one or more language representations can represent temporal quantities similar to one or more audio representations. For example, one or more language representations can represent time periods associated with audio data.

[0007] As described herein, in some examples, one or more language representations (e.g., vectors) can be generated to include dimensions associated with several languages ​​the model is capable of translating. For example, if the model is configured to translate 128 languages, one or more language representations may include 128 dimensions. Additionally, in some examples, one-hot encoding can be used for language recognition, such that one or more elements of one or more language representations associated with the target language may include a first value (e.g., 1), while one or more other elements of one or more language representations associated with one or more other languages ​​include a second value (e.g., 0). The model can then generate one or more combined representations by combining one or more audio representations with one or more language representations, such as by using one or more projection layers (and / or any other type of layer). As described herein, combined representations may include, but are not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation.

[0008] The model can combine one or more audio representations with one or more language representations using one or more techniques. For example, in some examples, the model can concatenate one or more language representations with one or more audio representations to generate one or more concatenated representations. In such examples, the model can then further process one or more concatenated representations to reduce the dimensionality back to one or more original audio representations while integrating language information with audio information. Additionally or alternatively, in some examples, the model can generate one or more additional language representations using one or more language representations and encoding a sinusoidal kernel (e.g., a fixed kernel matrix) of different frequencies associated with the target language. For example, one or more additional language representations can be generated by multiplying one or more language representations by a sinusoidal kernel to project language information to the same dimension as one or more audio representations. The model can then add one or more additional language representations to one or more audio representations (e.g., by masking one or more audio representations with one or more additional language representations). While these are just two example techniques for combining representations, in other examples, any other techniques can be used to combine representations.

[0009] The model can then process one or more combined representations (such as by using one or more decoders (and / or any other type of processing component)) to generate output data associated with the speech and presented in the target language. In some examples, the output data may represent tokens corresponding to the text, where the model and / or one or more additional processing components then process the tokens to generate the speech-associated text. However, in some examples, the output data may represent the actual text corresponding to the speech, such as a transcript of the speech. In any of these examples, one or more systems can then use the text to perform one or more tasks, such as providing the text to a user, further processing the text using additional processing components, and / or any other task.

[0010] While these examples describe processing audio data to generate text presented in a single target language, in other examples, similar processes can be used to process one or more instances of audio data to generate text presented in many target languages. For instance, for another target language, the model can generate one or more additional language representations associated with other target languages. The model can then process the audio data using one or more processes described herein, such as by using one or more additional language representations and one or more audio representations associated with the audio data, to generate text corresponding to the speech from the audio data and presented in another language. In other words, one or more systems can use an end-to-end model to process audio data representing speech presented in different languages ​​to generate translated text presented in multiple languages.

[0011] In some examples, one or more systems (and / or one or more other systems) can train a model to perform one or more procedures described herein. For example, one or more systems can train a model using training input data, such as audio data representing speech and a list of target languages ​​associated with the speech, and ground truth data representing text associated with the speech and presented in the target language. One or more systems can then process the training input data using the procedures described herein to generate output data representing the text. Additionally, one or more systems can determine one or more losses based at least on comparing the output text with the ground truth text, and update the model based at least on one or more losses. For example, one or more systems can update one or more weights and / or parameters of one or more encoders, one or more projection layers, one or more decoders, etc., based at least on one or more losses.

[0012] In some examples, one or more models described herein (e.g., machine learning models, deep neural networks, language models, LLM, VLM, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, neural networks, etc.) can be packaged as microservices, such as inference microservices (e.g., NVIDIA NIM), which may include containers (e.g., operating system (OS) level virtualization packages) that may include application programming interface (API) layers, server layers, runtime layers, and / or model "engines." For example, an inference microservice may include the container itself and one or more models (e.g., weights and biases). In some instances (e.g., where one or more machine learning models are small enough (e.g., have a sufficiently few parameters)), one or more models may be included within the container itself. In other instances (e.g., where one or more models are large), one or more models may be hosted / stored in the cloud (e.g., in a data center) and / or may be hosted on the ground and / or at the edge (e.g., on a local server or compute device, but outside the container). In these embodiments, one or more models can be accessed via one or more APIs, such as REST APIs. Thus, and in some embodiments, one or more machine learning models described herein can be deployed as inference microservices to accelerate the deployment of one or more models on any cloud, data center, or edge computing system, while ensuring data security.

[0013] For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployments with implementation software such as NVIDIA's Triton Inference Server), and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations for delivering low latency and high throughput for production applications (such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). One or more machine learning models described herein may be part of a microservice along with accelerated infrastructure capable of deployment using a single command and / or orchestrated and auto-scaled using a container orchestration system on accelerated infrastructure (e.g., on a single device up to data center scale). Thus, an inference microservice may include one or more machine learning models (e.g., already optimized for high-performance inference), inference runtime software for executing one or more machine learning models and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity, and / or other monitoring. In some embodiments, the inference microserver may include software for performing in-situ replacement and / or updates to one or more machine learning models. When a replacement or update is performed, the software performing the replacement / update may maintain user configurations for the inference runtime software and enterprise management software.

[0014] The systems and methods described herein can be, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vessels, ships, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, by way of example but not limited to: machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable application.

[0015] The disclosed embodiments can comprise a variety of different systems, such as motor vehicle systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing large language models (LLMs), systems implementing one or more visual language models (VLMs), systems implementing one or more multimodal language models, systems using or deploying one or more inference microservices, systems containing one or more machine learning models deployed in services or microservices and OS-level virtualization packages (e.g., containers), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0016] In some embodiments, the systems and methods described herein can be deployed in call or smart kiosk applications. For example, a kiosk, tablet, smart display, or other device may include one or more onboard processors (e.g., CPU, GPU, deep learning accelerator, SoC) and memory and / or storage devices (e.g., for storing models, image databases, etc.). In some embodiments, the kiosk / tablet / display may (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) communicate with one or more locally hosted servers / computing devices and / or with one or more remotely located (e.g., in one or more data centers) servers / computing devices. In these examples, the kiosk may use one or more APIs (such as, but not limited to, REST APIs) to communicate with one or more machine learning models (e.g., language models, LLM, VLM, MMLM, diffusion models, transformer models, NeRF, DNN, etc.) and / or image databases hosted on local and / or remote servers. For example, a call kiosk may deploy one or more machine learning models described herein to perform translations for the kiosk's users.

[0017] In one or more embodiments, the systems and methods described herein can be deployed in gaming applications. For example, game consoles, PCs, tablets, or other gaming devices may include one or more onboard and / or remote processors (e.g., CPUs, GPUs, deep learning accelerators, SoCs) and memory and / or storage devices (e.g., for storing game models, game assets, player data, etc.). These devices can use one or more machine learning models (e.g., diffusion models, transformer models, neural rendering field (NeRF) models, language models (e.g., LLM, VLM, MMLM, etc.), DNNs, etc.) to enhance gameplay, generate real-time dynamic content, and personalize the user experience based on in-game behavior or pre-stored player profiles. In some embodiments, the system can be deployed in a cloud gaming environment (e.g., NVIDIA's GeForce Now). In this case, client devices (e.g., smart displays, tablets, or game controllers) can be used to interact with the game, while one or more machine learning models and / or visual rendering may occur on one or more remote servers / computing devices (e.g., in one or more data centers). The language models, AI processing, and rendering described herein can operate in the cloud to process player input received from one or more end-user devices (e.g., based on controller, keyboard, mouse, joystick, AR / VR / MR / etc. input), generate appropriate in-game responses, render content, and send or transmit content to one or more end-user devices. During the reception and / or transmission of data with one or more end-users or edge devices, one or more data processing units (DPUs) and / or network interface cards (NICs) can be used. For example, one or more machine learning models described herein can be used to perform translations into any number of other languages ​​for users who speak any number of languages. In this way, game and / or other users can understand languages ​​and communicate across languages.

[0018] In some embodiments, the systems and methods described herein can be deployed in video conferencing applications. For example, video conferencing devices (such as dedicated conferencing units, computers, tablets, and / or smartphones) may include one or more onboard processors (e.g., central processing units, graphics processing units, deep learning accelerators, SoCs) and memory and / or storage devices (e.g., for storing video, audio, or other communication-related data). The system can use one or more machine learning models (e.g., diffusion models, transformer models, neural rendering field (NeRF) models, language models (e.g., LLM, VLM, MMLM, etc.)) to enhance video conferencing functionality, including real-time or near real-time transcription, speaker segmentation, language translation, automatic speech recognition (ASR), and / or background noise cancellation. In one or more embodiments, the system can enable users to interact with the video conferencing platform using natural language input. For example, users can issue voice commands to schedule, join, or leave a meeting or manage participants and screen sharing. During data reception and / or transmission with one or more end-users or edge devices, one or more data processing units (DPUs) and / or network interface cards (NICs) may be used. For example, one or more machine learning models described herein can be used to perform translations for users who speak any number of languages ​​into any number of other languages. In this way, video conferencing applications and / or other users can understand languages ​​and communicate across languages.

[0019] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA) (which may include one or more vector processing units (VPU)), direct memory access (DMA) systems and / or pixel processing engines (PPE)), hardware-based optical flow accelerators (OFA), SoCs, etc.), and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models, visual language models (VLM), large language models (LLM), visual language action (VLA) models, multimodal language models (MMLM), etc.), which enable the robotic system to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects, or navigating the environment using sensors (such as cameras, LiDAR, RADAR, ultrasonic sensors, etc.). This system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, RADAR, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server for computationally intensive tasks such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze optimized commands and distribute them across the entire fleet. In some embodiments, one or more machine learning models described herein (e.g., language models, VLM, VLA, LLM, MMLM, diffusion models, NeRF models, DNNs, etc.) can be used to enable the robot to perceive and reason about its environment and / or communicate with one or more other robots and / or humans in the environment. In some embodiments, the robot can communicate (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) with one or more locally hosted servers / computing devices and / or with one or more remotely located servers / computing devices (e.g., in one or more data centers). For example, one or more machine learning models described herein can be used to perform translations for the robot to communicate with humans / other robots in environments spanning any number of languages ​​to any number of other languages. In this way, the robot is able to understand languages ​​and communicate across languages.

[0020] In some embodiments, the systems and methods described herein can be deployed in in-vehicle infotainment (IVI) systems or in-cabin experience (IX) applications. For example, an infotainment system within a vehicle (e.g., a car, truck, drone, engineering equipment, robot, semi-autonomous vehicle, or autonomous vehicle) may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA) (which may include one or more vector processing units (VPUs)), direct memory access (DMA) systems and / or pixel processing engines (PPEs)), hardware-based optical flow accelerators (OFA), SoCs, etc.), memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models), and memory and / or storage devices (e.g., for storing entertainment content, navigation data, and user preferences). The system can use these processors to execute one or more machine learning models (e.g., language models) to achieve features such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services via network connectivity. In-vehicle infotainment systems can also utilize natural language processing (NLP) models to enable voice-based interaction. One or more machine learning models can be stored locally or accessed via one or more APIs connected to cloud services, allowing the system to process requests in real-time or near real-time. For example, one or more machine learning models described herein can be used to perform translations for users within the vehicle into any number of other languages. In this way, users and the vehicle can understand language and communicate across languages.

[0021] While examples using machine learning models such as neural networks may be described in this document, this is not intended to be limiting. For example, but not limited to, any of the various machine learning models and / or neural networks described herein can include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN), perceptrons, long short-term memory (LSTM) networks, multilayer perceptrons (MLP) networks. Networks, Deep Stacked Networks (DSN), Generative Pre-trained (GPT) models or networks, Feedforward networks, Radial Basis Function ANNs, Self-Organizing Maps (SOM), Kohonen Maps, Hopfield Networks, Boltzmann Machines, Deep Belief Neural Networks, Deconvolutional Neural Networks, Generative Adversarial Networks (GANs), Liquid Machines, Modular Neural Networks, Liquid Machines, Sequence-to-Sequence Models, Networks Using Transformer Architectures, State-Space Models (SSMs) (e.g., using Mamba architectures (e.g., Mamba-1, Mamba...) Networks such as 2, networks using selective state-space models, networks using structured state-space sequence models, diffusion models (e.g., diffusion probability models, fraction-based generative models, etc.), neural radiation field (NeRF) models, Gaussian blob models, Kolmogorov-Arnold networks (KAN), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLM), visual language models (VLM), multimodal language models (MMLM), large action models (LAM), visual language action (VLA) models, etc., and one or more machine learning models and / or other types of machine learning models.

[0022] In some embodiments, one or more Transformer Engines (TEs) can be implemented. Transformer Engines can use micro-tensor scaling to optimize performance and accuracy, such as implementing 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FPC4) AI processing. For example, a Transformer Engine can combine 16-bit or 8-bit floating-point precision and 8-bit or 4-bit floating-point data formats with software algorithms to improve AI performance and capabilities. By reducing mathematical operations to 8 bits or 4 bits, TEs allow for faster training of larger networks without compromising accuracy. For example, TEs can include libraries for accelerating Transformer Models on processing devices such as GPUs to provide better performance with lower memory utilization during both training and inference. When TE is combined with other technologies, such as high-speed interconnects between nodes (e.g., using switches such as NVLink switches) and tensor kernels (which enable mixed-precision computation, such as microscale precision support), server clusters can train large networks (e.g., billions of parameters) at higher speeds. This allows support for tensor kernel precision of FP64, TF32, BF16, FP16, FP8, INT8, FP6, and FP4, as well as CUDA kernel precision of FP64, FP32, FP16, and BF16.

[0023] refer to Figure 1 , Figure 1 The illustration shows an example data flow diagram of a process 100 for translating speech into different text languages ​​using one or more multilingual machine learning models according to some embodiments of the present disclosure. It should be understood that the arrangements and other arrangements described herein are illustrative only. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or as alternatives to the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, a processor executing instructions stored in memory can perform various functions. In some embodiments, the systems, methods, and processes described herein can use... Figure 11 Example computing device 1100 and / or Figure 12 Example data center 1200 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the operation.

[0024] The process 100 may include receiving audio data 102, representing at least speech, from a user, and language data 104, representing the target language used to generate text corresponding to the speech. In some examples, the target language may include the same language as the speech, while in other examples, the target language may include a different language compared to the speech. Additionally, as described herein, the language data 104 may include input data indicating the target language, selection data indicating the choice of the target language, historical data indicating the historical target language used, and / or any other type of data and / or associated with the input data indicating the target language, the selection data indicating the choice of the target language, the historical data indicating the historical target language used, and / or any other type of data. Furthermore, the language may include, but is not limited to, English, Spanish, French, Arabic, German, Chinese, Japanese, Dutch, and / or any other type of language.

[0025] The process 100 may then include: processing the audio data 102 using one or more encoders 106 of one or more machine learning models; and generating one or more audio representations 108 corresponding to the audio data 102, at least based on the processing. As described herein, the encoder 106 may include any type of encoder configured to perform one or more of the processes described herein, such as an audio encoder, a FastConformer encoder, a one-hot encoder, a variational autoencoder, a recurrent neural network encoder, etc. Additionally, the audio representation 108 may include, but is not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation of the audio data. For example, one or more audio representations 108 may include one or more vectors that capture features of one or more portions of the speech represented by the audio data, such as in a multidimensional embedding space.

[0026] The process 100 may further include: processing audio data 104 using one or more encoders 110 of one or more machine learning models; and generating one or more language representations 112 corresponding to the target language indicated by the language data 104, at least based on the processing. As described herein, encoders 110 may include any type of encoder configured to perform one or more of the processes described herein, such as one-hot encoders, ordinal encoders, binary encoders, tag encoders, frequency encoders, etc. Additionally, audio representations 108 may include, but are not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation of the audio data. For example, in some examples, one or more language representations 112 may include one or more vectors, wherein each vector includes dimensions corresponding to several languages ​​for which one or more machine learning models are configured to translate. In these examples, one or more elements of one or more vectors associated with the target language may include a first value, such as one (and / or any other value), while one or more elements of one or more vectors associated with one or more other languages ​​may include a second value, such as zero (and / or any other value).

[0027] for example, Figure 2 Examples of generated audio representation 202 (which may include and / or resemble audio representation 108) and speech representation 204 (which may include and / or resemble speech representation 112) according to some embodiments of the present disclosure are illustrated. As shown, audio representation 202 may include several vectors 206(1)-(M) contiguously linked together (which may also be referred to in the singular as "one vector 206" or in the plural as "a plurality of vectors 206") (wherein only three are labeled for clarity). For example, each vector 206 may be associated with a different temporal instance. Additionally, vector 206 may include dimensions 208(1)-(N) (which may also be referred to in the singular as "one dimension 208" or in the plural as "a plurality of dimensions 208"). As described herein, vector 206 may include any number of dimensions 208, such as 512 dimensions (and / or any other number of dimensions). In some examples, audio representation 202 may correspond to acoustic embedding.

[0028] like Figure 2As further illustrated, the language representation 204 may include several vectors 210(1)-(M) connected end-to-end (which may also be referred to in singular form as "one vector 210" or in plural form as "multiple vectors 210") (where only three are labeled for clarity). For example, each vector 210 may again be associated with a different time instance. Additionally, vectors 210 may include dimensions 212(1)-(O) (which may also be referred to in singular form as "one dimension 212" or in plural form as "multiple dimensions 212"). As described herein, in some examples, dimensions 212 may be based on the number of languages ​​that one or more machine learning models are configured to translate and / or can be configured to translate in the future. For example, each dimension 212 may be associated with a corresponding language such that if one or more machine learning models are configured to translate 128 languages ​​(and / or any number of any number of dimensions for any number of languages), then there will be 128 dimensions.

[0029] exist Figure 2 In the example, different shades within elements of audio representation 202 and / or language representation 204 can represent different values ​​of the elements. For example, and regarding language representation 204, elements associated with the fourth dimension 214(4) can include first values, such as one (and / or any other value), while elements associated with the other dimensions 214(1)-(3) and 214(5)-(O) include second values, such as zero (and / or any other value). This is because in Figure 2 In the example, the user may have selected the target language represented by the fourth dimension 212 (4) for translating the speech associated with the audio representation 202. Although Figure 2 The example only illustrates selecting a single target language for translation, but in other examples, users can select any number of target languages.

[0030] Return to reference Figure 1 For example, the process 100 may include: processing one or more audio representations 108 and one or more language representations 112 using one or more projection layers 114 of one or more machine learning models; and generating one or more combined representations 116 based at least on the processing. Although Figure 1 The example illustration depicts the generation of one or more combined representations 116 using one or more projection layers 114. However, in other examples, any other type of layer of one or more machine learning models can be configured to perform a similar process described herein to generate one or more combined representations 116. Additionally, the combined representation 116 may include, but is not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation.

[0031] As described herein, one or more projection layers 114 can use one or more techniques to generate one or more combined representations 116. For example, Figure 3 The illustration depicts a first technique for combining audio representation 202 and speech representation 204 using splicing, according to some embodiments of the present disclosure. As shown, one or more projection layers 114 can splice audio representation 202 and speech representation 204 to generate a spliced ​​representation 302. Additionally, Figure 3 The example illustration shows one or more projection layers 114 performing splicing by stacking audio representations 202 on top of speech representations 204 along the feature dimension. However, in other examples, one or more projection layers 114 may perform splicing using any other technique, such as stacking speech representations 204 on top of audio representations 202.

[0032] In some examples, such as Figure 3 As further illustrated, one or more projection layers 114 can further process the spliced ​​representation 302 to generate a fused representation 304, which includes dimensions smaller than those of the spliced ​​representation 302. For example, the dimensions of the fused representation 304 may correspond to (e.g., be the same as) the dimensions of the audio representation 202. When performing this process, one or more projection layers 114 can generate the fused representation 304 using learned relationships between language information and audio. Thus, in some examples, the spliced ​​representation 302 may include a combined representation 116, while in other examples, the fused representation 304 may include a combined representation 116.

[0033] Next, Figure 4 The illustration depicts a second technique using a combination of audio representation 202 and language representation 204 with sinusoidal transform according to some embodiments of the present disclosure. As shown, language representation 204 can be processed with respect to frequency matrix 402 to generate projection representation 404. As described herein, projection representation 404 can include, but is not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation of audio data. Additionally, one or more elements of frequency matrix 402 (e.g., each element of frequency matrix 404) can be encoded at one or more different frequencies to create unique patterns for one or more (e.g., each) language dimensions. For example, different frequency matrices can be generated for different target languages, wherein each frequency matrix includes unique encoded frequency patterns associated with the respective target language.

[0034] In some examples, the speech representation 204 can be multiplied by a frequency matrix 402 that includes a specific dimension, such that the projected representation 404 represents the speech information with the same dimensions as the audio representation 202. This is as follows: Figure 4 As further illustrated in the example, one or more projection layers 114 can (such as by masking features of audio representation 202 (and / or using any other technique)) add features of projection representation 404 to features of audio representation 202 to generate combined representation 406 (which may include combined representation 116). By performing this process, one or more projection layers 114 can preserve the original features of audio representation 202 through the combination of features.

[0035] Return to reference Figure 1 For example, process 100 may include: processing one or more combined representations 116 using one or more decoders 118 of one or more machine learning models; and generating text data 120 representing text corresponding to speech and presented in the target language, at least based on the processing. As described herein, decoder 118 may include any type of decoder configured to perform one or more of the processes described herein, such as a recurrent neural network transducer (RNNT) decoder, a connection-time classification (CTC) decoder, a tagging and duration transducer (TCT), etc. In some examples, text data 120 may represent a final text version of speech, such as transcribed text corresponding to speech. However, in other examples, text data 120 may represent a tagged text version of speech. In these examples, one or more machine learning models and / or one or more other processing components may process text data 120 to generate a final text version of speech.

[0036] for example, Figure 5 The illustration shows an example of performing ASR processing and translation processing using one or more machine learning models 502 according to some embodiments of the present disclosure. As shown, one or more machine learning models 502 may receive audio data 504 representing speech presented in French and input data 506 representing selections for translating the speech into English as input. Thus, one or more machine learning models 502 may perform one or more processes described herein to generate text data 508 representing text corresponding to the speech and presented in English.

[0037] For example, one or more machine learning models 502 may use one or more encoders 106 to process audio data 504 and generate one or more audio representations associated with audio data 504. One or more machine learning models 502 may also use one or more encoders 110 to process input data 506 and generate one or more language representations associated with the target language, English. Additionally, one or more machine learning models 502 may use one or more projection layers 114 to process one or more audio representations and one or more language representations to generate one or more combined representations. Then, one or more machine learning models 502 may use one or more decoders 118 to process one or more combined representations to generate text data 508 representing text corresponding to speech and presented in English.

[0038] As described in this paper, one or more machine learning models 502 can be trained to perform one or more processes described in this paper during the ASR processing and translation processing. For example, Figure 6 Examples of processes for training one or more machine learning models 602 (which may include and / or resemble one or more machine learning models 502) to perform ASR processing and translation processing are illustrated according to some embodiments of the present disclosure.

[0039] As shown, one or more machine learning models 602 can be trained using training input data 604. In some examples, training input data 604 may include audio data representing speech instances 606. For example, speech instances may include phrases such as "Can you tell me where to find this character?", "The performance will be interesting", "Please give me directions", and / or any other speech. Training input data 604 may also represent a list of languages ​​608 associated with the translated speech instance 606. For example, in some examples, speech instance 606 may be associated with a corresponding language 608 for performing translation. In some examples, training input data 604 may be real-world generated, synthetically generated, and / or any combination thereof.

[0040] One or more machine learning models 602 can be trained using training input data 604 and corresponding ground truth data 610. As shown, in some examples, ground truth data 610 may at least include text corresponding to speech instances 606 translated into language 608. In some examples, for each instance of training input data 604, such as each speech instance 606 having a corresponding language 608, there may be corresponding ground truth data 610 representing the desired text 612. In some examples, ground truth data 610 may be real-generated, synthetically generated, human-labeled, machine-labeled, and / or any combination thereof.

[0041] As shown, the process 600 may include: one or more machine learning models 602 processing training input data 604 to generate output data 614 corresponding to the training input data 604. For example, the output data 614 may represent estimated text corresponding to a speech instance 606 translated into language 608. The process 600 may then include: one or more training engines 616 using one or more loss functions to measure the loss (e.g., error) in the output data 614 by comparing it with ground truth data 610. For example, in some examples, one or more loss functions may measure the loss based at least on the difference between the estimated text represented by the output data 614 and the text 612 represented by the ground truth data 610. The one or more training engines 616 may then perform backpropagation computation to recursively compute the gradients of one or more loss functions with respect to training parameters in order to update the parameters and / or weights of one or more machine learning models 602, indicated by arrows from one or more training engines 616 to one or more machine learning models 604. For example, one or more training engines 616 can update the parameters and / or weights of one or more encoders, one or more projection layers, and / or one or more decoders of one or more machine learning models 602.

[0042] Figure 7 Examples of one or more systems 702 according to some embodiments of the present disclosure, capable of performing one or more processes described herein to translate speech, are illustrated. As shown, one or more systems 702 may include one or more processors 704 (which may include and / or be similar to one or more CPUs 1106 and / or one or more GPUs 1108), one or more communication interfaces 706 (which may include or be similar to one or more communication interfaces 1110), and / or memory 708 (which may include and / or be similar to memory 1104). Additionally, memory 708 may store at least one or more machine learning models 502, which one or more processors 704 may then execute to perform one or more operations described herein.

[0043] For example, and as illustrated, one or more systems 702 may receive audio data 712 (which may include and / or be similar to audio data 102) representing speech from a user, and language data 714 (which may include and / or be similar to language data 104) representing the target language into which the user wishes to translate the speech, from one or more user devices 710. One or more systems 702 may then perform one or more of the processes described herein to process the audio data 712 and language data 714 using one or more machine learning models 502 to generate text data 716 (which may include and / or be similar to text data 120) representing text corresponding to the speech and presented in the target language. Additionally, one or more systems 702 may send the text data 716 back to one or more user devices 710, enabling one or more user devices 710 to provide the text to a user.

[0044] Now, for reference Figure 8 and Figure 9 Each block of methods 800 and 900 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be executed by a processor that executes instructions stored in memory. These methods 800 and 900 can also be embodied as computer-usable instructions stored on a computer storage medium. These methods 800 and 900 can be provided by standalone applications, services (standalone or combined with another managed service), managed services, or plug-ins to another product, to name a few. Additionally, regarding... Figure 1 These methods 800 and 900 are described by way of example. However, these methods 800 and 900 may be additionally or alternatively performed by any system or any combination of systems, including but not limited to the systems described herein.

[0045] Figure 8 A flowchart of a method 800 for translating speech into text using a multilingual model, according to some embodiments of the present disclosure, is illustrated. At block B802, the method 800 may include generating one or more audio vectors associated with an embedding space using one or more encoders and based at least on audio data representing speech. For example, one or more encoders 106 may process audio data 102 representing speech to generate one or more audio vectors associated with an embedding space, wherein one or more audio vectors may be represented by one or more audio representations 108.

[0046] At box B804, the method 800 may include generating one or more language vectors corresponding to the target language. For example, one or more encoders 110 may process language data 104 representing the target language to generate one or more language vectors corresponding to the target language, wherein one or more language vectors may be represented by one or more language representations 112. As described herein, in some examples, one or more language vectors may include dimensions associated with several target languages ​​that can be used when performing speech translation. In such examples, one or more language vectors may include a first value of one or more elements associated with the target language and a second value of one or more other elements associated with one or more other languages.

[0047] At box B806, the method 800 may include generating one or more combined vectors based on at least one or more audio vectors and one or more language vectors. For example, one or more projection layers 114 may process one or more audio vectors and one or more language vectors to generate one or more combined vectors, wherein one or more combined vectors may be represented by one or more combined representations 116. As described herein, in some examples, one or more projection layers 114 may generate one or more combined vectors by concatenating one or more audio vectors with one or more language vectors and / or projecting one or more concatenated vectors back to the dimension associated with one or more audio vectors. In some examples, one or more projection layers 114 may generate one or more combined vectors by multiplying one or more language vectors by a frequency-associated matrix and then adding the result to one or more audio vectors.

[0048] At box B808, the method 800 may include: generating output data representing text corresponding to speech and rendered in the target language using one or more decoders and based at least on one or more combined vectors. For example, one or more decoders 118 may process one or more combined vectors to generate text data 120 representing text corresponding to speech and rendered in the target language. As described herein, in some examples, text data 120 may represent a final text version of speech, such as transcribed text corresponding to speech. However, in other examples, text data 120 may represent a tokenized version of speech. In these examples, one or more other processing components may process text data 120 to generate a final text version of speech.

[0049] At box B810, the method 800 may include: generating output associated with text. For example, text may be displayed to one or more users, text data 120 may be sent to one or more user devices, and then the one or more user devices may output the text, audio data representing the text may be used to output speech associated with the text, the text may be processed using one or more additional processing components, and / or any other type of output may be provided.

[0050] Figure 9 A flowchart of a method 900 for translating speech into text presented in a target language, according to some embodiments of the present disclosure, is illustrated. At block B902, the method 900 may include generating one or more audio representations associated with the speech. For example, one or more encoders 106 may process audio data representing the speech to generate one or more audio representations 108. As described herein, audio representations 108 may include, but are not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation of the audio data.

[0051] At box B904, method 900 may include generating one or more language representations associated with a target language. For example, one or more encoders 110 may process language data 104 representing the target language to generate one or more language representations 112 associated with the target language. As described herein, language representation 112 may include, but is not limited to, one or more vectors, one or more embeddings, one or more tensors, one or more kernels, one or more features, and / or any other type of representation. Additionally, one or more language representations 112 may include one or more first values ​​of one or more elements associated with the target language and one or more second values ​​of one or more other elements associated with one or more other languages.

[0052] At box B906, the method 900 may include generating one or more combined representations based on at least one or more audio representations and one or more speech representations. For example, one or more projection layers 114 may generate one or more combined representations 116 using one or more audio representations 108 and one or more speech representations 112. As described herein, in some examples, one or more projection layers 114 may generate one or more combined representations 116 by concatenating one or more audio representations 108 with one or more speech representations 112 and / or projecting one or more concatenated representations back to the dimension associated with one or more audio representations 108. In some examples, one or more projection layers 114 may generate one or more combined vectors 116 by multiplying one or more speech representations 112 by a frequency-associated matrix and then adding the result to one or more audio representations 108.

[0053] At box B908, method 900 may include: generating output data representing text corresponding to speech and rendered in the target language using one or more decoders and based at least on one or more combined representations. For example, one or more decoders 118 may process one or more combined representations 116 to generate text data 120 representing text corresponding to speech and rendered in the target language. As described herein, in some examples, text data 120 may represent a final text version of speech, such as transcribed text corresponding to speech. However, in other examples, text data 120 may represent a tokenized version of speech. In these examples, one or more other processing components may process text data 120 to generate a final text version of speech.

[0054] Example language model In at least some embodiments, language models such as Large Language Models (LLM), Visual Language Models (VLM), Multimodal Language Models (MMLM), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD formats such as OpenUSD), and / or the like based on context provided in input prompts or queries. In embodiments, these language models may be considered “large” because they are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—e.g., millions or billions of parameters. LLM / VLM / MMLM / etc. can be implemented for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), extracting insights from data (e.g., text, images, videos, etc.), and generating new text / images / videos / etc. in a user-specified style, tone, and / or format. In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein may be specifically designed for text processing, while in others, a multimodal LLM may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio, 2D and / or 3D data (e.g., USD format) and / or video. For example, a Visual Language Model (VLM) or more specifically a Multimodal Language Model (MMLM) may be implemented to accept images, video, audio, text, 3D designs (e.g., CAD) and / or other input data types and / or generate or output images, video, audio, text, 3D designs and / or other output data types.

[0055] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented using different techniques to understand and generate outputs (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, converter architectures (e.g., architectures relying on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. may also include one or more diffusion blocks (e.g., noise reduction blocks). The LLM / VLM / MMLM / etc. of this disclosure may include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc., including encoder and decoder components (e.g., T5 (Text-to-Text Transformer)), can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting and any architecture type (including, but not limited to, those described herein) can be implemented depending on the specific implementation and the task performed using LLM / VLM / MMLM / etc.

[0056] In various embodiments, LLM / VLM / MMLM / etc. can be trained using unsupervised learning, where LLM / VLM / MMLM / etc. learns patterns from a large amount of unlabeled text / audio / video / image / design / USD / etc. data. Due to the extensive training, in these embodiments, the model may not require task-specific or domain-specific training. An LLM / VLM / MMLM / etc. extensively pre-trained on a large amount of unlabeled data can be referred to as a base model and can excel at various tasks, such as question answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as cue tuning, fine-tuning, retrieval augmentation generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers to tune or adjust cues or labels to bias the language model towards a specific task or domain), and / or using optimization models for specific tasks and / or other fine-tuning or customization techniques for specific domains.

[0057] In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using LLM / VLM / MMLM / etc., and / or to prevent the output or presentation of information generated by LLM / VLM / MMLM / etc. (e.g., displays, audio outputs, etc.). In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with the model's inputs and / or outputs. For example, these "protective" models can be trained to identify "safe" or otherwise okay or desired inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a particular application / implementation. Therefore, the LLM / VLM / MMLM / etc. disclosed herein are unlikely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, insecure, out of scope, and / or unwanted for a particular application / implementation.

[0058] In some embodiments, an LLM / VLM / etc. can be configured or able to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations where the model is not ideally suited, the model may have instructions for accessing one or more plugins (e.g., third-party plugins) to help process the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such an example, when at least part of the prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least part of the response requires mathematical computation, the model can access one or more mathematical plugins or APIs to help solve the problem, and then the response from the plugins and / or APIs can be used in the model's output. This process can be repeated (e.g., recursively) an arbitrary number of iterations, using any number of plugins and / or APIs, until a response to each query / question / request / process / operation / etc. can be generated in response to the input prompt. Therefore, models can rely not only on their own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (such as APIs, plugins, etc.).

[0059] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple hints provided to the same language model or instances of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, the same input query and hints (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents, for example, providing more than one hint to constrain, guide, or otherwise influence the style, content, or character of the provided output. In one or more exemplary non-limiting embodiments, the same language model can be required to provide output corresponding to different roles, perspectives, characters, or different knowledge bases, as defined by the provided hints.

[0060] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated proxies of at least one language model, and / or provided to two or more prompts for at least one language model can be further processed, such as aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or proxy) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be required to generate or otherwise obtain output about the input source material, wherein the output is associated with the input source material. Such association can include, for example, generating captions or text portions embedded (e.g., as metadata) within the input source text or image. In one or more embodiments, the output of the language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, the language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, wherein the text or image is annotated to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a curatorial dataset, for example, but not limited to this.

[0061] Figure 10A This is a block diagram of an example generative language model system 1000 suitable for implementing at least some embodiments of the present disclosure. Figure 10A In the example shown, the generative language model system 1000 includes a retrieval augmentation (RAG) component 1092, an input processor 1005, a tokenizer 1010, an embedding component 1020, a plugin / API 1095, and a generative language model (LM) 1030 (which may include LLM, VLM, multimodal LM, etc.).

[0062] At a high level, the input processor 1005 can receive input 1001, which includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, generic scene descriptor (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 1030 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, input 1001 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 1001 may include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., tabular format, JSON, or XML). In the generative LM In some implementations of 1030 capable of handling multimodal input, input 1001 can combine text (or text that may be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking raw input text as an example, input processor 1005 can prepare the raw input text in various ways. For example, input processor 1005 can perform various types of text filtering to remove noise from the relevant text content (e.g., special characters, punctuation marks, HTML tags, stop words, portions of images, portions of audio, etc.). In examples involving stop words (common words that often have little semantic meaning), input processor 1005 can remove stop words to reduce noise and allow the generative LM 1030 to focus on more meaningful content. Input processor 1005 can apply text normalization, for example, by converting all characters to lowercase, removing accent marks, and / or handling special cases (such as abbreviations or abbreviations) to ensure consistency. These are just a few examples; other types of input processing can be applied.

[0063] In some embodiments, RAG component 1092 (which may include one or more RAG models, and / or may be performed using generative LM 1030 itself) may be used to retrieve additional information to be used as part of input 1001 or a prompt. RAGs can be used to enhance input to LLM / VLM / MMLM / etc. with external knowledge to make the answer to a specific question or query or request more relevant, for example, where specific knowledge is required. RAG component 1092 may obtain this additional information (e.g., basic information such as basic text / images / videos / audio / USD / CAD / etc.) from one or more external sources, and then feed it along with the prompt to LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.

[0064] For example, in some embodiments, in addition to the data retrieved using RAG component 1092, input 1001 may also be generated using query or model inputs (e.g., questions, requests, etc.). In some embodiments, input processor 1005 may analyze input 1001 and communicate with RAG component 1092 (or in embodiments, RAG component 1092 may be part of input processor 1005) to identify relevant text and / or other data to provide to generative LM 1030 as additional context or information sources, typically from which responses, answers, or outputs 1090 are identified. For example, when the input indicates that a user is interested in the required tire pressure for a particular brand and model of vehicle, RAG component 1092 may use a RAG model, for example, to perform a vector search in the embedding space to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the owner's manual for that particular vehicle brand and model. Similarly, when a user revisits the chatbot related to a specific product sale or service, the RAG component 1092 can retrieve previously stored conversation history (or at least its summary) and provide the previous conversation history, along with the current inquiry / request, as part of the generative LM 1030 as input 1001.

[0065] RAG component 1092 can use various RAG techniques. For example, naïve RAG can be used, where documents are indexed, chunked, and applied to an embedding model to generate embeddings corresponding to chunks. User queries can also be applied to this embedding model and / or another embedding model of RAG component 1092, and the embeddings of chunks can be compared with the embeddings of the query to identify the most similar / relevant embeddings to the query. These most similar / relevant embeddings can be provided to generative LM 1030 to generate output.

[0066] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.

[0067] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.

[0068] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which can lead to a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide the model with structured entity information by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, or it may be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can summarize the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.

[0069] In any embodiment, the RAG component 1092 may implement plugins, APIs, user interfaces, and / or other functions to perform RAG. For example, LLM / VLM / MMLM / etc. may use graph RAG plugins to run queries on knowledge graphs to extract relevant information to feed into the model, and may use standard or vector RAG plugins to run queries on vector databases. For example, the graph database may interact with the plugin's REST interface, thus decoupling the graph database from the vector database and / or the embedded model.

[0070] The tokenizer 1010 can segment (e.g., processed) text data into smaller units (tags) for subsequent analysis and processing. Depending on the implementation, the tags can represent individual words, sub-words, characters, audio / video / images, etc. Word-based tokenization divides the text into individual words, treating each word as a separate tag. Sub-word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 1030 to understand morphological changes and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate tag, enabling the generative LM 1030 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, the tokenizer 1010 can transform (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.

[0071] Embedding component 1020 can use any known embedding technique to transform discrete tokens into semantically meaningful (e.g., dense, continuous vector) representations. For example, embedding component 1020 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.

[0072] In some implementations where input 1001 includes image data / video data, etc., input processor 1005 may resize the data to a standard size compatible with the format of the corresponding input channel and / or normalize pixel values ​​to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 1020 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where input 1001 includes audio data, input processor 1005 may resample the audio file to a consistent sampling rate for uniform processing, and embedding component 1020 may use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where input 1001 includes video data, input processor 1005 may extract frames or apply resizing to extracted frames, and embedding component 1020 may extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where input 1001 includes multimodal data, the embedded component 1020 can use techniques such as early fusion (concatenation), late fusion (sequential processing), and attention-based fusion (e.g., self-attention, cross-attention) to fuse representations of different types of data (e.g., text, images, audio, USD, video, design, etc.).

[0073] Other components of the generative LM 1030 and / or generative LM system 1000 may use different types of neural network architectures depending on the implementation scheme. For example, a transducer-based architecture (such as the one used in models like GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feedforward network that processes the output of the self-attention layer, applying a nonlinear transformation to the input representation and extracting higher-level features. Some non-limiting example architectures include transducers (e.g., encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning), etc. Therefore, depending on the implementation scheme and architecture, the embedded component 1020 can apply the encoded representation of the input 1001 to the generative LM 1030, and the generative LM 1030 can process the encoded representation of the input 1001 to generate an output 1090, which may include response text and / or other types of data.

[0074] As described herein, in some embodiments, the generative LM 1030 may be configured to access or use (or be able to access or use) plugins / APIs 1095 (which may include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations where the generative LM 1030 is not ideally suited, the model may have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 1092) to access one or more plugins / APIs 1095 (e.g., third-party plugins) to help process the current input. In such an example, when at least part of the prompt is related to a restaurant or weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs), sending at least part of the prompt related to a particular plugin / API 1095 to the plugin / API 1095, which may process the information and return an answer to the generative LM 1030, which may use the response to generate output 1090. This process can be repeated (e.g., recursively) an arbitrary number of iterations and repeated using any number of plugins / APIs 1095 until an output 1090 that resolves each query / question / request / process / action / etc. from input 1001 is generated. Therefore, the model can rely not only on its own knowledge acquired from training on a large dataset and / or from data retrieved using the RAG component 1092, but also on the expertise or optimized properties of one or more external resources (e.g., plugins / APIs 1095).

[0075] Figure 10B This is a block diagram of an example implementation scheme, where the generative LM 1030 includes a converter encoder-decoder. For example, suppose the input text (e.g., "Who discovered gravity") is tokenized (e.g., by...) Figure 10A The tokenizer 1010 is used for tokens such as words, and each token is encoded (e.g., by...). Figure 10A The embedding component 1020 is a corresponding embedding (e.g., of size 512). Since these token embeddings do not typically represent the position of the tokens in the input sequence, positional encoding can be added to each token embedding using any known technique to encode the order relation and context of the tokens in the input sequence. Thus, the (e.g., the resulting) embeddings can be applied to one or more encoders 1035 of the generative LM1030.

[0076] In the example implementation, encoder 1035 forms an encoder stack, where each encoder includes a self-attention layer and a feedforward network. In the example converter architecture, each token (e.g., a word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, pass each vector through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used. For example, to compute a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. The self-attention score for a token pair can be computed by taking the dot product of the query vector and the corresponding key vector, normalizing the resulting score, multiplying by the corresponding value vector, and summing the weighted value vectors. The encoder can apply multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate context vectors that encode the input. Attention projection layer 1040 can transform the context vectors into attention vectors (keys and values) for decoder 1045.

[0077] In the example implementation, decoder 1045 forms a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network. Similar to encoder 1035, in the example converter architecture, each token (e.g., a word) flows through a separate path in decoder 1045. During the first pass, decoder 1045, classifier 1050, and generation mechanism 1055 can generate a first token, and generation mechanism 1055 can apply the generated token as input during a second pass. This process can be repeated cyclically, generating tokens (e.g., words) and adding them to the output of the previous pass, and in subsequent passes applying token embeddings of positionally encoded composite sequences as input to decoder 1045, generating one token at a time (called autoregression) until a symbol or token representing the end of the response is predicted. In each decoder, the self-attention layer is typically restricted to focusing only on preceding positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In the example implementation, the encoder-decoder attention layer operates similarly to the (e.g., multi-head) self-attention operation in encoder 1035, except that it creates its query from the layer below it and obtains keys and values ​​(e.g., matrices) from the output of encoder 1035.

[0078] Therefore, decoder 1045 can output some decoded (e.g., vector) representation of the input applied during a particular pass. Classifier 1050 can include a multi-class classifier comprising one or more neural network layers and a softmax operation that transforms logit probabilities into probabilities, the neural network layers projecting the decoded (e.g., vector) representation onto corresponding dimensions (e.g., one dimension for each supported word or token in the output vocabulary). Thus, generation mechanism 1055 can select or sample words or tokens based on corresponding predicted probabilities (e.g., selecting the word with the highest predicted probability) and append it to the output of the previous pass, thereby generating each word or token sequentially. Generation mechanism 1055 can repeat this process, triggering successive decoder inputs and corresponding predictions until a symbol or token representing the end of the response is selected or sampled, at which point generation mechanism 1055 can output the generated response.

[0079] Figure 10C This is a block diagram of an example implementation where the generative LM 1030 includes a decoder-only converter architecture. For example, Figure 10C The decoder 1060 can be used with Figure 10B The decoder 1045 operates similarly, except... Figure 10C Each decoder 1060 omits the encoder-decoder self-attention layer (because there is no encoder in this implementation). Therefore, the decoders 1060 can form a decoder stack, where each decoder includes a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or tag indicating the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., a corresponding embedding with positional encoding) can be applied to the decoder 1060. Figure 10B Similar to decoder 1045, each tag (e.g., a word) can flow through a separate path in decoder 1060, and decoder 1060, classifier 1065, and generation mechanism 1070 can use autoregression to generate one tag at a time sequentially until a symbol or tag indicating the end of the response is predicted. Classifier 1065 and generation mechanism 1070 can be... Figure 10B The classifier 1050 and generation mechanism 1055 operate similarly, wherein the generation mechanism 1070 selects or samples each consecutive output label based on the corresponding predicted probability and appends it to the output of the previous iteration, generating each label sequentially until a symbol or label representing the end of the response is selected or sampled. The architectures described herein, and others, are merely examples, and other suitable architectures may be implemented within the scope of this disclosure.

[0080] Example computing device Figure 11This is a block diagram of an example computing device 1100 suitable for implementing some embodiments of the present disclosure. The computing device 1100 may include an interconnect system 1102 directly or indirectly coupled to the following devices: memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communication interface 1110, input / output (I / O) ports 1112, input / output components 1114, a power supply 1116, one or more presentation components 1118 (e.g., one or more displays), and one or more logic units 1120. In at least one embodiment, one or more computing devices 1100 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1108 may include one or more vGPUs, one or more CPUs 1106 may include one or more vCPUs, and / or one or more logic units 1120 may include one or more virtual logic units. Thus, one or more computing devices 1100 may include discrete components (e.g., a full GPU dedicated to computing device 1100), virtual components (e.g., a portion of the GPU dedicated to computing device 1100), or combinations thereof.

[0081] although Figure 11 The various blocks are shown as connected via interconnect system 1102 using lines, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 1118 (such as a display device) may be considered I / O component 1114 (e.g., if the display is a touchscreen). As another example, CPU 1106 and / or GPU 1108 may include memory (e.g., memory 1104 may represent a storage device other than the memory of GPU 1108, CPU 1106, and / or other components). In other words, Figure 11 The computing devices described are for illustrative purposes only. No distinction is made between such categories as “workstation,” “server,” “laptop computer,” “desktop computer,” “tablet computer,” “client device,” “mobile device,” “handheld device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are considered within the scope of this category. Figure 11 Within the scope of computing devices.

[0082] Interconnect system 1102 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1102 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Fast Peripheral Component Interconnect (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 1106 may be directly connected to memory 1104. Further, CPU 1106 may be directly connected to GPU 1108. In cases where there is a direct or point-to-point connection between components, interconnect system 1102 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required to be included in computing device 1100.

[0083] The memory 1104 may include any computer-readable medium from a variety of computer-readable media. A computer-readable medium may be any available medium accessible by the computing device 1100. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0084] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1104 may store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Universal Disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 1100. As used herein, computer storage media does not include the signal itself.

[0085] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and includes any information transmission medium. The term "modulated data signal" can refer to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the above should also be included within the scope of computer-readable media.

[0086] CPU 1106 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. Each CPU 1106 may contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. CPU 1106 may contain any type of processor and may contain different types of processors depending on the type of computing device 1100 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1200, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as math coprocessors), computing device 1100 may also include one or more CPUs 1106.

[0087] In addition to or in lieu of one or more CPUs 1106, one or more GPUs 1108 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. One or more GPUs 1108 may be integrated GPUs (e.g., with one or more CPUs 1106) and / or one or more GPUs 1108 may be discrete GPUs. In embodiments, one or more GPUs 1108 may be coprocessors of one or more CPUs 1106. GPUs 1108 may be used by computing device 1100 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPUs 1108 may be used for general-purpose computing on a GPU (GPGPU). GPUs 1108 may include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. GPUs 1108 may produce pixel data of an output image in response to rendering commands (e.g., rendering commands received from CPUs 1106 via a host interface). GPU 1108 may include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 1104. GPU 1108 may include two or more GPUs operating in parallel (e.g., via links). The links may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 1108 may produce pixel data or GPGPU data for different portions of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for an analog image). Each GPU may include its own memory or may share memory with other GPUs.

[0088] In addition to or in lieu of CPU 1106 and / or GPU 1108, logic unit 1120 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPUs 1106, one or more GPUs 1108, and / or one or more logic units 1120 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 1120 may be a portion of one or more CPUs 1106 and / or GPUs 1108 and / or integrated into one or more CPUs 1106 and / or GPUs 1108, and / or one or more logic units 1120 may be discrete components or otherwise external to CPUs 1106 and / or GPUs 1108. In an embodiment, one or more of the logic units 1120 may be coprocessors of one or more of the CPU 1106 and / or one or more of the GPU 1108.

[0089] Examples of logic unit 1120 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree lateral unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or fast peripheral component interconnect (PCIe) element, etc.

[0090] Communication interface 1110 may include one or more receivers, transmitters, and / or transceivers that allow computing device 1100 to communicate with other computing devices via electronic communication networks (including wired and / or wireless communications). Communication interface 1110 may include components and functions that allow communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or wirelessband), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, logic unit 1120 and / or communication interface 1110 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 1102 to one or more GPUs 1108 (e.g., memory of one or more GPUs 1108).

[0091] I / O port 1112 enables computing device 1100 to be logically coupled to other devices including I / O component 1114, one or more presentation components 1118, and / or other components, some of which may be built into (e.g., integrated into) computing device 1100. Illustrative I / O component 1114 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, etc. I / O component 1114 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, pen recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 1200. Computing device 1100 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. Additionally, computing device 1100 may include an accelerometer or gyroscope that allows motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, computing device 1100 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.

[0092] Power supply 1116 may include a hardwired power supply, a battery power supply, or a combination thereof. Power supply 1116 may provide power to computing device 1100 to allow the components of computing device 1100 to operate.

[0093] The presentation component 1118 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 1118 may receive data from other components (e.g., GPU 1108, CPU 1106, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0094] Example Data Center Figure 12 An example data center 1200 that may be used in at least one embodiment of this disclosure is shown. The data center 1200 may include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.

[0095] like Figure 12 As shown, the data center infrastructure layer 1210 may include a resource coordinator 1212, grouped computing resources 1214, and node computing resources (“nodes CRs”) 1216(1)-1216(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CRs 1216(1)-1216(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CRs 1216(1)-1216(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CRs 1216(1)-1216(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of nodes CRs 1216(1)-1216(N) may correspond to virtual machines (VMs).

[0096] In at least one embodiment, the grouped computing resources 1214 may include individual groups of node CRs 1216 housed within one or more racks (not shown), or a plurality of racks housed within a data center in different geographical locations (also not shown). Individual groups of node CRs 1216 within the grouped computing resources 1214 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, a plurality of node CRs 1216, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0097] Resource coordinator 1212 may be configured or otherwise control one or more nodes CRs 1216(1)-1216(N) and / or groups of computing resources 1214. In at least one embodiment, resource coordinator 1212 may include a Software Design Infrastructure (“SDI”) management entity for data center 1200. Resource coordinator 1212 may include hardware, software, or some combination thereof.

[0098] In at least one embodiment, such as Figure 12As shown, framework layer 1220 may include job scheduler 1233, configuration manager 1234, resource manager 1236, and / or distributed file system 1238. Framework layer 1220 may include a framework of software 1232 supporting software layer 1230 and / or one or more applications 1242 supporting application layer 1240. Software 1232 or application 1242 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 1220 may be, but is not limited to, a free and open-source software web application framework (such as Apache Spark™ (hereinafter referred to as "Spark")) that can leverage distributed file system 1238 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 1233 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 1200. Configuration manager 1234 may be able to configure different layers, such as software layer 1230 and framework layer 1220 (which includes Spark and distributed file system 1238 for supporting large-scale data processing). Resource manager 1236 may be able to manage computing resources mapped to or allocated to clusters of distributed file system 1238 and job scheduler 1233, or allocated to support clusters of distributed file system 1238 and job scheduler 1233. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1214 in data center infrastructure layer 1210. Resource manager 1236 may coordinate with resource coordinator 1212 to manage these mapped or allocated computing resources.

[0099] In at least one embodiment, the software 1232 included in software layer 1230 may include software used in at least a portion of the nodes CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1238 of framework layer 1220. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.

[0100] In at least one embodiment, the application 1242 included in the application layer 1240 may include one or more types of applications used at least in part by nodes CRs 1216(1)-1216(N), grouped computing resources 1214, and / or the distributed file system 1238 of the framework layer 1220. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.

[0101] In at least one embodiment, any of the configuration manager 1234, resource manager 1236, and resource coordinator 1212 can implement any number and type of self-modification actions based on any amount and type of data obtained in any technically feasible manner. Self-modification actions can free the data center operator of data center 1200 from making potentially poor configuration decisions and may prevent underutilization and / or poor performance of the data center.

[0102] According to one or more embodiments described herein, data center 1200 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 1200 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1200 by using weight parameters computed through one or more training techniques, such as, but not limited to, those described herein.

[0103] In at least one embodiment, the data center 1200 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.

[0104] Example network environment A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 11 This is implemented on one or more instances of computing devices 1100—for example, each device may include similar components, features, and / or functions of one or more computing devices 1100. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 1200, examples of which are described in this document. Figure 12 To describe in more detail.

[0105] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0106] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.

[0107] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or application may respectively include network-based service software or applications. In embodiments, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer may be, but is not limited to, a free and open-source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").

[0108] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0109] One or more client devices may include the information described in this article. Figure 11 At least some of the components, features, and functions of one or more example computing devices 1100 described. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these depicted devices, or any other suitable device.

[0110] This disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, that are executed by a computer or other machine, such as a personal data assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This disclosure can also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.

[0111] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0112] This document describes in detail the subject matter of this disclosure to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the discloser has envisioned that the claimed subject matter may be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as suggesting any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

[0113] Example paragraph A: A method comprising: generating one or more audio vectors associated with an embedding space using one or more first encoders of one or more machine learning models and based at least on audio data representing speech; generating one or more language vectors corresponding to a target language using the one or more machine learning models; generating one or more combined vectors based at least on the one or more audio vectors and the one or more language vectors; generating output data representing text corresponding to the speech and presented in the target language using one or more decoders of the one or more machine learning models and based at least on the one or more combined vectors; and generating an output associated with the text.

[0114] B: As described in paragraph A, wherein: one of the one or more language vectors includes several elements associated with several target languages; a first element from the several elements associated with the target language includes a first value; and one or more second elements from the several elements associated with one or more other languages ​​include a second value.

[0115] C: The method described in paragraph A or paragraph B, wherein generating the one or more combined vectors includes: concatenating the one or more audio vectors with the one or more language vectors.

[0116] D: The method described in any of the segments AC further includes: generating one or more fusion features based at least on processing the one or more combined vectors using one or more projection layers of the one or more machine learning models, wherein the output data representing the text is generated using the one or more decoders and based at least on the one or more fusion features.

[0117] E: The method described in any of the paragraphs AD further includes: generating one or more second language vectors based at least on the one or more language vectors and a matrix representing one or more frequencies associated with the target language, wherein the generation of the one or more combined vectors is based at least on the one or more audio vectors and the one or more second language vectors.

[0118] F: The method described in paragraph E, wherein generating the one or more combined vectors is based at least on masking the one or more audio vectors using the one or more second language vectors.

[0119] G: The method as described in any segment AF further includes: receiving from a user equipment at least one of the audio data or input data representing the target language, wherein the generation of the one or more language vectors is based at least on the audio data or the input data.

[0120] H: The method described in any segment of segment AG, wherein: the output data represents one or more tokens corresponding to the text; and the method further includes: using the one or more machine learning models and based at least on the one or more tokens to generate the text corresponding to the speech and presented in the target language.

[0121] 1: A system comprising: one or more processors configured to: generate one or more audio representations associated with an embedding space, at least based on audio data representing speech; generate one or more language representations corresponding to a target language; generate one or more combined representations using the one or more audio representations and the one or more target representations; and generate output data representing text corresponding to the speech and presented in the target language using one or more machine learning models and at least based on the one or more combined representations.

[0122] J: The system as described in paragraph I, wherein: one of the one or more language representations includes several elements associated with several target languages; a first element associated with the target language from the several elements includes a first value; and one or more second elements associated with one or more other languages ​​from the several elements include a second value.

[0123] K: The system as described in paragraph I or paragraph J, wherein the one or more combined representations are generated at least based on splicing the one or more audio representations with the one or more language representations.

[0124] L: A system as described in any of the paragraphs IK, wherein the one or more processors are further configured to: generate one or more fusion features based at least on processing the one or more combined representations using one or more projection layers, wherein the output data representing the text is generated using the one or more machine learning models and at least based on the one or more fusion features.

[0125] M: A system as described in any of the paragraphs in IK, wherein the one or more processors are further configured to: generate one or more second language representations based at least on the one or more language representations and a matrix representing one or more frequencies associated with the target language, wherein the one or more combined representations are generated based at least on the one or more audio representations and the one or more second language representations.

[0126] N: A system as described in any paragraph of the IM, wherein the one or more processors are further configured to: receive from a user equipment the audio data and input data representing the target language, wherein the one or more language representations are generated at least based on the input data; and send to the user equipment the output data representing the text.

[0127] O: A system as described in any of the paragraphs IN, wherein: the output data represents one or more tokens corresponding to the text; and the one or more processors are further configured to use the one or more machine learning models and at least based on the one or more tokens to generate the text corresponding to the speech and presented in the target language.

[0128] P: The system as described in any of the segments IO, wherein the one or more processors are further configured to: generate one or more second audio representations associated with the embedding space based at least on the audio data; generate one or more second language representations corresponding to the second target language; generate one or more second combined representations using the one or more second audio representations and the one or more second target representations; and generate second output data representing second text corresponding to the speech and presented in the second target language using the one or more machine learning models and at least based on the one or more second combined representations.

[0129] Q: The system as described in any segment of the IP, wherein: the one or more audio representations include one or more audio vectors associated with one or more timestamps corresponding to the audio data; and the one or more language representations include one or more language vectors associated with the one or more timestamps.

[0130] R: A system as described in any paragraph of Section IQ, wherein the system includes at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for providing one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using a Systems that perform operations using one or more visual language models (VLMs); systems that perform operations using one or more multimodal language models; systems that perform one or more conversational AI operations; systems that generate synthetic data; systems that render at least one of virtual reality content, augmented reality content, or mixed reality content; systems that implement one or more multimodal language models; systems that use or deploy one or more inference microservices; systems that include one or more machine learning models deployed in services or microservices along with operating system OS-level virtualization packages (e.g., containers); systems that include one or more virtual machines (VMs); systems that are at least partially implemented in a data center; or systems that are at least partially implemented using cloud computing resources.

[0131] S: One or more processors, comprising: a processing circuitry system configured to: generate output data representing text corresponding to the speech and presented in the target language by using a machine learning model and based at least on concatenating one or more audio vectors associated with the speech and one or more language vectors associated with the target language; and generate output associated with the speech and presented in the target language.

[0132] T: One or more processors as described in paragraph S, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for providing one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); and a system for performing operations using large language models (LLMs). Systems that perform operations using one or more visual language models (VLMs); systems that perform operations using one or more multimodal language models; systems that perform one or more conversational AI operations; systems that generate synthetic data; systems that render at least one of virtual reality content, augmented reality content, or mixed reality content; systems that implement one or more multimodal language models; systems that use or deploy one or more inference microservices; systems that include one or more machine learning models deployed in services or microservices along with operating system OS-level virtualization packages (e.g., containers); systems that include one or more virtual machines (VMs); systems that are at least partially implemented in a data center; or systems that are at least partially implemented using cloud computing resources.

Claims

1. A method comprising: One or more first encoders using one or more machine learning models and based at least on audio data representing speech to generate one or more audio vectors associated with the embedding space; Use one or more machine learning models to generate one or more language vectors corresponding to the target language; One or more combined vectors are generated based on at least one or more audio vectors and one or more language vectors; Output data representing text corresponding to the speech and presented in the target language is generated using one or more decoders of the one or more machine learning models and based at least on the one or more combined vectors; as well as Generate output associated with the text.

2. The method of claim 1, wherein: One of the one or more language vectors includes several elements associated with several target languages; The first element associated with the target language from the plurality of elements includes a first value; and One or more second elements associated with one or more other languages ​​from the plurality of elements include a second value.

3. The method as described in claim 1, wherein, Generating the one or more combined vectors includes concatenating the one or more audio vectors with the one or more language vectors.

4. The method of claim 1, further comprising: One or more fused features are generated by processing the one or more combined vectors using at least one or more projection layers of the one or more machine learning models. The output data representing the text is generated using one or more decoders and is based on at least one or more fusion features.

5. The method of claim 1, further comprising: One or more second language vectors are generated based at least on the one or more language vectors and a matrix representing one or more frequencies associated with the target language. The generation of the one or more combined vectors is based at least on the one or more audio vectors and the one or more second language vectors.

6. The method of claim 5, wherein, The generation of the one or more combined vectors is based at least on masking the one or more audio vectors using the one or more second language vectors.

7. The method of claim 1, further comprising: Receive at least one of the audio data or input data representing the target language from the user equipment. The generation of the one or more language vectors is based on at least one of the audio data or the input data.

8. The method of claim 1, wherein: The output data represents one or more tags corresponding to the text; as well as The method further includes: using the one or more machine learning models and based at least on the one or more tags to generate the text corresponding to the speech and presented in the target language.

9. A system comprising: One or more processors, said one or more processors being used for: Generate one or more audio representations associated with the embedding space, based at least on audio data representing speech; Generate one or more language representations corresponding to the target language; One or more combined representations are generated using the one or more audio representations and the one or more target representations; as well as Output data representing text corresponding to the speech and presented in the target language is generated using one or more machine learning models and based at least on the one or more combined representations.

10. The system of claim 9, wherein: One of the one or more language representations includes several elements associated with several target languages; The first element associated with the target language from the plurality of elements includes a first value; and One or more second elements associated with one or more other languages ​​from the plurality of elements include a second value.

11. The system of claim 9, wherein, The one or more combined representations are generated at least based on concatenating the one or more audio representations with the one or more language representations.

12. The system of claim 9, wherein, The one or more processors are also used for: One or more fused features are generated by processing the one or more combined representations using at least one or more projection layers. The output data representing the text is generated using one or more machine learning models and based on at least one or more fused features.

13. The system of claim 9, wherein, The one or more processors are also used for: One or more second language representations are generated based at least on the one or more language representations and a matrix representing one or more frequencies associated with the target language. The one or more combined representations are generated based at least on the one or more audio representations and the one or more second language representations.

14. The system of claim 9, wherein, The one or more processors are also used for: Receive audio data and input data representing the target language from a user equipment, wherein the language representation is generated based at least on the input data; as well as The output data representing the text is sent to the user equipment.

15. The system of claim 9, wherein: The output data represents one or more tags corresponding to the text; and The one or more processors are further configured to use the one or more machine learning models and at least based on the one or more tags to generate the text corresponding to the speech and presented in the target language.

16. The system of claim 9, wherein, The one or more processors are also used for: At least based on the audio data, generate one or more second audio representations associated with the embedding space; Generate one or more second language representations corresponding to the second target language; One or more second combined representations are generated using the one or more second audio representations and the one or more second target representations; as well as Second output data is generated using one or more machine learning models and based at least on one or more second combined representations, representing a second text that corresponds to the speech and is presented in the second target language.

17. The system of claim 9, wherein: The one or more audio representations include one or more audio vectors associated with one or more timestamps corresponding to the audio data; as well as The one or more language representations include one or more language vectors associated with the one or more timestamps.

18. The system of claim 9, wherein the system is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system that provides one or more cloud gaming applications; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system that implements one or more multimodal language models; A system that uses or deploys one or more inference microservices; A system that includes a service or microservice in which one or more machine learning models are deployed together with an operating system (OS) level virtualization package (e.g., a container); A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

19. One or more processors, comprising: Processing circuit system, the processing circuit system being used for: The machine learning model is used to generate output data representing text corresponding to the speech and presented in the target language, based at least on concatenating one or more audio vectors associated with the speech and one or more language vectors associated with the target language. as well as Generate output that is associated with the speech and presented in the target language.

20. One or more processors as claimed in claim 19, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system that provides one or more cloud gaming applications; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system that implements one or more multimodal language models; A system that uses or deploys one or more inference microservices; A system that includes a service or microservice in which one or more machine learning models are deployed together with an operating system (OS) level virtualization package (e.g., a container); A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.