REDUCE RESPONSE TIMES IN CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
DE102026102306A1Undetermined Publication Date: 2026-07-23NVIDIA CORP
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-23
Smart Images

Figure 00000000_0000_ABST
Abstract
In various examples, response times (e.g., latencies) associated with conversational artificial intelligence (AI) systems can be reduced by sharing partial results between modules or components of the conversational AI systems. For instance, instead of waiting for an automatic speech recognition (ASR) system to finish converting a user utterance into text data, the systems of the present revelation can receive candidate prefixes for the utterance from the ASR system and use a language model to predict the complete utterance based on the candidate prefixes, as well as generate responses to the predicted utterances. If additional information is received (e.g., remaining parts of the utterance), the language model can update the predicted utterance and / or the response.Furthermore, in some cases the systems of the present disclosure can begin forwarding the response to a text-to-speech (TTS) system before the language model finishes generating the response.
Need to check novelty before this filing date? Find Prior Art