Stored alternative speech hypotheses carry across devices and dialog sessions to correct misrecognitions with less user repetition and compute load.
Pretrained TTS and ASR models synthesize paired data to improve low-resource language accuracy with less training data.
Buffered recognition data lets multi-channel speech interaction reuse partial results to cut delay and resource use in multi-user scenarios.
A web server maps voice commands to voice-enabled webpages, preserving continuous site interaction when visual or keyboard-based pages block voice use.
Harmonic components above 410 Hz help distinguish phonetic notes and chords, improving vowel-focused speech recognition and synthesis.
A two-stage wakeword detector uses local DSP detection and a companion-trained user model to improve accuracy while lowering battery use.
Token-level confidence and silence checks flag low-accuracy video captions for manual review while high-confidence captions publish automatically.
Real-time audio and screen analysis automates sharing at the right moment while reducing interruptions and accidental exposure of sensitive content.
Transforms CNN stride and dilation parameters to balance compute load and accuracy with one model across devices.
Jointly trained SVDF encoder-decoder layers cut hotword spotting complexity and resource use while improving multilingual streaming detection.
Voice, term-group, and remark analysis work together to score each utterance in real time and guide conversations with timely cues.
Compacting or batching short audio streams before ASR transcription cuts fixed per-file charges while preserving accurate text output.
Machine-learned forced cough signatures enable biometric authentication while flagging respiratory anomalies through baseline mismatch detection.
Dynamic task routing keeps suitable AI tasks on the client device to protect sensitive data while limiting compute load and network transfer.
A content-biased C-DSVAE preserves phonetic structure in speaker and content embeddings, improving zero-shot voice conversion stability.
Training on videos generated by known deepfake tools helps ML models identify tool-specific artifacts and improve deepfake source detection.
When spoken replies are unclear without a screen, the system detects re-requests and rebuilds responses from recognized text for clearer feedback.
Machine-learned acoustic models adapt to home noise and dialect, improving recognition of voice commands and environmental sounds.
A single prosody and vocoder pipeline adapts acoustic features across sampling rates to generate flexible speech waveforms with less model complexity.
Prosody-driven backchannel timing helps voice assistants manage turn-taking naturally without fixed timers that slow or disrupt dialogue.
Multiple ASR hypotheses are selectively rescored with an LLM to improve domain-specific transcription accuracy without separate model training.
When an LLM cannot finish a network troubleshooting task, the system routes it to a subject matter expert to improve accuracy and response time.
A review interface lets users correct past speech interactions, feeding ASR, NLU, and TTS models to improve accuracy over time.
Converts log-Mel speech features into spike trains so spiking neural networks can run real-time recognition on power-constrained edge devices.
ATIS runway data is decoded, confirmed with air traffic control, and then used to build a safer contingency landing trajectory.
Preliminary decoding with a static class-based model triggers dynamic language models only when class terms appear, cutting speech transcription latency.
Real-time end-of-speech prediction lets a two-pass ASR model finalize long utterances faster while reducing deletion errors.
A two-stage keyword-based voice check authenticates users earlier, cutting latency while preserving robust secondary verification.
Separating speech into linguistic and nonlinguistic vectors improves emotion estimation by distinguishing expressed and intrinsic emotions.
Synthetic audio generation builds paired speech training data from semantic representations and speaker embeddings, reducing reliance on rare parallel data.
A unified speech recognition model uses contextual-mode distillation to improve streaming accuracy while reducing latency, complexity, and compute costs.
Speech-driven commands trigger LEDs, sound, or vibration so users can quickly identify the right mobile devices in a uniform-looking group.
Audio transcripts are analyzed against selected teaching variables to deliver immediate, objective educator feedback without lengthy manual observation.
Enrollment and test embeddings let streaming keyword spotting detect custom phrases more accurately with limited data and less unnecessary processing.
Real-time speech analysis selects the next text segment for voice synthesis, enabling seamless speaker handoff without manual control.
Voice intent is translated into appliance-specific control commands to reduce transmission errors and improve home appliance control reliability.
Real-time voice commands identify verified products in audiovisual content and complete secure purchases without leaving the stream.
A two-stage neural network suppresses noise through low-dimensional speech frames while preserving clarity and recognition features in real time.
When rewind alone fails, replay adds captions, voice-frequency boosting, and slower playback to make unclear dialogue easier to understand.
A client-side generative AI model splits tasks between device and server execution to reduce network load, resource conflicts, and privacy risk.
A two-stage wake-word check uses local audio screening and server neural verification to cut misrecognition without heavy CPU and memory use.
Voice conversion shifts lecturer speech toward well-recognized voice patterns, improving ASR transcripts in noisy, echo-prone lecture rooms.
A shared encoder with wakeword-specific decoders cuts compute load, enabling multiple wakewords on resource-limited devices.
Pseudo-corrections keep on-device ASR models from forgetting expired user corrections, improving recognition accuracy without long-term memory growth.
Synthetic paired speech data reduces scarce parallel-data collection for voice conversion training while preserving prosody, timing, and speaker privacy.
User-mediated threshold updates help hotword detection balance missed activations and false triggers while limiting extra audio processing and privacy risk.
A secondary wake-word engine detects false triggers and temporarily deactivates network microphones to avoid chimes, playback interruptions, and wasted compute.
A speech recognition system selects acoustic models based on conversational context to decode utterances.
A speech recognizer and command generator activate specific input terminals on a display apparatus based on voice keywords.
A speech recognition system selects a language model tailored to user intonation from multiple stored options.
A system automates speech recognition training data generation using multiple ASR engines and confidence score evaluation.
Partial word lists compile into phoneme graphs to reduce computational load during speech recognition.