Stored alternative speech hypotheses carry across devices and dialog sessions to correct misrecognitions with less user repetition and compute load.
Pretrained TTS and ASR models synthesize paired data to improve low-resource language accuracy with less training data.
Buffered recognition data lets multi-channel speech interaction reuse partial results to cut delay and resource use in multi-user scenarios.
A web server maps voice commands to voice-enabled webpages, preserving continuous site interaction when visual or keyboard-based pages block voice use.
Harmonic components above 410 Hz help distinguish phonetic notes and chords, improving vowel-focused speech recognition and synthesis.
A two-stage wakeword detector uses local DSP detection and a companion-trained user model to improve accuracy while lowering battery use.
Token-level confidence and silence checks flag low-accuracy video captions for manual review while high-confidence captions publish automatically.
Real-time audio and screen analysis automates sharing at the right moment while reducing interruptions and accidental exposure of sensitive content.
Transforms CNN stride and dilation parameters to balance compute load and accuracy with one model across devices.
Jointly trained SVDF encoder-decoder layers cut hotword spotting complexity and resource use while improving multilingual streaming detection.
Voice, term-group, and remark analysis work together to score each utterance in real time and guide conversations with timely cues.
Compacting or batching short audio streams before ASR transcription cuts fixed per-file charges while preserving accurate text output.
Machine-learned forced cough signatures enable biometric authentication while flagging respiratory anomalies through baseline mismatch detection.
Dynamic task routing keeps suitable AI tasks on the client device to protect sensitive data while limiting compute load and network transfer.
A content-biased C-DSVAE preserves phonetic structure in speaker and content embeddings, improving zero-shot voice conversion stability.
Training on videos generated by known deepfake tools helps ML models identify tool-specific artifacts and improve deepfake source detection.
When spoken replies are unclear without a screen, the system detects re-requests and rebuilds responses from recognized text for clearer feedback.
Machine-learned acoustic models adapt to home noise and dialect, improving recognition of voice commands and environmental sounds.
A single prosody and vocoder pipeline adapts acoustic features across sampling rates to generate flexible speech waveforms with less model complexity.
Prosody-driven backchannel timing helps voice assistants manage turn-taking naturally without fixed timers that slow or disrupt dialogue.
Multiple ASR hypotheses are selectively rescored with an LLM to improve domain-specific transcription accuracy without separate model training.
When an LLM cannot finish a network troubleshooting task, the system routes it to a subject matter expert to improve accuracy and response time.
A review interface lets users correct past speech interactions, feeding ASR, NLU, and TTS models to improve accuracy over time.
Converts log-Mel speech features into spike trains so spiking neural networks can run real-time recognition on power-constrained edge devices.
ATIS runway data is decoded, confirmed with air traffic control, and then used to build a safer contingency landing trajectory.
Preliminary decoding with a static class-based model triggers dynamic language models only when class terms appear, cutting speech transcription latency.
Real-time end-of-speech prediction lets a two-pass ASR model finalize long utterances faster while reducing deletion errors.
A two-stage keyword-based voice check authenticates users earlier, cutting latency while preserving robust secondary verification.
Separating speech into linguistic and nonlinguistic vectors improves emotion estimation by distinguishing expressed and intrinsic emotions.
Synthetic audio generation builds paired speech training data from semantic representations and speaker embeddings, reducing reliance on rare parallel data.
A unified speech recognition model uses contextual-mode distillation to improve streaming accuracy while reducing latency, complexity, and compute costs.
Speech-driven commands trigger LEDs, sound, or vibration so users can quickly identify the right mobile devices in a uniform-looking group.
Audio transcripts are analyzed against selected teaching variables to deliver immediate, objective educator feedback without lengthy manual observation.
Enrollment and test embeddings let streaming keyword spotting detect custom phrases more accurately with limited data and less unnecessary processing.
Real-time speech analysis selects the next text segment for voice synthesis, enabling seamless speaker handoff without manual control.
Voice intent is translated into appliance-specific control commands to reduce transmission errors and improve home appliance control reliability.
Real-time voice commands identify verified products in audiovisual content and complete secure purchases without leaving the stream.
A two-stage neural network suppresses noise through low-dimensional speech frames while preserving clarity and recognition features in real time.
When rewind alone fails, replay adds captions, voice-frequency boosting, and slower playback to make unclear dialogue easier to understand.
A client-side generative AI model splits tasks between device and server execution to reduce network load, resource conflicts, and privacy risk.
A two-stage wake-word check uses local audio screening and server neural verification to cut misrecognition without heavy CPU and memory use.
Voice conversion shifts lecturer speech toward well-recognized voice patterns, improving ASR transcripts in noisy, echo-prone lecture rooms.
A shared encoder with wakeword-specific decoders cuts compute load, enabling multiple wakewords on resource-limited devices.
Pseudo-corrections keep on-device ASR models from forgetting expired user corrections, improving recognition accuracy without long-term memory growth.
Synthetic paired speech data reduces scarce parallel-data collection for voice conversion training while preserving prosody, timing, and speaker privacy.
User-mediated threshold updates help hotword detection balance missed activations and false triggers while limiting extra audio processing and privacy risk.
A secondary wake-word engine detects false triggers and temporarily deactivates network microphones to avoid chimes, playback interruptions, and wasted compute.
A speech recognition system selects acoustic models based on conversational context to decode utterances.
A speech recognizer and command generator activate specific input terminals on a display apparatus based on voice keywords.
A speech recognition system selects a language model tailored to user intonation from multiple stored options.
A system automates speech recognition training data generation using multiple ASR engines and confidence score evaluation.
Partial word lists compile into phoneme graphs to reduce computational load during speech recognition.
Sequence-level optimization of deep belief network weights and language model scores overcomes local optima trapping during training.
A vehicle onboard computer system converts graphical user commands into voice instructions for external electronic devices.
A dynamic finite state transducer adjusts arc weights based on user context to customize speech recognition without loading separate models.
A voice conversion system selects a standard speaker from a TTS database to synthesize an intermediate base voice for subsequent morphing.
A speech recognition system constrains accessible audio metadata entries to reduce memory requirements.
A character-level emotion detection network converts raw audio signals into sentiment data without determining speech words.
A speech recognition device compresses power spectrum parameters to reduce memory bit width and hardware costs.
Multimodal deep learning models detect student cognitive engagement through visual and audio data streams.
A vehicle speech recognition system extends the listening window when a user manipulates a manual input device.
A semantic analyzer applies theme vocabulary-semantic relationship data sets to correct sentences, resolving mixed-language recognition errors.
Statistical classifier analyzes IP address, location, and search history alongside audio features to automatically select the optimal speech recognition model.
Runtime pitch detection categorizes speakers to adjust filter bank frequencies, resolving accuracy and computation time trade-offs.
A network of sound recognition devices forms acoustic beams to capture voice commands from specific directions.
Pre-extracting voice features reduces system complexity while maintaining high adaptability to user preferences.
A unified data model merges speech acts and observation history to maintain conversation continuity across chat, email, and voice interfaces.
A voice dialogue system detects skip signals to automatically execute high-priority operations based on stored execution history.
A dialog enhancement system separates audio into speech and non-speech channels using spatial analysis to apply targeted peaking filters.
Finite state transducers embed rule codes along arcs to resolve the loss of hierarchical information during automatic speech recognition processing.
Processor judges device conditions to output speech prompts, avoiding unwanted interruptions during user interaction.
Electronic device extracts keywords from call audio and displays them for immediate search.
A vehicle speech recognition system adapts acoustic models using real-time interference profiles to isolate voice commands from engine and external noise.
A central controller converts speech audio to text via automatic speech recognition, enabling natural language device control without manual setup.
Intelligent list reading segments data items based on detected request specificity, resolving information overload in voice interactions.
Segmenting the learning process isolates speaker and channel variations, resolving estimation accuracy issues caused by mixed environmental factors.
A classifier selects warping factors using formant frequencies and pitch to normalize speech data for speaker recognition.
Voice recognition system captures geoscience data via speech, resolving contamination risks and manual entry errors during microscope analysis.
A layer trajectory LSTM model uses a depth processing block to scan hidden states across time layers.
A vehicle-mounted device selects pictograms from a correspondence table using speech recognition data.
An aircraft instrumentation controller generates electronic transcripts of radio communications for flight crew reference.
A steering wheel microphone system uses a controller to cancel noise components by detecting phase differences between signals.
A smart device speech control method detects control instructions by matching them against the present operation scene.
A skill categorization system classifies software components by trust level to enable automated invocation.
Context-expanded Transformer processes multiple utterances simultaneously using self-attention to adapt hidden vectors for long audio recordings.
A job command generation device corrects text data errors from voice recognition to generate accurate commands.
Terminal tag alignment synchronizes pre-prepared text segments with audio signals, eliminating speech-to-text conversion errors and improving caption accuracy.
Information processing apparatus creates print data from voice keywords and user identification.
A speech detection system calculates audio signal grades using segment and block amplitude values to identify speech presence in real time.
A voice device merges touch and voice inputs to generate control instructions for target applications.
An adaptive equalization system adjusts speech spectral shapes to enhance intelligibility without voicing decisions.
A speech detection system extracts audio features to identify desired segments within an incoming stream.
Graph-based feature extraction resolves data sparseness by enriching intent recognition with lexical, semantic, and syntactic relationships.
A universal speech dialog system links transactions across background applications using a unified specification.
A smart terminal switches voice roles by recognizing user instructions and applying distinct role attributes to generate interactive responses.
A wake-on-voice processor generates a key-phrase model from user-provided sub-phonetic units to detect spoken commands.
A natural language processing system translates user utterances into executable SQL statements for dataset analysis.