A dynamically-adjustable listening timeout system adapts speech recognition parameters to user input patterns.
A conversion method transforms non-back-off language models into back-off formats using entropy-based pruning to optimize speech decoding.
A hybrid speech data processing system digitizes vehicle inputs into packets and transmits them via wireless channels to a recognition server.
A speech synthesis device uses a style encoder to extract emotional characteristics from reference audio for natural output.
A speech classification probability calculation unit estimates cluster membership using a generative model.
Emulation software creates virtual voice assistant devices to enable scalable load testing, resolving hardware resource constraints during development.
A mobile voice platform processes acoustic signals via automated speech recognition to identify and initiate user service requests.
A parallel uplink and downlink architecture streams speech data and recognition results simultaneously over separate HTTP connections.
A keyword confirmation system detects contiguous silence fragments in audio data to verify effective keywords.
A speech emotion recognition model calculates distance metrics between sample embeddings and pre-registered category representations.
Element-wise multiplication of encoder and prediction vectors improves interaction modeling between acoustic and language models.
Display apparatus extracts similar words based on reliability values and presents a selection list for user confirmation.
Unsupervised task vectors adapt primary ASR models to new speech tasks, eliminating the need for transcribed training data.
A multichannel acoustic signal processing method calculates inter-channel feature similarity to select and separate only relevant input channels.
A customizable communication system enables users to select regional dialect and language preferences during voice command interactions.
A speech recognition system selects model configurations to balance accuracy and latency.
A local speech recognition engine adapts its grammar using high-confidence results from a server-based system to maintain accuracy during connectivity loss.
Phonetic posteriorgrams from speaker-independent ASR map to acoustic features via bidirectional LSTM networks, eliminating parallel data requirements.
A phoneme lattice indexing unit structures voice data frames to enable precise keyword matching and representative word extraction.
A parameter generation model creates dialect parameters applied to a trained speech recognition model, resolving low accuracy for complex dialects.
A voiced sound interval detection device clusters multidimensional power spectrum vectors to determine signal-to-noise ratios for accurate voice segmentation.
A speaker verification system uses co-location data to identify users across multiple devices.
Multiple beamformers mix directional and non-directional power spectral density signals to enlarge the voice control sweet spot.
An adaptive dialog system dynamically selects output candidates based on classifier probability distributions.
PCEN normalization stabilizes input energy to reduce bias in streaming inference while maintaining low latency.
A speech recognition device extracts fundamental frequency and spectrum parameters from digital audio signals to identify the speaker.
A system generates customized video summaries by separating content into semantic clips and scoring relevance based on user preferences.
A speech recognition system extracts partial syllable sequences to identify unknown words in input audio.
A handheld device transmits audio signals to a remote server for processing.
An audio guidance generation device accumulates competition messages and synthesizes explanatory speech to convey real-time event data.
A voice interaction mediator routes user audio inputs to multiple agents and converts control signals between them.
A media playback system dynamically adjusts audio speed based on listener profiles to optimize consumption.
A speech recognition apparatus adjusts acoustic models based on device state information to improve command signal accuracy.
A bookmark application monitors voice streams to detect keywords and create stored command templates.
Prosody feature extraction improves speech recognition by integrating emotional and intentional aspects into acoustic analysis.
A control apparatus dynamically adjusts the identification level of a speech section detector to optimize voice recognition accuracy.
A communication system generates personalized vocabularies by tagging network words with weights.
A post-processing speech system uses domain-specific grammars to verify natural language recognition results.
A voice interface system processes spoken commands to execute drilling operations.
Block-wise irregular pruning reduces neural network weight matrix size to accelerate automatic speech recognition inference.
Continuous audio analysis identifies voice commands directly, eliminating artificial interruptions and improving natural communication flow for users.
Triggered attention aligns frame-synchronous and label-synchronous decoders, reducing output delays while maintaining high transcription accuracy.
A device loads specific voice trigger phrases based on detected environmental context.
A speech forced alignment model evaluation method calculates time accuracy scores by comparing predicted phoneme timestamps against reference values.
A physiological cochlear model extracts place-based and time-based vectors from speech signals.
A system generates finite state grammars from sample phrases using a tree-data structure.
A smart detection circuit analyzes syllable and frequency characteristics to identify voice commands.