A vocal user interface model adapts online to speaker-specific pronunciations using Non-negative Matrix Factorization and Maximum A Posteriori estimation.
A voice response system updates its artificial intelligence model using non-verbal user feedback signals.
Voice activity detection posteriors weight frame-level features to resolve speaker recognition degradation from uniform pooling.
Bidirectional language models calculate forward and backward scores to resolve accuracy issues from unidirectional modeling.
Adaptive time and frequency smoothing estimates noise floors, reducing speech distortions during variable acoustic conditions.
A mobile computing device activates a trigger word detection subroutine to enable voice command control without manual input.
A digital assistant determines target application availability to execute user tasks on a terminal device.
A speech recognition acoustic model predicts pronunciation duration to skip intermediate frames.
A teaching agent monitors learning actions to generate labels dynamically.
A hybrid language model combines character-level and word-level components to enable open-vocabulary speech recognition.
Hierarchical parse tree scoring isolates low-confidence phrases, reducing annotation errors and limiting user repetition requests.
Fused audio training generates a universal neural network model that removes background sound without repeated learning or complex hardware.
Acoustic feature extraction identifies indoor scenes using ambient sound, bypassing GPS signal instability.
A neural network isolates speech signals from background noise by estimating frequency subbands and reconstructing clean audio.
A process engine converts audio signals into sign language animation frames.
Annotated speech mappings enable a voice API interface to translate acoustic inputs into executable requests without traditional hardware.
A voice activity detection method computes energy, spectral centroid, and tonality features from sub-band signals to determine signal activity.
A speech processing device removes reverberation components weighted by segment-specific influence degrees to enhance recognition accuracy.
A dialog engine assigns dynamic cancellation scores to spoken terms based on user actions.
A voice energy detection circuit adjusts oversampling ratios to optimize power usage.
Electronic device prompts users for missing speech parameters via a graphical interface to complete tasks.
A speech recognition system generates N-best lists for character inputs and compares pronunciation data to verify word candidates.
Trained neural networks extract voice emotion features from audio streams to classify customer emotional states in real time.
Voice keyword extraction updates channel maps, resolving recognition accuracy issues for diverse user utterances.
A wake-up network identifies speech using a garbage word list and approximate pronunciation data for portable devices.
A parallel model training platform generates domain expert models and averages their parameters to create a unified inference model.
A speech recognition control part determines whether to execute real-time processing on voice data from telephone calls.
A speech recognition system converts distorted vocalizations into text for real-time communication assistance.
A speech recognition system segments streaming audio by detecting linguistic boundaries in decoded text rather than relying on fixed time intervals.
A speech recognition system transforms feature vectors to classify pronunciation variations invariantly.
A pre-processing apparatus detects trailing silence in speech signals and adjusts the duration based on stored reference periods.
Clusters user speech data via social graph demographics to build tailored acoustic models, resolving pronunciation variations across age and accent groups.
Audio front end generates fixed and adaptive gain outputs to resolve wakeword detection versus device arbitration conflicts.
A system generates acoustic event profiles from natural language descriptions without requiring audio samples.
A subword-based end-to-end speech recognition apparatus simplifies network structure by combining estimated subword sequences into recognizable words.
A media playback system processes voice commands to identify registered users via linked profiles.
Parallel speech recognition systems abort incomplete tasks when confidence thresholds are met, reducing processing time while maintaining accuracy.
A vehicle-mounted voice recognition apparatus adjusts guidance output based on user interaction metrics.
A dual mode speech recognition system combines local and remote engines to process queries.
An audio environment manager detects surrounding sound events to automatically pause or lower music playback volume.
Dual classifiers generate decision data to select encoders, reducing misclassification artifacts.
A multimedia playback apparatus extracts frame images to query an AI model for keyword data and adjusts playback parameters automatically.
Classifying speech samples into speaker-dependent classes to extract and combine identity features, reducing confusion between unknown and known speakers.
A peripheral device control circuit integrates information from other devices to generate summary data, reducing central processing load.
A conversation-based language model adjusts n-gram probabilities using external text from multiple participants.
A computing device combines audio and image capture to identify user inputs through speech recognition and facial expression analysis.
Semi-tied covariance modules map correlated features into uncorrelated spaces for fMLLR processing, resolving robustness issues in speech recognition.
Transforms source acoustic models via LDA and VTLN to improve transcription accuracy for languages with limited training data.
Segmented neural networks compute temporal correlation vectors between frames to improve classification accuracy while reducing processing latency.
A voice interactive device detects trigger words to route content to specific services.