Analog-to-information circuitry extracts sparse sound features directly from signals to enable continuous monitoring with minimal energy usage.
Statistical analysis of recurrent neural network outputs determines optimal reset timing and context duration for speech recognition systems.
A speech processing system extracts formant-based feature vectors from digitized phonemes to estimate speaker age.
A coordinator device samples audio input and sends stop instructions to external devices.
Segmenting voice signals into phoneme sequences reduces false positives in real-time fraud detection by improving matching precision.
Unambiguous triggers like physical buttons enable microphones only when needed, preventing accidental activation.
A digital assistant detects dialog impasses to trigger learning sessions that adjust speech recognition parameters based on user clarification inputs.
A speech recognition system selects speaker-specific profiles based on identity and location parameters.
Electronic device rearranges multiple user intents into a coherent response sequence based on domain classification.
A voice recognizing apparatus superimposes dynamic white noise onto input signals to enhance recognition precision in multi-speaker environments.
Segmented voice processing reduces power consumption by activating high power circuits only after low power detection confirms wake-up keywords.
Distance detection dynamically constrains the set of recognized expressions, reducing false recognition rates caused by background noise interference.
A conditional multipass automatic speech recognition system processes audio frames through multiple grammar-based passes to handle diverse voice inputs.
Runtime hints dynamically select custom vocabulary lists to resolve transcription accuracy issues in domain-specific speech recognition.
A speech recognition device integrates an end-to-end neural network with a weighted finite-state transducer decoder to process acoustic features.
An ensemble byte pair encoder system merges phone and character models to improve speech recognition accuracy.
A speech recognition system encodes audio fragments during acquisition using a streaming structure to minimize computational overhead.
Low-power modules preprocess audio to trigger higher-performance units, resolving the trade-off between energy consumption and detection accuracy.
A speech recognition system selects transcription hypotheses using user context information to improve accuracy.
A processor configures a noise reduction space around a reference device to control nearby electronics.
Acoustic processing apparatus estimates sound source direction to separate specific signals from multi-channel inputs.
Aligns word and speaker confusion networks from multiple devices to resolve synchronization bottlenecks in ad-hoc meeting transcription.
A deep neural network fuses signal processing and classification layers to optimize model parameters for speech recognition.
A multichannel end-to-end speech recognition framework integrates neural beamforming to process acoustic signals directly into text.
Variable touch areas expand for low-probability components, reducing correction time and driver distraction in vehicles.
A voice search system organizes audio data into a phoneme uniterm tree structure for efficient content retrieval.
Parallel processing of equivalent sub-graphs within a weighted finite state transducer framework reduces response delay and power consumption on small devices.
A speech recognition system automatically learns new words and pronunciations by detecting unknown phoneme subsequences and tokenizing them for dictionary updates.
Display intermediate speech transcription results alongside final outputs on user devices to streamline the editing workflow.
A speech processing apparatus calculates an acoustic diversity degree to compensate recognition feature values for speaker identification.
Segmenting acoustic modeling across hierarchical linguistic levels reduces computational complexity while maintaining high recognition precision.
A speech recognition system estimates user training resources using historical accuracy data.
A deep recurrent neural network system uses stacked long short-term memory layers to process acoustic data for speech recognition.
A command keyword engine processes voice inputs locally to execute playback commands without cloud transmission.
Training an antialias filter with speech recognition loss prevents aliasing during audio down-sampling, preserving data integrity.
A speech model trained on human scores automates text-to-speech evaluation, reducing time and labor costs while maintaining reliability.
A speaker detection system adapts to individual voices by identifying unique speech features for each user.
Automated distortion classification system processes audio signals to generate acoustic feature vectors and identify specific noise types.
A two-pass automatic speech recognition model uses a shared encoder to reduce computational resources while maintaining accuracy.
A portable gateway device translates voice commands into control actions for building management systems via serial interfaces.
A voice-based transaction terminal system enables consumers to order store items through a natural language chatbot while fueling.
Autocorrelation analysis estimates decay rates to selectively mitigate reverberation, improving speech recognition accuracy across changing acoustic conditions.
Ultrasonic nasal cavity detection replaces complex multi-sensor arrays, achieving over 81.8% accuracy while reducing device complexity.
Machine learning algorithms analyze multi-source user data to identify vernacular language attributes and segment audiences for targeted marketing.
Automated skill chaining reduces user input requirements by allowing registered components to invoke dependent actions without explicit commands.
Interpolating runtime weights in a combined language model reduces perplexity by 25% and improves accuracy over static FSTs.
Synthetic mixed language datasets train a speech recognition model, improving fidelity, accuracy, and latency for Cantonese and English.
A predictive feature extraction method combines linguistic and statistical information to map noisy text.
A natural language processing model modifies user inputs using a knowledge base of previous interactions to execute computing tasks.
Dynamic acoustic models detect co-articulation errors in numeric sequences to improve speech recognition accuracy.