A hybrid voice activity detection system combines a deep neural network with digital signal processing statistical post-processing to estimate speech probabilities.
A two-stage audio classifier system combines initial and refined confidence scores to resolve false identification of rap music as speech.
Distillation techniques transfer training states between neural networks, enabling efficient acoustic model generation that handles noisy speech data.
A driving assistance apparatus recognizes driver voice commands and displays the interpreted process content on a meter for verification.
A dual-mode voice control system switches between operate-to-speak and directly-speak modes based on user operations and microphone states.
A voice detection model uses recurrent neural networks to process audio features.
A cyclic buffer queue separates speech data processing to resolve accuracy drops caused by partial content interception during wakeup word detection.
A voice-controlled data system generates phonetic transcriptions from language identification tags to enable accurate media file selection via speech commands.
Trainable semantic model classifies speech inputs using user information to resolve accuracy versus complexity trade-offs.
A streaming voice conversion system partitions audio data into independent segments processed in parallel by a multi-core processor.
Language adaptation system augments vocabulary with domain-specific terms to improve speech recognition accuracy.
Maximum likelihood methods compute offset and scaling factors to normalize speech feature vectors without preliminary transcripts.
A vehicle communication system localizes processed voice output to specific passengers using directional speakers and seat position data.
A secondary microphone captures spoken user inputs while the primary device handles phone calls, maintaining voice command availability.
Language skill profiles replace speaker-specific data to reduce computational complexity and memory consumption while maintaining speech recognition accuracy.
A minimal user-specific language model replaces generic components with tailored n-grams to eliminate unused data and reduce perplexity.
A neural vocoder system maps acoustic features to generate updated audio signals with precise pitch and rhythm control.
A universal remote control system manages multiple electronic devices through integrated speech recognition and signal generation.
A speech intent recognition device extracts user intent using skill level and dialog context information to provide tailored feedback.
A voice processing system pre-executes substitute word requests to reduce response latency.
A weighted singleton histogram table generates an optimally sorted list of elements based on computed popularity scores.
A dynamic automatic speech recognition system selects language-specific models to process digital audio input on edge devices.
Context identification modules dynamically load specified vocabularies, resolving transcription accuracy issues caused by general database limitations.
Digital assistant adjusts audio volume based on speaker identity profiles to prioritize relevant utterances.
A spoken language understanding method detects analysis abnormalities by checking predefined word slot values against intent information.
A voice action system generates intents from developer data to enable new application commands.
A controller selects voice tags based on ambient noise levels to enhance speech clarity before transmission.
An integrated circuit impulsive noise detector uses a neural network to classify speech versus noise via spectral features, reducing false detection rates.
A universal speech model estimates speaker transforms from input data to modify independent models for generating specific voice characteristics.
Audio processing device analyzes speech data to determine user demographics and interests without manual input.
A portable sound recognition apparatus extracts peak values from audio signals and calculates statistical data for efficient processing.
A speech recognition device combines noise suppression with acoustic model adaptation to improve accuracy.
Incremental device-based acoustic model adaptation using multi-model trees resolves the trade-off between reliability and data collection time.
Scripts running on telnet clients associate data identifiers with screen fields to process input and generate speech output.
Parsing grammar files into syllable networks resolves inflexibility in vehicle-mounted systems by supporting multiple grammars without cloud dependency.
A computing device generates voice profiles from acoustic features to segment user-specific speech utterances.
A direction detector assesses user viewing orientation to adapt response output forms, resolving hands-free interaction bottlenecks.
A bi-layer audio watermarking system embeds multi-segment eigenvectors to enable accurate detection via self-correlation.
A passive training method updates speaker-dependent models using spoken utterances detected by a speaker-independent model.
Aggregated multi-scale CNN extracts speech features using parallel convolution paths and time-frequency transforms.
A character-based encoder-decoder network generates feature vectors to predict words from input sequences.
Acoustic model compensator adapts speech recognition models using estimated additive and convolutive distortion factors.
A speech recognition engine generates n-grams from candidate transcriptions and weights them using user search history to select the intended query term.
A generation device produces a weighted finite state transducer embedding loops to recognize slow utterances.
A voice modulation graph displays transcribed text at corresponding audio positions to support visual editing.
A speech recognition system generates inaudible audio signals to request functions from other capable systems.