Constraining speaker turn detection to word boundaries reduces false positives and improves statistics estimation for audio diarization.
Activity predictor forecasts user intent using sensor data, reducing false alarms in noisy environments.
A depth camera computes scene depth maps to determine speaker distance for selecting acoustic models.
Pitch synchronous analysis aligns acoustic feature extraction with pitch marks to resolve phoneme duration mismatch in speech synthesis.
A dynamic noise adaptation model uses Bayesian selection to weight speech inputs against a null noise baseline.
Dynamic parameter adjustment based on real-time noise estimation reduces false triggers and latency in low-power speech recognition systems.
A probabilistic generative model extracts speaker embeddings to assign labels.
A non-linear multilayer perceptron merges micro-modulation and cepstral features to filter noise and improve speech recognition accuracy.
A view hierarchy file maps speech text to screen elements for dynamic interaction.
A terminal matches voice information with preset data to authenticate users and execute commands.
Distributing phone states across local caches enables parallel processing of acoustic frames for faster model alignment.
A dialog system detects discrepancies between user model predictions and speech recognition results to repair data.
Local feedback data personalizes speech recognition tasks, resolving the contradiction between server efficiency and individual user preference adaptation.
Calculating word coincidence ratios enables partial utterance recognition, resolving the trade-off between usability and exact text matching.
A speech processing system automatically captures a second utterance to correct an erroneous transcribed word.
A spoken dialog system analyzes prosodic cues to determine utterance prominence for accurate speech recognition.
A data processing device generates visual speaker representations from voice signals using modality transfer functions.
A voice response system manages concurrent user interactions through dynamic occupation transitions.
A reduced user dictionary creation unit extracts necessary words from a full user dictionary to support speech recognition processing.
A noise suppressor classifies audio signals to improve speech encoding accuracy.
A parameter prediction device inputs environmental characteristics and target evaluation values into a model to generate optimal control parameters.
Phoneme language model segments speech signals into connection phonemes to reduce vocabulary search space and improve recognition speed.
A unified model aligns enhancement and recognition objectives to resolve far-field accuracy drops caused by low signal-to-noise ratios.
Iterative voice conversion training augments low-resource datasets, resolving accuracy-efficiency trade-offs in speech processing.
Automated agents represent users in virtual meetings, resolving delays when users must manage conflicting schedules and active participation.
A speech recognition engine enables user-configurable commands via a wireless headset.
Segmenting audio into speech and non-speech components improves resolution accuracy while managing processing complexity.
Locale-specific hotword classifiers identify speech language to select the correct recognition model.
A distributed voice recognition system segments processing between client and server to handle acoustic and language tasks.
Processing audio data extracts metadata to identify shared dialogues, conserving network resources and processor cycles.
A pre-signal detector wakes analog-to-digital converters only when ambient audio meets a condition.
A speech recognition model construction method aligns transcripts from multiple speakers along a time axis to synthesize mixed speaker data.
Clustering phones into garbage units builds a decoding network that reduces memory usage and improves wake-up accuracy.
Dynamic model generation captures frame dependencies without offline clustering, improving recognition reliability while reducing system latency.
A service orchestration layer transmits personalized pending user actions to equipment during hold periods.
A mobile communication terminal monitors voice sessions to detect and display key words on a screen.
A speech-to-text system uses neural networks to recognize user utterances by processing machine-generated text as context.
A display unit presents function items in two languages simultaneously to support multilingual operation.
A display controlling unit manages speech recognition information visibility based on detected voice input operations.
A vehicle speech controller adjusts dialogue response times based on real-time acceleration data to manage driver interaction.
A speech recognition apparatus adjusts processing speed based on detected user features and environmental conditions.
Probabilistic Linear Discriminant Analysis normalizes speech representations to isolate emotion-dependent features from speaker identity.
A voice recognition system selects models based on signal-to-noise ratio.
Combining finite state transducers with segmented language models reduces computational time while maintaining high accuracy for domain-specific vocabulary.
A voice detector calculates signal-to-noise ratios for frequency sub-bands and applies a non-linear power function to form a single decision value.
A server extracts audio feature values to match preset models and determine forewarning levels for notification output.
Local audio firewall processing identifies wake words before transmission, preventing sensitive data exposure to remote servers.
Natural language query arrangement processes descriptive phrases to identify surrounding objects, reducing user burden from memorizing exact names.
On-device interrupt architecture detects speech with low latency to reduce false activations while maintaining seamless user interaction.
A semiconductor integrated circuit device dynamically adjusts standard pattern ranges and recognition accuracy parameters for speech signals.