A speech processing method combines speech features with text bottleneck features using a trained unidirectional LSTM model.
A phoneme-based speech analytics system identifies meaningful phrases by analyzing audio sequences without generating text transcripts.
A context-aware routing system segregates speech processing into domains to improve interaction accuracy.
Semantic filtering removes redundant ASR hypotheses to resolve disambiguation bottlenecks and improve user experience.
A voice communication system verifies message duration against a preset proportion to detect incomplete recordings before transmission.
A hybrid parametric and exemplar-based model extracts high-level structures to enhance prosodic expressiveness in speech synthesis.
A speech recognition device selects acoustic models using multi-sensor data to improve accuracy.
Grammar testing filters invalid speech hypotheses, reducing computational load during recognition.
An intermediary noise suppression model analyzes audio-environment metrics to improve speech recognition accuracy in noisy environments.
Modifying second scaling vectors with user utterance data personalizes speech recognition accuracy while reducing training time and computational resources.
Prosody models resolve flat speech issues by training separate style models that select directives based on context confidence.
Embedded browser routes audio through a WebRTC integration layer to enable voice calls on devices without SIM cards.
A communication intermediary system integrates third-party accounts to enable cross-platform messaging and calling sessions.
A subtitle generation method extracts text from video backgrounds to update generated subtitles.
A voice processing method divides long audio segments by detecting breath locations to maintain accurate phrase boundaries.
Multiple voice receivers feed a detector that separates main speech from background noise, routing signals to distinct channels for targeted elimination.
A spoken word generation system uses mode detection to switch between training and recognition states.
A voice-enabled device dynamically updates its local speech recognition model using detected voice actions to expand vocabulary.
A speech processing device filters in-vehicle bus data and uses code phrases to adjust vehicle elements, reducing driver distraction from manual controls.
Unsupervised segmentation and recursive dynamic time warping detect repeated phrases in dialog systems regardless of word order variations.
Merging voice and handwriting probabilities resolves recognition accuracy issues caused by complex individual handwriting styles.
Information processing system extracts voice commands from broadcast audio using interference signals for precise text conversion.
A speech recognition learning system adapts acoustic models using converted audio signal vectors to improve synthesis accuracy.
Direct processing of multi-frame audio feature vectors into a preset intent recognition model eliminates transcription errors from background noise and accents.
A voice authentication system processes spoken commands to generate payment authorization requests for merchant devices.
A voice identification system selects meeting audio samples to establish speaker voiceprints without dedicated enrollment sessions.
A speech retrieval system detects coinciding segments by comparing word and phoneme recognition results to calculate evaluation values.
Segmenting context-free grammars into a back-off module handles out-of-grammar utterances, resolving the contradiction between recognition speed and accuracy.
A multi-factor audio watermarking system detects media content to suppress automated assistant query processing.