A DNN estimates vocal emotion and adjusts speaker scoring for accurate identification during laughter or angry shouts.
Precomputed zone blocking matrices separate target speech from interference and guide RTF estimation in vehicle microphone arrays.
This case reuses selected spatial parameters across frequency bands, reducing metadata bitrate while preserving spatial audio quality.
This case combines audio and video features with framed processing to separate speakers amid noise, music, and reverberation.
Speech metrics guide machine learning predictions that transform reference speech and synchronize media for tailored training.
This case correlates aspirated and alveolar-palatal frication sounds with unusual signals to locate forgery in compressed audio.
Hearing-based band quantization preserves audio quality while limiting wireless traffic.
RF-domain variability, tuning offsets, modulation, and noise are simulated to train clearer separation of overlapping radio signals.
Separate style and content encoders improve low-bitrate speech reconstruction while enabling speaker-style manipulation and privacy.
Voiceprint identification switches accounts on shared devices and visibly signals when the speaker differs from the logged-in user.
The encoder adjusts quantization by frequency sub-band to preserve spatial audio quality while meeting a defined bit budget.
This decoder separates short-term content from long-term speaker style to improve speech coding efficiency and privacy at low bitrates.
Joint GAN enhancement in a dynamic range reduced domain restores spatial information and audio quality across coded multi-channel signals.
The case matches spoken and stored names in a shared phonetic space to improve IVA caller identification accuracy.
Training with speech, room impulse responses, and noise helps preserve directional cues during binaural enhancement.
This case analyzes ultrasonic bursts above 30 kHz to distinguish gunshots, reduce false detections, and locate events.
This case combines multi-channel and time-frequency transformers with perceptual assessment to restore audio fidelity efficiently.
Correlation-based gain adjustment improves reverberation quality in multi-channel encoding.
A neural network predicts enhanced transform coefficients for adaptive blocks, reducing preecho distortion in decoded media signals.
Voiceprint authentication identifies the speaker, switches shared-device accounts, and signals mismatched users on screen.
A mixer, mute unit, and switch selectively route sound signals so required applications receive input without global cutoffs.
The framework adapts metadata precision and bitrate across audio objects to preserve immersive quality and improve noise robustness.
The model aligns anomaly degrees for frequent and infrequent normal sounds, improving acoustic anomaly detection accuracy.
Personalized models isolate each speaker while proximity-based mixing limits echoes and crosstalk among co-located devices.
Speaker embeddings and Gaussian age distributions use label-distribution loss to capture ambiguity and improve generalization.
This decoding approach uses previous audio channels to generate adaptive noise for zero-valued bands in multichannel signals.
A six-microphone array and face-guided beamforming isolate voices while detecting critical ambient sounds for timely warnings.
Time-domain correlation and frequency-domain energy checks improve very short pitch lag coding for speech and music signals.
A portable sound signal lets networked devices report unique IDs, linking them to mounting locations even when visual labels are obscured.
This case adapts BCC interleaving and LDPC tone mapping to non-consecutive dRU tones in 6 GHz LPI networks.
A parametric gain function models 3 dB distance attenuation for volumetric sources, improving XR audio accuracy with low complexity.
A latent-domain GAN and attention transformer target high compression while retaining audio fidelity and lowering storage overhead.
A first quantizer and lattice-quantized residual subvector reduce LSF coding bits while predicting higher-order audio coefficients.
Normalize each time-frequency bin against its surrounding mean to improve fingerprint completeness across the audio spectrum.
A spectrogram encoder, waveform encoder, and reference audio help neural networks reconstruct intelligible speech from noisy recordings.
This decoder approach reconstructs lost audio frames by reversing tracked coefficient signs, improving tonal quality without added delay.
External and internal microphones support alternating noise and echo cancellation, producing clearer voice signals for recognition.
Spectral and temporal-spatial analysis guides entropy-maximized manifolds for smaller files, lower bandwidth, and preserved audio quality.
Remap screen-related audio objects for accurate playback on non-centered screens.
Decibel-based dialog and talkover detection identifies human-answered calls while avoiding costly speech-to-text processing.
ASR and linguistic pattern analysis identify non-static query types, improving interpretation without manual clicks.
Biometric user identification retrieves stored media preferences, reducing manual inputs, playback errors, and device energy waste.
Context embeddings narrow voiceprint searches for cross-device authentication.
Jointly encoding residual signals helps reconstruct four audio channels with high quality while lowering bit rate and coding redundancy.
Confidence-based noise suppression improves speech perception while limiting signal distortion.
Mixing and prediction gains prioritize the primary channel while supporting efficient regeneration of scaled non-primary channels.
This case combines softmask smoothing, frequency gating, panning, and phase control to improve clarity and source levels in stereo mixes.
This case combines image sensing and local AI with an audio assistant to detect known people, reduce false positives, and adapt interaction.
A transition-aware correction process replaces improbable quantizer indices to reduce distortion and improve decoded audio quality.
Hierarchical clustering and re-segmentation address unknown speakers, overlap, and variable environments in meeting audio.