Complex gain and phase-preference rules preserve panning continuity and spatial relationships when downmixing multichannel audio to stereo.
A shared reference LPC parameter cuts inter-channel redundancy, lowering multi-channel audio bit rate and coding complexity.
Selecting representative virtual speakers by vote values cuts 3D audio encoding complexity while preserving sound field accuracy and compression.
A three-packet scene structure maps cells to audio data for real-time spatial rendering with efficient metadata updates and lower decoder burden.
Audio-based flash detection counts treatment pulses despite distraction or background noise, helping prevent skin overtreatment.
Previous-frame status guides time-domain concealment mode selection to reconstruct lost audio frames with low complexity and no added delay.
Swapped speech spectral bands create low-level masking noise that cuts intelligibility in adjacent audio zones without loud playback.
Interim media encoding and prevalidated security tokens cut bandwidth and playback delay while protecting network content access.
Encodes audio objects with acoustic environment metadata so one bitstream supports dynamic listener poses with lower data and rendering load.
Residual spectral envelope coding preserves background-noise detail so decoder comfort noise sounds closer to the original during speech pauses.
A neural model reconstructs missing high-frequency content from narrowband audio, improving perceived sound quality without complex hardware changes.
A GMM-guided neural VAD improves speech detection in noisy multi-channel audio and enables post-filter noise reduction with less distortion.
Dynamic audio block sizing and redundancy control cut XR group session latency while preserving continuity under packet loss and congestion.
A one-time passphrase sent to a personal device enables spoken authentication through a conversational interface while limiting eavesdropping risk.
Frame-by-frame threshold updates improve stereo correlation detection and de-correlation choice, raising lossless audio compression rate.
Sound source position filters candidate speakers before voiceprint matching, improving role separation accuracy when voices are similar.
FiLM-based personal VAD improves target-speaker detection in streaming ASR without added latency, even when enrollment is skipped.
Motion-sensor-guided rotation of spatial components avoids costly ambisonics translation, improving VR audio coding efficiency and playback quality.
Starting phase modulation adds watermark payload bits through phase analysis, increasing data capacity without harming media perception.
Spectral-center filtering restores low frequencies and trims midrange distortion in ear-canal voice signals affected by loose fit and ANC.
Selective prediction of spaced harmonic spectral coefficients cuts audio coding complexity while improving reconstruction quality.
DUET peak analysis removes overlapping time-frequency points from mixed recordings, yielding purer separated voices for speech recognition.
A linear filter, compact DNN, and nonlinear post-filter cut residual noise and speech distortion in low-power single-channel audio.
Delayed crosstalk and correlation-based weighting create a mono downmix that preserves arrival cues and improves audio coding efficiency.
Phase-aware weighting of microphone-array spectra suppresses reverberation and noise while improving far-field speech recognition.
Speech forecasting splits hearing-aid audio processing across devices to cut delay and power use while improving noise reduction and intelligibility.
A modular encoder-separator-decoder architecture uses latent representations and small-context convolutions to separate overlapping speech more accurately.
Biometric unlock with accessory or voice fallback cuts redundant input, reducing authentication time, cognitive burden, and battery drain.
Call-specific watermark detection verifies authentic voice call audio and helps block fraud from recorded or manipulated speech.
A frequency-axis CNN generates denoising masks for mixed voice spectra, cutting checkerboard artifacts while supporting real-time audio enhancement.
User control data is embedded as a distinct packet in the audio stream, enabling compatible cross-device decoder and renderer interaction.
A superframe bitstream combines downmix audio and reconstruction metadata to cut bandwidth while preserving flexible immersive playback.
Quantization-aware latent compensation helps neural audio codecs cut decoder input distortion and improve decoded audio quality at limited bit rates.
XCSPE-based peak tracking separates mixed source signals in real time while avoiding precise sensor geometry requirements.
Rotating aligned sound-source directions cuts spatial parameter bit use, helping immersive audio fit limited codec bitrates.
Speech existence probability guides steering and weight vectors to separate target speech from noise with higher extraction accuracy.
A single conditioned neural network uses sampled loss weights to enhance diverse audio signals while cutting separate model and memory overhead.
Cross-correlation pattern detection identifies coincident stereo capture and biases ITD search toward zero to reduce reverberation-driven instability.
Real-time intelligibility targets tune speech enhancement to changing noise, reducing artifacts, delay, and power use in near-end audio.
Outlier sample overwrite limits burst-error artifacts in low-latency audio streams without adding FEC complexity or silicon area.
Parametric filtering and virtual loudspeaker rendering convert low-channel FOA recordings into higher-order ambisonics with sharper playback.
Fusing audio, image, and structured features by segment improves highlight recognition accuracy without excessive video processing time.
Phase-aware speaker detection and content analysis keep program meeting files updated, reducing missed progress and interrupted tracking.
Scalar quantization plus vector-quantized residuals cuts audio data while preserving sound quality through staged encoding and lossless residual coding.
Frequency-domain spectral checks validate time-domain pitch estimates, reducing multiplication errors and improving noisy signal detection.
Transformer-based speaker embeddings and nearest-neighbor thresholds improve real-time voice identification and mimicry detection.
Real-time anti-voice feedback reduces stuttering by decoupling self-voice recognition while preserving natural pitch and voice quality.
Independent bitstream blocks combine joint-coded audio with device delay and gain data to keep wireless immersive playback low-latency and robust.
Adaptive decorrelation control adjusts filter length to match scene variation, improving reverberation and ambiance in decoded spatial audio.
Bitstream-carried parameter updates let decoder neural networks adapt to bitrate and framerate changes while preserving media processing quality.