Geometric time warping, amplitude scaling, and a denoising vocoder turn mono audio into binaural output without binaural training data.
Allocating audio sources along voxel-derived line segments enables realistic 3D extent rendering with lower complexity in XR audio scenes.
Parallel audio sessions and voice detection let the control panel identify the active device faster during emergencies and cut response delays.
Frequency-dependent allpass filtering adds elevation and spatial cues to monaural audio without coloring output channels.
Neural manifold mapping and entropy maximization compress audio and video while extending frequency range and preserving standard format compatibility.
Audio is mapped into a higher-dimensional spectral domain so AI transformers can remove redundancy and raise compression without degrading fidelity.
Temporal prediction and residual latent coding cut audio redundancy, enabling lower-bitrate speech transmission with low delay.
Selects frame-specific IPD extraction modes from signal features to preserve phase information while controlling audio coding resources.
Bitstream-carried neural network update parameters let decoders adapt to bitrate and framerate changes while preserving media processing quality.
Cross-channel in-air and in-ear microphone processing detects clipping and pseudo signals, then restores speech for clearer recognition.
Transient-aware block grouping and neural encoding improve multi-channel audio compression quality and reconstruction for sharp signal changes.
High-band tone detection and noise floor parameters help recover tone components more accurately and improve decoded audio quality.
A visual listening-space interface lets creators place sound images directly and derive stereo gain values for accurate audio-video alignment.
A hierarchical neural model generates filter-bank audio directly, avoiding phase reconstruction while improving frequency-domain integration and parallel synthesis.
Frame-level high-frequency gain uses monaural decoded audio to restore stereo high-frequency energy and improve sound quality.
Upmixed monaural and common signals are combined to refine stereo decoded audio when separate code streams leave useful information unused.
Pre-encoding the frame overlap lets stereo audio coding match monaural delay while preserving embedded multichannel decoding.
A scene tree encodes 3D audio source positions with relative origins and identifiers to cut bitrate while preserving spatial rendering accuracy.
Microphone and sensor feedback detects attention-critical sounds and temporarily overrides noise cancelation without interrupting audio playback.
A dual-filter audio pipeline recreates the rate-change filter on demand to avoid data loss and keep playback smooth during speed adjustments.
Switching among spectral patching algorithms extends audio bandwidth while reducing blocking artifacts and avoiding extra domain transforms.
A segmented IVAS bitstream uses headers, metadata, and EVS payloads to support immersive audio modes without excessive decoding complexity.
Shared spatiotemporal covariance estimation links dereverberation and source separation to cut computation and improve convergence.
A two-stage ML pipeline predicts acoustic features, applies spectral masking, and generates studio-quality audio with less noise and reverberation.
Dynamic TDAC block switching improves audio coding efficiency by balancing frequency resolution, time resolution, and pre-echo control.
A speaker-channel bed plus object metadata preserves legacy playback while enabling personalized immersive audio rendering.
CNN-RNN audio encoding preserves timing information while quantizing features to cut encoded data volume and maintain decoded audio quality.
Independent left and right channel extraction and fusion reduce inter-channel correlation and improve positioning in multichannel upmixing.
Modulo differential coding and shared probability tables cut bitrate and memory use while preserving audio object reconstruction quality.
Power-spectrum-based LP parameter conversion enables seamless frame sampling-rate switching while lowering codec complexity and preserving sound quality.
Separating low- and high-frequency audio features enables lower-bandwidth encoding while preserving reconstruction quality.
Random loss-weight conditioning lets one audio neural network handle speech, music, codecs, and bitrates with lower compute and memory use.
Adversarial encoder-decoder reconstruction conceals long audio packet losses with deterministic output and fewer perceptible VoIP distortions.
ResCNN voice fingerprint training with softmax and liveness checks improves remote exam candidate verification and reduces impersonation risk.
A hybrid Kalman-DNN AEC approach suppresses nonlinear echo, adapts to changing echo paths, and preserves target audio in double-talk.
Delta coding and Rice encoding improve lossless compression of small data sets, cutting transfer time and power use without data loss.
Interchannel correlation and phase-opposition indicators guide downmix mode selection to preserve mono energy and phase quality.
Spectral shaping and filtered aliasing-cancellation signals enable smoother cross-domain frame transitions with good audio quality and moderate bitrate overhead.
Independent left and right channel extraction reduces inter-channel correlation and improves positioning accuracy in stereo audio upmixing.
A second frame syntax flag tells the decoder when FAC data is present, preventing parsing failures after frame loss during mode switching.
Adaptive virtual speaker reselection cuts frame-to-frame fluctuation in 3D audio encoding, improving reconstructed sound quality with lower data demand.
Downmix-based dialogue separation reconstructs multichannel speech with projection coefficients, avoiding phase distortion and heavy training.
Shared LP filter coefficients and adaptive channel mixing preserve stereo quality and intelligibility at low bit-rates in complex audio scenes.
Selective neural gain prediction and wind filtering suppress one-channel wind noise while preserving desired audio content and spatial balance.
Reusing primary-channel coding parameters lets stereo audio keep intelligibility and spatial quality at lower bit-rates and low delay.
A mixed IEC 60958 audio stream carries compressed and linear PCM signals together to preserve quality, cut decoding delay, and simplify playback.
Multiple offset-frequency comparisons help reconstruct high-frequency audio bands with better harmonic balance at lower bit rates.
Speech spectrograms and pose sequences are combined to generate sign language video that preserves emotion without gloss-heavy annotation.
Neural architecture search automates hyperparameter tuning for sound source separation models, cutting redesign effort and compute use.
Domain-invariant features and one-class learning cut false alarms and preserve speech deepfake detection under codec and channel distortions.