Perceptual whitening and adaptive joint processing cut bitrate for arbitrary 3D audio channels while preserving audio quality.
Adaptive DSHT rotation decorrelates HOA channels before perceptual coding, reducing noise unmasking and side information overhead.
Image-based sound source location data adjusts coherence and spread metadata to improve multi-microphone spatial audio reproduction.
Adjusts elevation panning and filter coefficients to reduce audio image distortion when 3D audio is downmixed from non-standard heights.
Voice spectrogram matching extracts phonetic indicators during calls to verify identity, block unauthorized access, and save processing resources.
Layer-specific HOA extension payloads improve partial sound field reconstruction and decoding quality when higher layers are missing or invalid.
Adjusts voice authentication conditions by utterance length or syllable count to reduce fixed speaking-time burdens while preserving accuracy.
Sliding-window emotion detection and face animation models cut lag while keeping multi-user facial movements synchronized to speech.
Integer winTDAC and INTMDCT enable real-time switching between lossy and lossless audio codecs with low overhead and no perceived noise.
Perceptual channel ranking lets HOA SPAR coding balance bitrate limits, metadata size, and low-latency immersive audio quality.
Self-supervised features and a generative model recover cleaner speech from distorted coded audio without raising bitrate.
Adjusts direct and indirect sound per source, giving speech stronger reverb reduction to improve intelligibility without flattening non-speech audio.
Adaptive bitrate allocation between downmix channels and spatial metadata cuts codec overhead and bit wastage in immersive audio encoding.
Precomputed acoustic transfer functions suppress controller and environmental noise while preserving voice quality with lower latency.
A two-stage denoising chain combines ML spectrogram cleanup with stationary-noise removal to cut noise while preserving phase and harmonics.
LPC bandwidth sharpening extends filter impulse response to conceal lost LFE audio frames while reducing low-frequency rumble artifacts.
A shared audio encoder with separate SID and SSD heads blocks synthetic speech spoofing while avoiding the overhead of separate models.
Spatial parameters are remapped through covariance matrices to mix coded audio with lower latency, less computation, and preserved sound quality.
Saliency-based ambisonics assigns higher-order encoding to attended sound sources and lower-order encoding to others, cutting latency and power use.
A new SBR frame class uses transient position cues to localize fine grid areas, cutting pre-echo, bit use, and decoding delay.
Modular decoding classes turn unaligned PER cockpit voice recorder hex data into human-readable objects for faster aircraft analysis.
Adaptive spectral patching switches among harmonic, SBR, and non-linear methods to extend bandwidth with fewer artifacts and lower complexity.
A purified common signal from monaural and stereo decodes improves stereo output quality without abandoning independent decoding paths.
Transfers voice-initiated payments from a shared smart assistant to the right personal device, keeping card data private while preserving convenience.
Multiple directional cues and covariance synthesis cut bitrate for object-based audio while preserving spatial fidelity and reducing artifacts.
A unified geometry format with domain conversion cuts redundant AR/VR bitstreams while preserving identical data for multiple processing blocks.
Transient-only padding before spectral phase modification prevents temporal aliasing and preserves audio quality with lower processing overhead.
AI-generated sub-flow hopping lets call representatives switch topics in real time, cutting claim delays and manual communication steps.
A higher-order Ambisonics decoder uses region-specific panning and a pseudo-inverse matrix to improve stereo localization and suppress negative side lobes.
Switching among spectral patching algorithms improves high-frequency audio reconstruction while avoiding blocking artifacts and excess processing.
Priority speech sources get stronger indirect-audio reduction, improving intelligibility while preserving non-speech quality and spatial realism.
Mix audio streams with different sample rates in the MDCT domain to avoid sample rate conversion delay and preserve audio quality.
Edge acoustic encoders compress audio into embedding vectors for central classification, reducing data load and protecting privacy.
Spatially weighted submixes let object-based audio work with channel processing while preserving metadata for immersive playback.
Native AI in the EHR turns provider speech and patient data into charts and forms, cutting manual entry errors and charting time.
Synchronized far-field and near-field microphones feed a speech enhancement network to remove room reverberation and echo in conferencing audio.
Voice biometrics replace manual PIN entry for emergency personnel, cutting call setup time, reducing errors, and strengthening access security.
Audio embeddings and similarity metrics merge same-talker profiles or split mixed ones, improving automatic enrollment accuracy.
Personalized 3D avatars link animations, audio, and detected emotions to deliver nuanced, consistent expression across chat, video, and audio.
External microphones on a separate device isolate target speech with acoustic fingerprints, improving hearable clarity in noisy group settings.
FIFO buffers and polynomial kernels add temporal convolution to spatiotemporal neural networks, cutting edge AI compute and power use.
Classifying sound and non-sound audio lets the SoC switch resource configurations to cut power use without losing processing capability.
Band-pass filtering, spectral subtraction, and AI models isolate abnormal poultry sounds from noise in high-density houses for earlier disease detection.
Direct sound filtering after AEC cuts residual echo in multi-speaker audio, improving voice wake-up and call quality.
Estimated parameter replacement limits error propagation in multichannel predictive decoding while preserving bandwidth and audio quality.
Probability-based control points let pseudorandom animations track audio characteristics for dynamic overlays in messaging.
Passive enrollment, embedding extraction, and adaptive clustering keep multi-speaker voice profiles current while reducing false authentication.
Buffers audio packets by timestamp and description index so multiple description streams are preserved and decoded audio quality improves.
Binary data is encoded by comparing paired frequency amplitudes, improving reconstruction after compression with minimal audible impact.
Adaptive hangover frame selection marks background-noise frames in SID signaling, improving comfort noise quality without hurting DTX efficiency.