Separating speech and silence packets lets a base station predict call quality and tune network parameters without costly packet inspection.
QMF-domain BRIR truncation uses reverberation-based subband filter lengths to cut binaural rendering load while preserving sound quality.
A base speaker-channel layer preserves full-range playback on legacy decoders, while object extensions and metadata enable personalized immersive mixes.
Frame-level signal type selection and channel ratio quantization reduce stereo image drift and improve synthesized audio clarity.
Combining per-direction spatial audio parameters within each time-frequency tile cuts transmission bits while preserving immersive audio quality.
A modular IVAS bitstream structure with common, tool, metadata, and EVS payload sections cuts codec complexity while supporting immersive audio formats.
Separate voiceprint and spoofprint embeddings help detect unknown synthesized speech attacks while preserving robust speaker recognition.
Dynamic masking thresholds based on band energy and hearing sensitivity improve bit allocation and preserve audio quality across level changes.
A delay unit aligns QMF waveform subbands with SBR metadata, enabling low-delay bitstream splicing without re-sampling.
A hierarchical filter-bank generative model predicts magnitude and phase together, enabling parallel audio processing without Griffin-Lim reconstruction.
An audio-sensing charging case detects predefined environmental sounds and alerts ANC earphone users without draining earbud battery life.
Encodes shared static and dynamic acoustic metadata for multiple listening poses, cutting immersive audio data rate while preserving rendering flexibility.
Prototype-buffer copying and overlap-add fill lost MDCT analysis windows, reducing unreliable IFFT endpoint artifacts in audio decoding.
Adaptive phase and spectrum adjustment improves lost audio frame concealment, reducing discontinuities and tonal artifacts during burst errors.
Client-side browser echo suppression cuts feedback and noise before cloud distribution, improving audio clarity in large virtual meetings.
Multiple biofeedback sensors and machine learning assess participant trustworthiness in real time while providing a confidence metric.
Frequency band-wise noise mixup helps tiny DNN speech enhancement improve intelligibility and perceptual quality on edge devices.
By embedding scene audio metadata in the bitstream, the decoder skips local generation and speeds binaural rendering for XR audio.
ASR encoder features replace clean reference signals to estimate speech intelligibility with lower complexity and faster evaluation.
A compact GRU and self-attention speech enhancer cuts on-device compute demands while preserving noise suppression and clean magnitude estimation.
Preconfigured user profiles let a shared device identify the requester and choose the right media service faster with fewer inputs and lower battery use.
Embedded scene-audio metadata lets the decoder render binaural signals faster without local metadata generation or external retrieval.
Adaptive neural separation adjusts how many sound sources are extracted, preserving key audio while limiting computation and power use.
Non-acoustic sensor fusion classifies mobility scenes so mobile audio processing can adapt in changing environments and improve dialogue intelligibility.
Transfers voice-initiated payments from a shared smart assistant to the right personal device, keeping card data private while enabling flexible payment options.
Coordinated uplink and downlink filtering suppresses background noise while preserving voice quality in real-time audio capture and playback.
An asymmetric loss and adjusted speech mask reduce noise while preserving fricatives, laughter, and applause in low-SNR audio.
Truncated BRIR subband filters enable real-time binaural rendering of mixed channel and object audio with lower computational load.
Adaptive smoothing and subband feature analysis help reconstruct missing high frequencies more accurately for clearer music playback.
A single learning model separates vocals, guitar, piano, and noise to cut training cost and memory use in limited-capacity audio devices.
Bit-efficient spatial audio encoding compares sub-band direction data with audio object metadata to cut bitrate while preserving audio quality.
Constraining spatial, gain, and volume adaptation in output metadata preserves renderer flexibility while improving audio quality-to-bitrate ratio.
Virtual loudspeaker rendering and HOA normalization bound absolute gain values, enabling minimum-bit coding for random access and lower bitrate.
Low-frequency speech is prioritized as an independent substream, enabling recovery under packet loss while optional high-frequency data restores quality.
Audio signatures tied to subjects let users replay exact content segments and search related media without manual rewinding or name recall.
Machine learning turns motion and audio sensor data into high-quality target sound, avoiding wearable microphones, extra weight, and voice leakage.
Dynamic watermark strength keeps embedded markers detectable after audio processing, helping monitor corruption without unnecessary audio quality loss.
Personalized cry models use selected high-quality audio to improve baby cry assessment accuracy even without constant server connectivity.
An intermediate audio format and render-feedback training loop improve multichannel object separation, convergence, and time continuity.
Automatically shares photos by proximity, group context, or facial match to cut manual steps and keep shared media easy to find.
A three-tier DirAC approach splits low-, mid-, and high-order Ambisonics synthesis to limit artifacts and preserve spatial audio quality at low bitrates.
Masked feature-sequence learning improves voice conversion by preserving time-frequency structure while converting nonverbal and paralanguage cues.
Tile- and subband-based tonal coding restores high-frequency components more accurately while keeping audio coding bit rate low.
Latent audio encoding and forecasting cut hearing-aid processing delay while improving noise suppression and speech intelligibility.
A multi-stage model isolates singing vocals, extracts fakeprint artifacts, and improves deepfake detection and singer attribution in mixed audio.
AI analyzes recorded sales pitches to deliver personalized training feedback and help recruiters match candidates to specific campaigns.
Universal adversarial audio masks emotional cues from smart-speaker speech while preserving wake-word use and transcription accuracy.
Two-pass Siamese training conditions each frame on past separations, improving real-time speech separation while reducing training-inference mismatch.
Locally generated acoustic representations let nearby devices authenticate a user by shared location while keeping sensitive speech off the network.
Predicts and fills missing high-frequency spectrum regions during decoding to balance energy and improve reconstructed audio quality.