Back-to-back non-windowed DCTs cut overlap processing, reduce block discontinuities, and support smaller audio files with better playback.
Complex-number vectors replace separate real and imaginary filtering steps to speed audio processing while preserving accuracy in game scenes.
Real-time voice emotion recognition replaces delayed facial cues to assign more accurate content emotion ranks during events.
A parametric distance-dependent gain model corrects unnatural level and frequency changes when rendering volumetric audio sources in XR.
Combining harmonic and mixing patches fills high-frequency spectral gaps while preserving the original audio envelope and quality.
Correlating fingerprint and voice traits with family-member biometrics helps identify subjects missing from the database with higher confidence.
Combining filtered and original audio analysis signals with a maximum function preserves narrow peak magnitude while detecting broad peaks accurately.
Audio signature matching identifies botnet calls and routes them to a separate queue, protecting PSAP operators from DDoS disruption.
Low-latency time-domain separation adjusts target and noise gains to meet a hearing-based SNR while preserving environmental cues.
Embeds watermarks into selected playback frames in real time, enabling terminal tracing while limiting audio quality loss.
Padded transient blocks prevent temporal aliasing and waveform dispersion during phase-based bandwidth extension while limiting extra computation.
Acoustic voiceprints replace manual speaker separation, improving diarization accuracy and transcription speed in live audio interactions.
Dynamic truncation in a pre-configured GAN decoder improves speech enhancement by balancing audio quality and output variety.
Weighted blending of monaural and stereo decoded signals improves stereo output quality without requiring complex joint decoding.
A structured latent space separates clean and noise features to compress audio waveforms with lower memory and bandwidth use.
Weighted scoring of articulatory events helps speech analysis distinguish disease-indicative units and improve physiological state evaluation.
Using encoding bandwidth instead of input bandwidth improves perceptual entropy estimates and stabilizes audio frame bit allocation.
A unidirectional ML audio pipeline removes noise in real time on smartphones and IoT devices without future-frame access or added latency.
Embedded audio control instructions let lighting devices respond without Bluetooth or Wi-Fi pairing, reducing latency and sync issues.
Combining recurrent and nonrecurrent audio coding paths removes both long- and short-term redundancy for better compression and reconstruction.
Distance-aware loss weighting and task-ID embeddings help isolate a target speaker from single-channel noisy speech.
Dual patching combines harmonic and mixing patches to fill upper-band gaps and improve low-bitrate audio spectral quality.
Frequency band compression and cyclic feature mapping cut speech enhancement model complexity while preserving real-time communication quality.
Partitioned alpha-matrix processing computes critical HR filter sections with higher precision to cut rendering load while preserving spatial audio quality.
Standardized appearance, action, expression, voice, and interaction calibration makes digital humans more realistic, accurate, and responsive.
Virtual directional microphone signals let ICA separate more sound sources than physical microphones, cutting hardware cost.
Precomputed voice fingerprints and AI tagging identify speakers by time segment, enabling faster, more accurate content navigation.
A modular cloud platform combines emergency alerts, telehealth, and two-way messaging to improve reliable communication for deaf users.
Dynamic bitrate control preserves Bluetooth multipoint links while balancing bandwidth use and audio quality across connected devices.
Voiceprint vectors guide a neural network to separate target speech from interfering voices and ambient noise in complex audio.
Two Kalman filters separate singer voice, playback vocals, and music to suppress feedback, echo, and noise in hands-free karaoke.
Rate-distortion pulse selection in audio encoding balances signal error and computation to improve fixed codebook speech quality.
Selective caller-profile release gives PSAP staff only incident-relevant data, improving situational awareness while protecting privacy.
Two patching algorithms rebuild missing high-frequency content while preserving the spectral envelope for better low-bitrate audio quality.
A shared STFT spectral masking approach cleans speech and RF signals while cutting hardware footprint, cost, and latency.
Bit allocation shifts from initial coding to refinement for tonal frames, improving bit-use accuracy and audio quality.
User-triggered background sound softening reduces noise and music interference during video playback while preserving dialogue clarity.
By amending stereo width metadata in a compressed bitstream, the decoder widens playback while avoiding Fourier transforms and extra memory load.
Frame-level speech effectiveness feedback trains the network to suppress residual noise in non-speech segments without added computational load.
Edge ML on light-mounted microphones filters city audio locally and sends only gunshot spectrograms for accurate multilateration.
Adaptive downmixing concentrates sonic elements in a primary channel, reducing correlated multi-channel audio data while preserving decoding quality.
Helper networks screen bad biometric inputs and spoof attempts before model training, improving private authentication accuracy.
A single-channel downmix plus dry and wet upmix parameters reconstructs multichannel audio while reducing bandwidth and memory use.
Separate neural networks encode waveform and spectro-temporal audio components to improve quality, lower data rate, and reduce complexity.
By adjusting spatial audio metadata in the compressed bitstream, the decoder widens stereo imaging without Fourier-heavy processing.
Pseudo-label interpolation and unsupervised training enrich audio samples, improving separation accuracy on mismatched speech without manual labeling.
Graph-based scene representation links object nodes with interaction edges and audio cues to align audio-video data for recognition and anomaly tasks.
Frequency-oversampled harmonic transposition improves transient response and high-frequency resolution without window-switching artifacts.
Attribute and stream correspondence metadata lets receivers decode only needed channel or object audio, cutting processing load.
A neural network generates auxiliary audio from downmix and transient cues to improve multichannel upmixing with less smearing and distortion.