Custom Psychoacoustic Model Encoding for Hearing Profiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding technologies, such as MP3, use a generic 'golden ears' psychoacoustic model that fails to account for individual hearing profiles, leading to suboptimal listening experiences and inefficient lossy compression for hearing-impaired individuals, who require more mental effort to discern sounds in noisy environments due to increased masking and reduced perceptually relevant information.
Innovation Solution
The development of custom psychoacoustic models based on user-specific hearing profiles, derived from tests like suprathreshold and threshold tests, to optimize perceptual coding and digital signal processing, allowing for enhanced compression and improved listening experiences by removing irrelevant audio information and adjusting masking and hearing thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a generic 'golden ears' psychoacoustic model is used for audio encoding, then the encoding process is simple and file sizes are reduced, but the listening experience deteriorates for hearing-impaired individuals due to increased masking and reduced perceptually relevant information
Solution Approach 1:
The patent applies parameter changes by adapting the psychoacoustic model parameters based on individual hearing profiles. Instead of using fixed generic parameters, the system modifies masking thresholds, critical band widths, and other psychoacoustic parameters to match the specific hearing characteristics of each user, thereby improving listening experience without significantly complicating the encoding process
Solution Approach 2:
The patent implements dynamics by making the psychoacoustic model adaptive rather than static. The encoding system dynamically adjusts its parameters based on the user's hearing profile, allowing the model to respond to individual hearing impairments. This dynamic adaptation resolves the contradiction by maintaining encoding efficiency while improving reliability for hearing-impaired users
2Reliability
If a custom psychoacoustic model based on individual hearing profiles is used, then the listening experience improves for hearing-impaired individuals, but the encoding process becomes more complex and computationally intensive
Solution Approach 1:
The patent applies preliminary action by obtaining and storing user hearing profiles in advance before the actual audio encoding process. Hearing tests and profile measurements are performed beforehand, and the resulting data is saved for reuse during encoding. This eliminates the need to perform complex hearing assessments during each encoding operation, thereby reducing real-time computational complexity while maintaining improved listening experience
Solution Approach 2:
The patent uses copying by creating a digital representation of the user's hearing profile that can be reused across multiple encoding operations. Instead of repeatedly performing complex hearing tests, the system copies and applies the stored hearing profile data to multiple audio files, significantly reducing the computational burden of custom psychoacoustic modeling while maintaining reliability for hearing-impaired individuals
3Loss of substance
If perceptually irrelevant information is discarded to reduce data rate, then file size is reduced, but the amount of perceptually relevant information for hearing-impaired individuals is further reduced due to their higher masking thresholds
Solution Approach 1:
The patent applies parameter changes by adjusting the masking threshold parameters used to determine which audio information can be discarded. For hearing-impaired users, the system uses elevated masking thresholds that reflect their reduced auditory sensitivity, allowing more information to be retained in frequency regions where normal-hearing models would discard it. This resolves the contradiction by reducing file size according to actual user needs rather than generic standards
Solution Approach 2:
The patent implements local quality by applying different information retention strategies to different frequency regions based on the user's hearing profile. Instead of uniformly discarding information across all frequencies, the system selectively retains more information in frequency bands where the hearing-impaired user has elevated thresholds and reduced sensitivity, while still compressing regions where their hearing is relatively preserved. This local differentiation reduces file size while preserving perceptually relevant information
Data Source
Figure 1A~1B
Figure 2
Figure 3
AI summary
Systems and methods are provided for modifying an audio signal using custom psychoacoustic methods, for encoding the audio signal. A user's hearing profile is first obtained. Subsequently, a sample of the audio signal is split into frequency components. Next, masking and hearing thresholds are obtained from the user's hearing profile and applied to the frequency components of the audio sample, wherein the user's perceived data is calculated. User's imperceptible audio signal data is then disregarded. The audio sample is quantized and the resulting transformed audio sample encoded.