Statistical Audiogram Processing
By integrating user hearing threshold data with population statistics and Bayesian estimation, the method improves the accuracy of audiogram estimation on mobile devices, addressing issues of unsupervised testing and device variability.
Patent Information
- Application Number
- JP2025505720
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2023-07-28
- Publication Date
- 2025-08-26
AI Technical Summary
Existing methods for assessing hearing loss on mobile devices are prone to errors due to lack of supervision, environmental noise, and unknown headphone frequency responses, leading to inaccurate audiogram measurements.
A method that combines user hearing threshold data with statistical data from a population, including calibration and noise data, using Bayesian maximum a posteriori (MAP) estimation to improve audiogram estimation accuracy.
Enhances the accuracy of audiogram estimation by correlating measurements across frequencies, ears, and devices, reducing errors associated with unsupervised testing and unknown device characteristics.
Smart Images

Figure 2025528067000001_ABST
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 394,017, filed August 1, 2022, and U.S. Provisional Application No. 63 / 438,669, filed January 12, 2023, which are incorporated herein by reference in their entireties.
[0002] [Technical field] SUMMARY The present disclosure relates to techniques for estimating an audiogram for a user of a media playback device, and techniques for estimating calibration data for a media playback device.
[0003] Although some embodiments are described herein with particular reference to that disclosure, it should be understood that the disclosure is not limited to such applications but is applicable in a broader context. [Background technology]
[0004] Any discussion of background art throughout this disclosure should not be taken as an admission that such art is widely known or forms common general knowledge in the art.
[0005] Hearing loss is a common problem caused by noise exposure, disease, and aging. People with hearing loss may find it difficult to converse. They may also have difficulty understanding dialogue in media content, fully enjoying music, and effectively interacting with systems that include user interfaces and feedback mechanisms that rely on voice (e.g., voice assistants, voice command or control). There are several types of hearing loss, ranging from mild hearing loss, which results in decreased sensitivity to high-pitched sounds, to severe hearing loss, which results in the inability to hear anything except at very high sound pressure levels. Hearing loss develops gradually as people age. Because age-related hearing loss, also known as presbycusis, develops gradually, people with presbycusis may not realize that they have lost some of their ability to hear low sounds.
[0006] Hearing loss is traditionally assessed by an audiologist using carefully calibrated, specialized equipment. The assessment is performed in a soundproof environment. The audiologist determines, among other things, a so-called audiogram, which describes the amount of hearing loss (hearing loss) in decibels as a function of frequency. This process involves measuring hearing thresholds, expressed in sound pressure level (dB SPL), using test signals such as sine waves or band-limited noise signals. The audiogram, or specified hearing loss as a function of frequency, is then determined by calculating the difference between the measured hearing thresholds and those of a healthy ear. This process requires accurate calibration data so that the digital signal levels present in the sound generator or test app (e.g., levels specified for a digital full-scale signal, or dB re FS) can be converted to acoustic sound pressure levels reproduced by the headphones used during the assessment.
[0007] 1 illustrates an example of an audiogram estimation process by audiogram estimation system 100. The following hearing loss estimation process is repeated for a range of frequencies (e.g., at multiple discrete frequencies) and for both ears of a listener (e.g., the process requires user / listener input in response to multiple separate stimulus signals for each ear).
[0008] According to one example of this audiogram estimation process, an audio generator 110 generates a signal at a specific frequency that is played over headphones to a listener (user) 10. The listener 10 indicates (in the form of a subjective response 120) whether or not they can hear the signal and determines a threshold sound level 130 in the digital domain, expressed as dB re FS (Full Scale). This threshold 130 is then corrected (using calibration data 140) for the frequency response of the headphones and playback equipment, and a corresponding threshold 150 is calculated in dB SPL. Finally, the normal hearing level 160 at the specific frequency is subtracted to calculate the hearing loss (at the specific frequency) that goes into the audiogram 170.
[0009] The resulting audiogram can be applied to a hearing loss compensation system (e.g., a media playback device such as a mobile phone, television, set-top box, computer, etc.), such as a hearing aid or an application that applies hearing loss compensation to the playback of media content based on the audiogram data. An example of a system 200 is shown schematically in FIG. 2. The system 200 receives as input audio content 20 corresponding to an input audio signal 210 and calculates the signal level of the audio input 210 in the digital domain in a signal level calculation block (or module) 220. This process is typically performed in two or more subbands (not shown). Then, for each subband, a hearing loss compensation gain 250 is calculated in a hearing loss compensation calculation block (or module) 230 based on the user's audiogram data 270 and the digital signal level of the subband. In this step, device level (such as device playback level or volume setting) and frequency response calibration data (e.g., data characterizing acoustic output versus device level), i.e., device calibration data 270, are essential to ensure that the digital signal level can be converted to an associated acoustic sound pressure level before calculating the hearing loss compensation gain 250. In the final stage, in a hearing loss compensation gain application block (or module) 260, the hearing loss compensation gain 250 is applied to (a sub-band of) the audio signal 210 to generate a reproduced audio signal 280, which is sent to the headphones, earphones, or hearing aid transducer of the listener (user) 10.
[0010] Traditionally, the process of obtaining an audiogram and setting up hearing aids is performed by an audiologist. Hearing loss measurements are performed in a supervised manner under laboratory conditions and provide an accurate estimate of hearing loss. However, recently, mobile device applications have been introduced that attempt to measure and compensate for hearing loss. To obtain accurate measurements, these apps typically require the consumer (user) to wear limited headphones (e.g., headphones with a known frequency response and known calibration data) in a quiet environment while performing a hearing test. Similar to an audiologist assessment, the process used by mobile apps is repetitive (e.g., requiring user input in response to many separate stimulus signals for each ear), cumbersome (e.g., requiring significant concentration, prone to error, difficult to follow steps), and inefficient (e.g., consuming time and associated computing resources). The process of assessing hearing loss is often referred to as "onboarding" or "enrolment."
[0011] Measuring hearing loss in a consumer domain setting presents significant challenges, including: Lack of supervision: In contrast to a process performed by an audiologist, the onboarding process on mobile devices is typically unsupervised, increasing the risk of errors and reducing the accuracy of test results. · Environmental noise: Consumers may not have access to a quiet enough environment to accurately measure their audiogram, reducing the accuracy of the measured audiogram. · Unknown playback levels: With a large number of mobile devices, the exact conversion from the digital signal level in the app to the sound pressure level produced by the headphones can vary significantly, introducing further inaccuracies into the estimated audiogram. Unknown headphone frequency response: The conversion of digital signal levels to sound pressure levels in mobile devices depends on the frequency response of the headphones used during the measurement. Due to the vast number of headphone types and brands, the exact frequency response of headphones is often unknown at the time of onboarding.
[0012] As a result of these challenges, estimated hearing loss for a particular ear and frequency is prone to significant error or inaccuracy compared to assessments performed by audiologists.
[0013] Therefore, there is a need for improved techniques for hearing loss estimation. There is a particular need for such techniques that compensate for at least one of lack of supervision, environmental noise, unknown playback levels, and unknown headphone frequency response. There is a further need for techniques that allow for estimating device calibration data or headphone frequency response for use in subsequent hearing loss estimation, including for use in cloud-based or server-based settings. Summary of the Invention
[0014] In view of the above, the present disclosure provides a method for estimating an audiogram for a user of a media playback device, and a method for estimating calibration data for a media playback device, as well as corresponding apparatus, programs, and computer-readable storage media having the features of the respective independent claims.
[0015] According to a first aspect of the present disclosure, there is provided a method for estimating an audiogram for a user of a media playback device. The method may include acquiring user hearing threshold data for the user, where the user hearing threshold data may indicate one or more frequencies and a hearing threshold for one or both ears. The hearing threshold of the user hearing threshold data may be related to a hearing threshold measurement. The hearing measurement may be acquired by the media playback device. The method may further include acquiring sample hearing threshold data, where the sample hearing threshold data may indicate a hearing threshold of a personal sample set (e.g., a population). The sample hearing threshold data (statistical hearing threshold data, population hearing threshold data) may indicate a pre-stored hearing threshold of the personal sample set. The method may further include acquiring at least one of sample calibration data and sample noise data, where the sample calibration data may indicate a frequency response of the media playback device sample set (e.g., a population). Furthermore, the sample noise data may indicate a variation in playback level and / or a variation in user response in the process of measuring the hearing threshold. The method may further include determining an estimate of an audiogram for the user based on at least one of the user hearing threshold data, the sample hearing threshold data, and the sample calibration data and the sample noise data. The method may further include outputting the determined estimate of the audiogram and / or the set of compensation gains associated with the determined estimate of the audiogram, for example, to an application capable of playing audio on a media playback device or transmitting it to another device (e.g., a server or a cloud-based service).
[0016] By relying on statistical data on the hearing thresholds of a population of individuals, and also on statistical data on device calibration and / or noise, the proposed method can correlate measurements at different frequencies and / or in different ears, thereby improving the accuracy of the audiogram estimation, which can particularly address deficiencies in audiogram estimation due to unknown device data, lack of supervision, and non-ideal testing environments.
[0017] In some embodiments, determining the audiogram estimate may be further based on normal hearing data indicating expected hearing thresholds in the absence of hearing loss.
[0018] In some embodiments, determining the estimate of the audiogram may include applying relative weights to the user hearing threshold data and the sample hearing threshold data based on at least one of the sample calibration data and the sample noise data.
[0019] This allows the relative impact of pre-stored audiogram data to be measured based on the overall quality of the measured hearing thresholds, ensuring that high quality measurements are not "watered down" by statistical data while at the same time ensuring that the negative impact of low quality measurements is minimized.
[0020] In some embodiments, determining the audiogram estimate may be based on a Bayesian maximum a-posteriori (MAP) estimation technique.
[0021] The MAP estimation technique provides a reliable tool for using prior knowledge (eg, population audiogram and population calibration data, and expected noise) to improve the estimation of a user's audiogram.
[0022] In some embodiments, the user's hearing threshold data may indicate hearing thresholds at multiple frequencies for the left and right ears.
[0023] In some embodiments, the hearing threshold of the user's hearing threshold data may be represented by a digital signal level on a playback device.
[0024] In some embodiments, obtaining the user's hearing threshold data may include outputting, by the media playback device, multiple audio signals (audio test signals) of different frequencies (at different sound pressure levels, such as gradually increasing and decreasing sound pressure levels). The multiple audio signals may have different frequencies, and there may be one such signal for each frequency for each ear. The obtaining may further include receiving user input in response to the output audio signals. The obtaining may further include generating the user's hearing threshold data based on the received user input.
[0025] Thus, the method can obtain a user's subjective responses to an audio test signal and determine the user's hearing threshold based on these subjective responses.
[0026] In some embodiments, the sample hearing threshold data may indicate information about the audiogram for the individual sample set. Additionally or alternatively, the sample hearing threshold data may indicate the mean and covariance of the audiogram for the individual sample set. Additionally or alternatively, the sample hearing threshold data may indicate the mean and covariance of the hearing threshold for each frequency and ear for the individual sample set. For example, the sample hearing threshold data may indicate the mean and covariance of a hearing threshold vector for the individual sample set. Each entry in the sample hearing threshold vector may be associated with a given frequency and ear. Thus, the sample hearing threshold vector may be of dimension (2N F )×1, where 2N F (or K as defined below) is the number of frequencies.
[0027] In some embodiments, the sample calibration data may indicate the mean and covariance of frequency responses for a sample set of media playback devices. For example, the sample calibration data may indicate the mean and covariance of frequency response vectors for a sample set of media playback devices. Each entry in the frequency response may be associated with a given frequency and ear. For example, the frequency response vector may be of dimension (2N F )×1, where N F is the number of frequencies. The mean is the dimension (2N F )×1 vector, where each entry represents the average frequency response for each frequency-ear pair. Thus, the covariance can be expressed as, for example, a vector of dimension (2N F )×(2N F ) matrix.
[0028] In some embodiments, the audiogram estimate ŷ is given by:
number
[0029] In some embodiments, the personal sample set may be selected based on at least one user attribute of the user. For example, the personal sample set may be selected based on at least one user attribute of the user's age, gender, or location. The user attribute may be derived, for example, from user input or from information about the user from other sources.
[0030] Thereby, the sample hearing threshold data may be selected such that the user's actual hearing threshold is likely to be similar to the sample hearing threshold data.
[0031] In some embodiments, the step of determining an audiogram estimate may be performed at the media playback device or at a server device in communication with the media playback device.
[0032] In some embodiments, the method may further include receiving audio data for playback on a media playback device. The method may further include determining a set of compensation gains based on the determined estimate of the audiogram. The method may further include generating hearing-optimized audio data by applying the determined set of compensation gains to the audio data. The method may further include rendering the hearing-optimized audio data for playback.
[0033] In some embodiments, determining the set of compensation gains may be further based on the received audio data.
[0034] According to another aspect of the present disclosure, a method for estimating calibration data for a media playback device is provided. The calibration data can be indicative of a frequency response of the media playback device. The method can include obtaining first sample hearing threshold data. The first sample hearing threshold data can be indicative of a hearing threshold for a first personal sample set and associated with a given device type. The method can further include obtaining second sample hearing threshold data. The second sample hearing threshold data can be indicative of a hearing threshold for a second personal sample set different from the first personal sample set. The second sample hearing data need not be associated with a single given device type. For example, the second sample hearing data can include multiple subsets of sample hearing data, each subset associated with a respective device type. The method can further include determining an estimate of the calibration data based on the first sample hearing threshold data, the second sample hearing threshold data, and normal hearing data indicative of hearing thresholds expected in the absence of hearing loss.
[0035] This allows calibration data for a given device type (e.g., the type of media playback device the user is currently using) to be estimated, for example, for use in more accurate audiogram estimation or to be stored for later use.
[0036] In some embodiments, the method may further include obtaining user hearing threshold data for a user of the media playback device. The method may further include determining an estimate of the user's audiogram based on the user hearing threshold data, the estimate of the calibration data, and the normal hearing data.
[0037] In some embodiments, the user hearing threshold data, the first sample hearing threshold data, and the second sample hearing threshold data may each indicate hearing thresholds for one or more frequencies and one or both ears.
[0038] In some embodiments, the user's hearing threshold data may indicate hearing thresholds for multiple frequencies for the left and right ears. Similar comments may apply to the first and second sample hearing threshold data.
[0039] In some embodiments, the hearing threshold of the user's hearing threshold data may be represented by a digital signal level on a playback device.
[0040] In some embodiments, the second sample hearing threshold data may indicate information regarding an audiogram for the second individual sample set. Additionally or alternatively, the second sample hearing threshold data may indicate an average audiogram for the second individual sample set. Additionally or alternatively, the second sample hearing threshold data may indicate an average hearing threshold for each frequency and ear for the second individual sample set.
[0041] In some embodiments, the second personal sample set may be selected based on at least one user attribute of the user.
[0042] In some embodiments, the step of determining an audiogram estimate may be performed at the media playback device or at a server device in communication with the media playback device.
[0043] In some embodiments, the method further includes obtaining updated user hearing threshold data for the user for a second media playback device different from the media playback device, the updated user hearing threshold data indicating an updated hearing threshold for the given frequency. The method may further include determining an offset between the user hearing threshold at the given frequency indicated by the user hearing threshold data and the updated hearing threshold. The method may further include determining a second calibration data estimate for the second media playback device based on the calibration data estimate and the determined offset.
[0044] In some embodiments, the method further includes obtaining updated user hearing threshold data for the user for a second media playback device different from the media playback device, the updated user hearing threshold data indicating updated hearing thresholds for a plurality of given frequencies. The method may further include determining an estimate of an offset between the calibration data and second calibration data for the second media playback device based on the user hearing thresholds at the plurality of given frequencies indicated by the user hearing threshold data, the updated hearing thresholds, second sample hearing threshold data, and sample noise data. Here, the sample noise data may indicate variations in user response in the process of measuring the hearing threshold. The method may further include determining an estimate of the second calibration data based on the estimate of the calibration data and the determined estimate of the offset.
[0045] In some embodiments, determining the estimate of the second calibration data may be based on a Bayesian maximum a posteriori (MAP) estimation technique.
[0046] Aspects of the present disclosure may be implemented via a computing device. The device may include at least one processor and a memory coupled to the processor. The processor may be adapted to perform methods according to aspects and embodiments of the present disclosure. For example, the memory may store instructions that, when executed by the at least one processor, cause the computing device to perform methods according to aspects and embodiments of the present disclosure.
[0047] Aspects of the present disclosure may be implemented via a computer program. When the instructions of the computer program are executed by a processor (or computing device), the processor can perform aspects and embodiments of the present disclosure. A computer-readable storage medium may store the program. Such computer-readable storage media may include memory devices as described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, various novel aspects of the subject matter described in this disclosure may be implemented via one or more computer-readable storage media having software stored thereon.
[0048] It should be noted that the methods and systems, including the preferred embodiments, outlined in this disclosure can be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this disclosure can be combined in any manner. In particular, the features of the claims can be combined with each other in any manner.
[0049] It will be understood that apparatus features and method steps may be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding apparatus (or system), and vice versa, as will be understood by those skilled in the art. Furthermore, it will be understood that statements made above with respect to a method apply equally to a corresponding apparatus or system, and vice versa. [Brief explanation of the drawings]
[0050] Exemplary embodiments of the present disclosure are now described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a schematic diagram of an example audiogram estimation system. [Figure 2] 1 illustrates a schematic diagram of an example hearing loss compensation system. [Figure 3]1 illustrates a schematic diagram of an example of a hearing loss compensation system according to an embodiment of the present disclosure. [Figure 4] FIG. 1 shows an example of an average population headphone frequency response and an average population audiogram. [Figure 5] 1 is a flowchart illustrating an example method for estimating an audiogram of a user of a media playback device, according to an embodiment of the present disclosure. [Figure 6] 1 is a flowchart illustrating an example method for estimating calibration data for a media playback device, according to an embodiment of the present disclosure. [Figure 7] 1 illustrates a schematic diagram of an example of a system for determining an audiogram using a cloud-based database infrastructure, according to an embodiment of the present disclosure. [Figure 8] 10 is a flowchart illustrating another example method for estimating calibration data for a media playback device, according to an embodiment of the present disclosure. [Figure 9] 1A-1C illustrate example audiograms using different techniques, according to embodiments of the present disclosure. [Figure 10] 1A-1C illustrate example audiograms using different techniques, according to embodiments of the present disclosure. [Figure 11] 1A-1C illustrate example audiograms using different techniques, according to embodiments of the present disclosure. [Figure 12] 1A-1C illustrate example audiograms using different techniques, according to embodiments of the present disclosure. [Figure 13] 1 is a block diagram that schematically illustrates an example of a computing device for performing methods according to embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0051] In view of the above needs, and to solve the problems outlined above and significantly improve the accuracy of audiogram estimation in consumer settings, techniques and systems for estimating and / or applying audiograms for hearing loss compensation are described herein.
[0052] Broadly speaking, these techniques and systems combine threshold measurements with one or more additional probabilistic sources of information to reduce error in the estimated audiogram. These one or more additional sources of information may include, for example: Thresholds measured at frequencies other than the one at which they were measured. It has been found that both headphone calibration data and audiogram data are not independent across frequencies. This means that threshold measurements at, for example, a frequency of 500 Hz also provide probabilistic information or likelihood regarding thresholds at other frequencies. Therefore, in contrast to traditional methods that measure hearing loss at each frequency independently, probabilistic combination of threshold data across all frequencies tested can improve accuracy. Thresholds measured in one ear (e.g., the left ear) provide probabilistic information about the hearing loss in the other ear (e.g., the right ear). Combining threshold data from both ears can provide a more accurate estimate of the audiogram. Demographic data. Combining threshold data with statistical population information with similar demographic attributes (e.g., gender, age, and / or location information) can provide a more accurate estimate of the audiogram. Thresholds measured during separate measurement sessions. The user can repeat one or more measurements at different times or in different environments. Probabilistic combination of such multiple measurements can improve the accuracy of the resulting audiogram or estimated hearing loss. Thresholds measured with different devices (e.g., different mobile phones and / or different headphones or earphones). Probabilistic combination of measurements across multiple devices improves audiogram estimation by reducing bias due to specific, often unknown, device calibration characteristics. Thresholds obtained by other users using the same device. Even if calibration data for a particular device (mobile phone and / or earphones) is not available, measurement data from multiple users using the same device, preferably with similar demographic data, can improve the audiogram estimation process. Population information on the frequency response of headphones. Even if the headphones or earphones used during the audiogram acquisition are unknown, the statistical properties of the frequency response of a population of headphones or earphones can provide information on how the frequency response varies on average and correlates with frequency. This can help improve the accuracy of audiogram estimation. Population audiogram information. The availability of audiogram data for a consumer population can provide information on how audiograms correlate across frequency and between ears, allowing for more accurate estimation of audiogram data.
[0053] Thus, rather than using hearing loss estimates for each frequency and ear independently, the techniques proposed by this disclosure broadly use at least one of multiple measurements at different frequencies and / or different ears, headphones and / or devices, and / or users, and combine that information with overall population characteristics of audiograms and / or headphone frequency responses to calculate a more accurate estimate of hearing loss at each frequency and each ear.
[0054] 3 illustrates schematically an example of an audiogram estimation system 300 for hearing loss estimation. The system 300 may be suitable for estimating an audiogram for a user of a media playback device, for example in a consumer setting.
[0055] In the above-described system 100, an audio generator 310 generates a signal 315 of a particular frequency that is played over headphones to a listener (user) 10. The listener 10 indicates (in the form of a subjective response 320) whether or not the signal is audible and determines a threshold audio level 330 in the digital domain, expressed as dB re FS (Full Scale). This is done for one or more frequencies (e.g., for multiple frequencies) and one or more ears (e.g., left and right ears). It will be understood that multiple threshold audio levels are determined. Preferably, threshold audio levels are determined for multiple frequencies and for the left and right ears. The resulting audio levels are represented by or included in user hearing threshold data.
[0056] An index of hearing loss 360 for listener 10 is determined based on user hearing threshold data by statistical optimization block (or module) 350. This process can use normal hearing data that indicates expected hearing thresholds in the absence of hearing loss. The index of hearing loss 360 can be stored or compiled in the form of an audiogram 370 for listener 10.
[0057] According to the techniques of the present disclosure, the statistical optimization block 350 (probabilistically) combines the user hearing threshold data with statistical information 340 (e.g., sample hearing threshold data, and / or sample calibration data and sample noise data, as described in more detail below) to improve the accuracy of the estimated audiogram 370.
[0058] The statistical information 340 can be stored locally (e.g., on a mobile device, on a playback device, etc.) or remotely (e.g., in a distributed system, a cloud service, etc.). Similarly, the optimization process can be performed locally or as a cloud-based service. Furthermore, user-generated threshold data can contribute to statistical information that is collected (e.g., locally or remotely) to improve the process of estimating an audiogram.
[0059] Exemplary Embodiments Combining Measurements Across Frequency and / or Binaurally This starts from a set of measurements of hearing thresholds (e.g. formed by or included in user hearing threshold data), e.g. expressed as digital signal levels (dB re FS)x, including (e.g. composed of) threshold measurements for two ears {l, r} and K frequencies, e.g. given by:
number
[0060] It is understood that, in general, a measurement set includes at least two measurements for different frequencies and / or different ears. Without intended limitation, reference may be made throughout this disclosure to measurements for multiple frequencies and for both ears.
[0061] While the mobile devices and headphones (which may collectively be referred to as playback devices) used during measurements are unknown, it is assumed that statistical information describing the population of headphones and mobile devices typically used by consumers is accessible. More specifically, the headphone frequency response population information (e.g., in the form of sample calibration data as described in more detail below) includes a mean frequency response vector μ that represents (e.g., contains) the average of the overall frequency response h for K frequencies across the two ear cups. h → , and the frequency response covariance matrix Σ using the expectation value operator <.> h For example, the headphone frequency response population information may be given by:
number
[0062] Similarly, it is assumed that mean and covariance matrices are available for the audiograms of a population of users (e.g., in the form of sample hearing threshold data, as described in more detail below). Ideally, the mean and covariance matrices are adjusted for at least one user attribute, such as the user's age and / or gender. The audiogram mean and covariance matrix are respectively denoted by μ a → and Σ a It is expressed as:
[0063] As an example, μ as a function of frequency for the age group 40-60 years old h → (average population headphone frequency response) and μ a → The values of (average population audiogram) are shown in Figure 4. Graph 410 shows μ h → The graph 420 represents the value of μ a → Represents the value of
[0064] Finally, the playback level may vary from one mobile device (or playback device) to another, with a variance σ p 2 Furthermore, we assume that the response variance when measurements are repeated due to subjective rating fluctuations can be described by a zero-mean random variable with variance σ r 2 These variables can be represented by or included in sample noise data, for example, as described in more detail below.
[0065] The absolute threshold in quiet for normal hearing (i.e., no hearing loss) in dB SPL is μ n → It is shown as follows.
[0066] Throughout this disclosure, the term "calibration data" (e.g., sample calibration data) can be used to describe the transfer from digital signal level to acoustic sound pressure level, including both the sensitivity and frequency response of the headphones and the effect of the mobile device that generates the electrical signal. As noted above, the combination of a mobile device (e.g., a mobile phone) and headphones can be referred to as a playback device.
[0067] Given the above-mentioned prior population information (e.g., sample hearing threshold data), a listener's audiogram can be estimated from noise measurement data (e.g., user hearing threshold data) acquired together with unknown device calibration data (including frequency response data).
[0068] In particular, the user's audiogram → The Bayesian maximum a posteriori (MAP) estimate ŷ for y is given, for example, by:
number
number
[0069] As exemplified by equation (5), when determining the audiogram estimate ŷ, the techniques disclosed herein use the frequency response covariance matrix Σ h (e.g., contained in sample calibration data) and / or a zero-mean random variable σ r 2 and / or σ p 2 (e.g., including sample noise data) based on user hearing threshold data x → and the mean vector μ a → (e.g., contained in sample hearing threshold data)
[0070] Assuming the covariance matrix has non-zero off-diagonal elements, an estimate of the audiogram at frequency k does not depend only on the measurement at frequency k, but on all measurements across frequencies and across both ears. Thus, the techniques disclosed herein (probabilistically) correlate user hearing thresholds at different frequencies and / or different ears to determine an estimate of the listener's audiogram. This correlation can be based on sample hearing threshold data, and at least one of sample calibration data and sample noise data.
[0071] 5 illustrates an example method 500 for estimating an audiogram of a user of a media playback device in accordance with the above discussion. Method 500 includes steps S510-S540 that may be performed on the media playback device or on a server device in communication with the media playback device.
[0072] In step S510, user hearing threshold data for the user is obtained (e.g., determined, measured), which can be done as described above with reference to Figures 1 and 3.
[0073] In particular, the user hearing threshold data may indicate hearing thresholds for one or more frequencies and one or both ears. Preferably, the user hearing threshold data indicates hearing thresholds for multiple frequencies for the left and right ears. For example, the user hearing threshold data may indicate a user hearing threshold vector (i.e., a vector of user hearing thresholds, such as the vectors defined above). Each entry in the user hearing threshold vector may be associated with a given frequency and ear. Thus, the user hearing threshold vector may be of dimension (2N F )×1, where N F =K is the number of frequencies.
[0074] As described above, the hearing threshold of the user hearing threshold data can be expressed in digital signal levels, for example, in digital signal levels of a playback device.
[0075] The hearing threshold represented by or included in the user hearing threshold data may be related to a hearing threshold measurement, as described above. The hearing measurement may be obtained by the media playback device. For example, the hearing threshold or measurement may be obtained by performing the following steps A to C, which can be understood as examples of sub-steps of step S510.
[0076] Step A: Outputting multiple audio signals of different frequencies by the media playback device. For example, the multiple audio signals may have different frequencies, and for each frequency there may be one such signal for each ear. Furthermore, for each frequency and ear, audio signals of different levels may be output to determine the sound pressure level threshold at which the audio signal is audible to the user.
[0077] Step B: Receive user input in response to the output audio signals, which may relate to a user's subjective response as to whether the user can hear each audio output audio signal.
[0078] Step C: Generate user hearing threshold data based on the received user input.
[0079] In step S520, sample hearing threshold data is obtained. The sample hearing threshold data may be statistical data indicative of the hearing thresholds of the individual sample set. Thus, the sample hearing threshold data may be indicative of or related to pre-stored hearing thresholds of the individual sample set.
[0080] Additionally, the sample hearing threshold data may indicate information about the audiogram for the individual sample set. Specifically, the sample hearing threshold data may indicate information about the audiogram (e.g., Σ as defined above) for the individual sample set. a ) (e.g., the above μ a →) and variance (e.g., covariance) of the individual hearing threshold vector (i.e., a vector of hearing thresholds) for the individual sample set. Each entry in the sample hearing threshold vector can be associated with a given frequency and ear. Thus, the sample hearing threshold data can indicate the mean and variance (e.g., covariance) of the hearing threshold vector (i.e., a vector of hearing thresholds) for the individual sample set. Each entry in the sample hearing threshold vector can be associated with a given frequency and ear. Thus, the sample hearing threshold vector can be of dimension (2N F )×1, where N F =K is the number of frequencies.
[0081] It will be further appreciated that the personal sample set may be selected (e.g., compiled) based on at least one user attribute of the user. For example, the personal sample set may be selected based on at least one user attribute of the user's age, gender, or location. The user attribute may be derived from user input, e.g., other user-provided data, or user-related data obtained from an external source.
[0082] In step S530, at least one of sample calibration data and sample noise data is obtained.
[0083] Here, the sample calibration data may be statistical data describing the frequency response of a sample set of media playback devices. These frequency responses may be given for the entire playback device, including headphones. In other words, the frequency responses may link digital signal levels to sound pressure levels for multiple frequencies.
[0084] In some embodiments, the media playback device sample set may be selected (e.g., compiled) based on device attributes, device ID, etc. of the media playback device. For example, the media playback device sample set may be selected according to the headphone type (e.g., open, closed, earbud, cable-based, wireless) of headphones used by the user. Further, the media playback device sample set may be selected according to, for example, device type or ID or device brand. In general, any known characteristic or attribute of a media playback device may be used to select the media playback device sample set.
[0085] In some embodiments, the sample calibration data may indicate the mean and variance (e.g., covariance) of frequency responses for a media playback device sample set. Specifically, the sample calibration data may indicate the mean and variance (e.g., covariance) of a frequency response vector (i.e., a vector of frequency responses) for a media playback device sample set. Each entry of the frequency response may be associated with a given frequency and ear. For example, the frequency response vector may be of dimension (2N F )×1, where N F =K is the number of frequencies. h → ) is a function of dimension (2N F ) × 1 vector, where each entry represents the average frequency response for each frequency-ear pair. h ) is a function of dimension (2N F )×(2N F ) matrix.
[0086] Furthermore, the sample noise data may be used to account for variations in the playback level (e.g., σ as defined above) in the process of repeatedly measuring the hearing threshold. p 2 ) and / or the variability of the user response (e.g., σ defined above). r 2 ) may be statistical data showing the
[0087] The media playback device sample set may include a number of different media playback devices that a user may use to estimate their audiogram in a consumer environment, it being understood that the different media playback devices may relate to different combinations of media players, mobile phones, PDAs, handhelds, etc., headphones, earphones, etc.
[0088] In step S540, an estimate of the user's audiogram is determined (e.g., calculated) based on at least one of the user hearing threshold data, the sample hearing threshold data, and the sample calibration data and the sample noise data. Determining the estimate of the audiogram may further involve using normal hearing data (e.g., μ in Equation (4)) indicative of the hearing thresholds expected in the absence of hearing loss. n → ) For example, the predicted hearing threshold may be subtracted from the measured threshold of the user hearing threshold data to determine an estimate of the user's hearing loss.
[0089] Furthermore, as described above, determining an audiogram estimate in step S540 may include applying relative weights to the user hearing threshold data and the sample hearing threshold data based on at least one of the sample calibration data and the sample noise data. The relative weights may be adjusted to account for uncertainty in better measurement conditions (e.g., uncertainty in device calibration (e.g., Σ in Equation (5)). h ) is small) and / or the influence of noise (e.g., σ in Eq. (4) r 2 and / or σ p 2 )), the audiogram estimate is closer to an audiogram calculated directly from the measured user hearing thresholds (without frequency response and other calibration data (e.g., μ in Eq. (4)). h → )), and in worse measurement conditions, the population-averaged audiogram (e.g., μ in Eq. (4)) a → ) can be determined to be close to
[0090] Additionally, step S540 may (probabilistically) correlate user hearing thresholds at different frequencies and / or different ears (preferably at different frequencies and for both ears) to determine an estimate of the listener's audiogram. This correlation may be based on sample hearing threshold data, and at least one of sample calibration data and sample noise data.
[0091] For example, the determination of the audiogram estimate in step S540 may be based on a Bayesian maximum a posteriori (MAP) estimation technique that introduces the aforementioned correlation between user hearing thresholds at different frequencies and / or different ears. Referring to the example of equations (4) and (5), this correlation is expressed by the non-trivial matrix Σ h and / or σ p 2 It is introduced by the off-diagonal elements of the estimation matrix M present for J.
[0092] When using the MAP estimation technique, the audiogram estimate ŷ is given by equation (4), i.e.:
number
[0093] Although not shown in FIG. 5, the method 500 may further include outputting the determined estimate of the audiogram and / or the set of compensation gains associated with the determined estimate of the audiogram to, for example, an application capable of playing audio on a media playback device or transmitting it to another device.
[0094] Table 1 shows example results for the performance of audiogram estimation in dB (i.e., root mean square error) using four different methods. "No measurement" assumes no hearing loss (e.g., y^=0 → ). "Population average" uses the population average as the audiogram (e.g., y^=μ a → ). "Measurement" uses the measurement data directly as an audiogram (e.g., y^=x → -μ n → ). "Bayesian MAP" estimation (e.g., y^=M(x → -μ n → +μ h → -μ a → )+μ a → ).
[0095] The numbers in Table 1 represent the average error in dB, with lower values being better. Thus, the Bayesian MAP estimator provides a more accurate estimate of the audiogram than any of the other methods included. [Table 1] Performance of various audiogram estimation techniques [Table 1]
[0096] These examples assume that there is no prior information about the type of headphones used during the measurements. If there is information about the type of headphones, μ h → and Σ h By adjusting the value of , accuracy can be improved. For example, different statistical population values can be used for circum-aural and supra-aural headphones, earphones, etc.
[0097] Without loss of generality, it should be noted that the frequency index k can refer to different measurement frequencies, but equally well to repeated measurements at the same frequency, or a mixture of both. The Bayesian MAP estimator works regardless of which frequency is associated with index k, as long as all mean and covariance matrices are properly determined for the frequencies associated with measurement index k (i.e., as long as a consistent assignment of vector elements to frequency-ear pairs is used).
[0098] Exemplary Embodiments: Using an Estimated Audiogram for Hearing Optimization This embodiment can be easily combined with the above embodiments and implementations.
[0099] The user's estimated audiogram can be used to enhance audio reproduction for the user by compensating for individual hearing loss. This can be done in the same manner as described above in connection with Figure 2. For example, a method for enhancing audio reproduction, which can be performed in connection with method 500 described above or as a standalone method, can include the following steps A through D:
[0100] Step A: Receive audio data for playback on a user's media playback device.
[0101] Step B: Determine (e.g., calculate) a set of compensation gains based on the determined estimate of the audiogram. Determining the set of compensation gains can be further based on the received audio data. For example, the compensation gains can be applied in a multi-band compression framework, where the gains depend on the audiogram and the audio signal level of each frequency band.
[0102] Step C: Applying the determined set of compensation gains to the audio data (or an audio signal derived therefrom) produces perceptually optimized audio data.
[0103] Step D: Render the hearing-optimized audio data for playback to the user.
[0104] Exemplary Embodiments: Use of Demographic Statistics Estimating instrument calibration data This embodiment can be easily combined with the above embodiments and implementations.
[0105] Through a process where a large group of users determine their hearing thresholds, the following digital levels x are generated for known device IDs, denoted by D: D → It is assumed that a database of threshold vectors is generated, denoted by:
number
[0106] The statistical signal model x of these measurement thresholds D → is the unknown instrument frequency response x of instrument D h,D → , noise or variability in subjective responses n r → , normal hearing threshold μ n → , and the estimated actual audiogram y → The combination of is given as:
number
[0107] The system determines whether there are a sufficiently large number of thresholds x among users for a given device D. D → and assuming that the acquisition noise, audiogram, and calibration data are independent, the above equation can be rewritten in terms of expectation using the expectation value operator <.> as follows:
number
[0108] Therefore, the estimated instrument calibration data x h,D ^ can be obtained by the following formula:
number
[0109] In other words, the instrument calibration data x h,D → is the population average audiogram μ a → and normal hearing threshold μ n → , and all measurements x performed by many users on the same device D. D → This can be estimated by evaluating and calculating the average of the instrument calibration data x h,D → can be used in estimating the audiogram of a user (or any user) using a media playback device of type device D.
[0110] Therefore, the measured threshold x for a given user D → The audiogram estimate ŷ based on is given, for example, by:
number
[0111] 6 illustrates an example method 600 for estimating calibration data for a media playback device in accordance with the above discussion. The calibration data may indicate a frequency response of the media playback device. Method 600 includes steps S610-S630 that may be performed on the media playback device or on a server device in communication with the media playback device.
[0112] In step S610, first sample hearing threshold data is obtained (e.g., retrieved from a database), where the first sample hearing threshold data may represent the hearing thresholds of a first personal sample set and is associated with a given device type. Thus, in one embodiment, the first sample hearing threshold data may represent, for example, the inter-user threshold x for a given device D as defined above. D → a (large) number (e.g., a set) of, or its expected value <x D → > may also correspond to
[0113] In step S620, second sample hearing threshold data is obtained (e.g., retrieved from a database). The second sample hearing threshold data may be indicative of the hearing threshold of a second individual sample set that is different from the first individual sample set. The second sample hearing threshold data need not be associated with a single given device type. For example, the second sample hearing data may include multiple subsets of sample hearing data, each subset associated with a respective device type. For example, the second sample hearing threshold data may be indicative of information regarding an audiogram for the second individual sample set. Furthermore, the second sample hearing threshold data may be indicative of an average audiogram of the second individual sample set, or in other words, an average hearing threshold for each frequency and ear of the second individual sample set. Thus, in one embodiment, the second sample hearing threshold data may be indicative of, for example, a population mean audiogram μ as defined above. a → It may correspond to.
[0114] Again, the second personal sample set may be selected based on at least one user attribute of the user, which may be done, for example, similar to the selection described above in connection with step S520 of method 500.
[0115] In step S630, the calibration data is estimated (e.g., x h,D ^) is the first sample hearing threshold data (e.g., <x D → >), second sample hearing threshold data (e.g., μ a → ), and normal hearing data (e.g., μ n → ) is determined based on
[0116] In the above, similar to the description above in connection with method 500, the user hearing threshold data, the first sample hearing threshold data, and the second sample hearing threshold data may each represent a hearing threshold for one or more frequencies and one or both ears. Preferably, the user hearing threshold data represents hearing thresholds for multiple frequencies for the left and right ears. In this case, the other data defined above also relate to multiple frequencies and the left and right ears.
[0117] Estimating calibration data for media playback devices (e.g., x h,D ^), the method comprises obtaining user hearing threshold data (e.g., x^) of the user of the media playback device. D →) (e.g., expressed in digital signal levels of a playback device), and determining an audiogram estimate (e.g., ŷ) for the user based on the user hearing threshold data, the calibration data estimate, and the normal hearing data (not shown in the figures), where the former of these steps may proceed similarly to step S510 of method 500 described above, and the latter of these steps may proceed similarly to step S540 of method 500. In some embodiments, determining an audiogram estimate (e.g., ŷ) for the user may use equation (10) above, or equations (13) and (14) defined below.
[0118] Further, as mentioned above, the step of determining an audiogram estimate may be performed at the media playback device or at a server device in communication with the media playback device.
[0119] Cloud-based services and iterative improvements for hearing loss compensation The above-described embodiments can be implemented in a cloud-based manner.
[0120] An exemplary system 700 for determining an audiogram using a cloud-based database infrastructure is shown generally in FIG. 7. A user 710 determines a digital threshold 730, for example, using a mobile app 710 (e.g., running on a playback device). This can be done in the same manner as described above in the context of FIG. 3, based on the user's subjective response 720 to an audio signal 715 generated by the playback device. The mobile app 710 transmits the threshold 730 and corresponding device and / or user information 780 (e.g., device ID, user ID, age, and / or gender) to a central cloud-based server 790 that collects data from multiple users. The cloud-based service can then estimate device characteristics 740 of device D and the user's audiogram 770, as described above.
[0121] In particular, this architecture / method can be combined with the method described in the embodiment of method 500 (combining measurements across frequencies and / or across ears). For example, if the service has not yet collected many audiograms, or if a new, unknown device is being used by the user, the system can apply the method described in the previous embodiment to calculate the initial audiogram y0̂ for this particular user via the following formula:
number
[0122] Over time, if enough users take audiograms using the same device D, the system will generate a population-averaged frequency response μ h → Instead of having to rely on the estimated instrument frequency response x h,D You can update the user's audiogram using:
number
[0123] Each time a more accurate estimate of the audiogram data or device characteristics is determined by the system, the resulting improved audiogram and device characteristics / calibration data can be sent from the cloud-based service to the mobile device (playback device) or to a server in communication with the mobile device (playback device), so that the more accurate data can be used to improve the hearing loss compensation algorithm.
[0124] Clustering of equipment data The process of estimating instrument calibration data can be performed using two or more instruments D as long as there is confidence that the two or more instruments have similar characteristics. h,DThis can be further improved by calculating ^. Clustering can be applied to devices with similar characteristics, for example using k-means or multivariate Gaussian mixture models, or other clustering techniques based on the similarity of the measurement data obtained for the devices. Additionally, headphone type (earphones, over-ear, or on-ear headphones) can be used to improve the clustering process.
[0125] Exemplary Embodiment: User switches to different headphones In this example, a user performs an audiogram registration and measurement with headphones D1 (or, in general, device configuration D1) and obtains threshold data x D1 → Consider a use case where a user has obtained a calibration data set D2 and is now using the same process with a different pair of headphones D2 (or, in general, a device configuration D2). Thus, in this case, the calibration data associated with the device has not changed, but the calibration data associated with the headphones may have changed. Therefore, in order for the user to obtain an optimal hearing loss compensation process for the headphones D2, new calibration data needs to be obtained. One way is to use the method described in the previous example (Example: Use of Demographics), where a cloud-based system can provide calibration data for the new headphones D2.
[0126] Alternatively, a user can perform a reduced test at a limited number of frequencies, where new thresholds are evaluated and new calibration data is derived from them. For example, the thresholds can be measured at only one frequency using headphone D2. For headphones D1 and D2, the measured thresholds for a particular user l and frequency k can be calculated as x D1,l,k , x D2,l,k and assuming that the device calibration data for headphones D1 and D2 differ by a scalar d in the decibel domain, the calibration difference scalar can be expressed as x D1,l,k and x D2,l,k can be estimated from the difference in thresholds between:
number
[0127] If multiple measurements are made for headphone D2, a Bayesian MAP estimator can be used to calculate an estimate of the scalar d, for example via the following equation:
number
[0128] Therefore, this method uses the covariance matrix Σ a and subjective measurement (reproducibility) noise σ r 2 When measuring the threshold of headphone D2, which is represented by
[0129] 8 illustrates an example method 800 for estimating calibration data for a media playback device, consistent with the discussion above. The calibration data may indicate the frequency response of the media playback device (including particular headphones). Method 800 includes steps S810-S830, which may be performed on the media playback device or on a server device in communication with the media playback device. Steps S810-S830 of method 800 may be performed after steps of method 600, or may be performed as a stand-alone method.
[0130] In step S810, updated user hearing threshold data for a user is obtained for a second media playback device (e.g., a media playback device with headphones D2) that is different from the media playback device (e.g., a media playback device with headphones D1). The updated user hearing threshold data includes updated hearing thresholds for a given frequency (e.g., x D2,l,k ) may be indicated.
[0131] In step S820, the user hearing threshold (e.g., x) at a given frequency indicated by the user hearing threshold data is calculated. D1,l,k ) and updated hearing thresholds (e.g., x D2,l,k) is determined (eg, a scalar offset d).
[0132] In step S830, the calibration data is estimated (e.g., x h,D Based on ̂) and the determined offset (e.g., scalar offset d), an estimate of second calibration data for the second media playback device is determined, e.g., by adding the determined offset to the estimate of the calibration data.
[0133] Further, the updated user hearing threshold data at step S810 may indicate updated hearing thresholds for a plurality of given frequencies. Next, at step S820, an estimate of an offset between the calibration data and second calibration data of the second media playback device may be determined based on the user hearing thresholds at the plurality of given frequencies indicated by the user hearing threshold data, the updated hearing thresholds, second sample hearing threshold data, and sample noise data. The sample noise data may indicate variations in user response in the process of measuring the hearing threshold. Finally, at step S830, an estimate of second calibration data may again be determined based on the estimate of the calibration data and the determined offset estimate. The determination of the estimate of the second calibration data may be based on a Bayesian MAP estimation technique, as described above.
[0134] Transform Although the techniques described throughout this disclosure use least mean squares and Bayesian estimation techniques to improve audiogram accuracy and derive device calibration data, other methods can be used as well, including, but not limited to, fuzzy logic, Dempster-Shafer theory, imprecision probability, machine learning, k-means clustering, deep learning, neural networks, etc.
[0135] It is understood that the transmission, storage, and use of personal data such as audibility thresholds, device ID, demographic data, etc. can be or should be appropriately protected by dedicated privacy and security mechanisms.
[0136] The Bayesian estimation methods described throughout this disclosure are based on Gaussian distributions and lead to closed-form solutions. Note that for more complex distributions and data with unknown noise hyperparameters, various iterative algorithms exist for Bayesian MAP estimation in such conditions. Details are provided, for example, in Reference [1].
[0137] Further explanation: Calibration inference for hearing optimization problem Part of a hearing optimization solution is an onboarding step that measures hearing levels (HL) at different frequencies (e.g., 11 or 17 data points in some embodiments) to generate an audiogram. Traditional audiogram testing requires calibrated equipment to reproduce test tones at specific levels. Therefore, to reproduce the test on a device such as a mobile phone, the response characteristics of both the earphones and the mobile phone must be known (i.e., device-specific calibration data must be known). In the context of this disclosure, "calibration" can refer to (1) the sound pressure level (SPL) that the mobile phone and earphones (e.g., collectively referred to as playback devices) generate for a given signal level (e.g., specified at the OS / application level) and (2) the frequency response of the mobile phone and earphones. Because it is not feasible to measure every mobile phone and earphone on the market, the solution must be calibration-free.
[0138] Notably, the audiogram does not include any aspect of device calibration and only measures the user's hearing level. Test results that include device calibration are referred to in this disclosure as "audiogram + calibration."
[0139] Solution Typically, as soon as the mobile phone or earphones are changed to a different model, a full audiogram + calibration test must be repeated as follows: However, the techniques outlined in this disclosure are directed to an optimization process that can be performed without calibration, as long as the equipment used is the same as the one on which the test was performed.
[0140] FIG. 9 shows an example comparison of an audiogram for a known device (graph 910) and an audiogram plus calibration for a new device (graph 920) as a function of frequency.
[0141] The following solutions S1 to S4 may be related to the above embodiment and its implementation, as will be understood by those skilled in the art.
[0142] S1. Solution for SPL Calibration: Instead of performing a full "audiogram + calibration" test, a reduced test can be performed when a change to an unknown device is detected. By testing at a single frequency that was tested with a known device, a calibration offset between the known device and the new device can be generated. This allows the previous "audiogram + calibration" to be used for the new device. Information identifying the device, such as Bluetooth device ID, can be used to identify the device to which the updated test applies.
[0143] FIG. 10 shows an example comparison of an audiogram for a known device (graph 1010) and an audiogram + calibration for a new device with SPL recalibration (graph 1020) as a function of frequency.
[0144] S2. Solution for Separating SPL Calibration and Audiogram: The method in S1 does not allow for the separation of the "audiogram + calibration" results into earphone + mobile calibration and audiogram (i.e., separate device-pair compensation data and audiogram data cannot be derived from S1). To do this, a Bayesian inference model is constructed that uses update rules to incrementally estimate the model output as more information becomes available. As more and more users switch devices and perform update tests, the Bayesian inference model incorporates information for updating the SPL calibration for a given device. This information can be distributed to users who have not or have not yet performed update tests.
[0145] S3. SPL Calibration and Frequency Response Solution: The method in S2 allows for the separation of SPL calibration and audiogram + device frequency response. However, there is insufficient information to separate the device frequency response from the audiogram. Instead of performing a single point test as in S1, two frequency points can be tested. This allows the difference in SPL calibration between the test and two devices, as well as two different frequency points, to be used as inputs to the model, allowing the model to establish both SPL calibration and frequency response. An inference model can be constructed to separate the frequency response characteristics of the phone and earphones.
[0146] FIG. 11 shows an example comparison of the audiogram of a known device (graph 1110), the audiogram of a new device using SPL and response recalibration + adjustment (graph 1120), and the device response (graph 1130) as a function of frequency.
[0147] S4. Solution for SPL calibration, frequency response, and device characteristic variability: Some devices do not have consistent SPL or frequency response calibration across manufacturing runs. To solve this, a k-means clustering step can be introduced to collect similar devices together. A Bayesian inference model can use this information to learn which cluster a user's device belongs to in order to speed up onboarding.
[0148] FIG. 12 shows an example comparison of the audiogram of a known device (graph 1210), the audiogram of a new device using SPL and response recalibration plus adjustment (graph 1220), the device response (graph 1230), and the variation in the device response (graph 1230) as a function of frequency.
[0149] Initial Onboarding A full audiogram test is performed (e.g., all frequencies are tested at the left and right ear (L&R)). The measured hearing levels are split into an SPL calibration, and into audiogram values (hearing levels) for each frequency and device response values. Each value is assigned a probability (e.g., Gaussian) distribution. If the device is known, the known response and SPL can be used. In this case, the associated probability distribution may be narrow. If the device is unknown, the device response may be flat and a default SPL value can be used. The associated probability distribution is wide, making it more likely that an update will change the split between audiogram and calibration.
[0150] Audiogram and calibration updates using Bayesian inference Onboarding establishes the audiogram and device calibration and associated confidence. When a user switches to a new device (cell phone and / or earphones), a subset of frequency test points must be re-evaluated. As with onboarding, the device characteristics may or may not be known. If known, the device calibration has a high confidence level; if unknown, the confidence level is low (wide probability distribution). Bayesian inference can be used to update the probability distribution of the audiogram and calibration points for both the old and new devices. As more points are tested, the model becomes more certain about what portion of the test values should be assigned to the old / new device calibration and what should be assigned to the audiogram. Importantly, the probability distribution of the device calibration can be updated based on data from many users, allowing for faster updates to device characteristics.
[0151] Device Clustering As more users perform update tests, the model is able to see trends in data for the same device model identifier. Clustering within device identifiers allows the model to infer device variations that occur due to differences in manufacturing or other causes. As users perform more test updates, the model becomes more certain about which cluster a device belongs to, allowing for more accurate updates of device characteristics.
[0152] advantage In addition to the above advantages, the technology according to the present disclosure may have the following advantages. Calibration (e.g., measuring the SPL output and frequency response of each device) is no longer necessary. The proposed technology can work with any ear device, known or unknown. This simplifies both user onboarding and the development of hearing optimization solutions. The proposed technique allows for easy construction of a library of instrument responses and calibrations.
[0153] Apparatus for carrying out methods according to the present disclosure Finally, while methods according to embodiments of the present disclosure have been described above, the present disclosure also relates to devices (e.g., computer-implemented or computing devices, such as playback devices or server devices) for implementing the methods and techniques described throughout this disclosure. FIG. 13 illustrates an example of such a device 1300. In particular, the device 1300 includes a processor 1310 and a memory 1320 coupled to the processor 1310. The memory 1320 can store instructions for the processor 1310. The processor 1310 can also receive appropriate input data 1330 (e.g., statistical information, hearing thresholds, subjective responses from a user, etc.), depending on the use case and / or implementation. The processor 1310 can be adapted to perform the methods / techniques described throughout this disclosure (e.g., method 500 of FIG. 5 , method 600 of FIG. 6 , method 800 of FIG. 8 ) and generate corresponding output data 1340 (e.g., an audiogram estimate for the user, a hearing loss-compensated playback audio signal, etc.), depending on the use case and / or implementation.
[0154] It will be appreciated that the present disclosure further relates to corresponding computer programs and computer readable storage media storing such computer programs.
[0155] <Interpretation> Aspects of the systems described herein can be implemented in any suitable computer-based processing network environment for processing digital data (such as a standalone playback device, a server, a cloud environment, etc.). Portions of the system may include one or more networks including any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0156] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that various functions disclosed herein may be described in terms of their operation as hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media, using any number of combinations of register transfers, logical components, and / or other characteristics. The computer-readable media on which such formatted data and / or instructions are embodied include various forms of physical (non-transitory) non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.
[0157] Specifically, it should be understood that the embodiments may include hardware, software, and electronic components or modules, which, for purposes of explanation, may be shown and described as if the majority of the components were implemented solely in hardware. However, those skilled in the art will recognize, based on reading this detailed description, that in at least one embodiment, electronic-based aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) executable by one or more electronic processors, such as microprocessors and / or application-specific integrated circuits (ASICs). As such, it should be noted that a number of hardware- and software-based devices and a number of different structural components may be used to implement the embodiments. For example, a system, block, or module described in the context of Figures 1, 2, 13, 7, or 13 above may include one or more electronic processors, one or more computer-readable media modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0158] While one or more implementations have been described by way of example and in terms of specific embodiments, it is to be understood that the one or more implementations are not limited to the disclosed embodiments. To the contrary, the implementations are intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0159] It is also to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, the use of "including," "comprising," or "having" and variations thereof is meant to encompass the items listed thereafter and equivalents thereof, as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled" are used broadly and encompass both direct and indirect mounting, connecting, supporting, and coupling.
[0160] Enumerated exemplary embodiments Various aspects and implementations of the present disclosure can be understood from the following non-claimed enumerated example embodiments (EEE).
[0161] (EE1) A method of estimating a hearing threshold for a first user of a media playback device, said method comprising: obtaining user hearing threshold data corresponding to the first user (in some embodiments, a set of hearing thresholds for one or both ears of the user corresponding to one or more frequencies (e.g., thresholds expressed as digital signal levels); obtaining population hardware data corresponding to a plurality of media playback devices (in some embodiments, the hardware data includes data characterizing a population of listening devices (e.g., headphones, hearing aids, speakers) and media playback systems (e.g., mobile phones, televisions), data identifying the distribution of models, types, brands, etc. of listening devices and / or media playback devices); obtaining population hearing threshold data corresponding to a plurality of second users (in some embodiments, the hearing threshold data includes data characterizing hearing thresholds of people in a population (e.g., a population corresponding to a set of demographic characteristics); determining an estimated audiogram of the first user based on the user hearing threshold data, the population hardware data, the population hearing threshold data, and audiogram data representative of normal hearing (e.g., audiogram-based data indicating no hearing loss); A method comprising:
[0162] (EEE2) The method of EEE1, wherein the user hearing threshold data includes hearing threshold measurements expressed as digital signal levels (e.g., dB re FS) for two ears (e.g., left and right) for a plurality of frequencies (e.g., 2 frequencies, 11 frequencies, 17 frequencies, etc.).
[0163] (EEE3) The step of acquiring user hearing threshold data includes: outputting, at the media playback device, a plurality of audio signals (e.g., outputting audio tones or signals corresponding to a plurality of frequencies via headphones connected to the media playback device); receiving a user input corresponding to each of the plurality of audio signals (e.g., a user input indicating a perception of a particular audio signal); wherein the user hearing threshold data is based on data corresponding to the received user input.
[0164] (EEE4) A method according to any one of EEE1 to EEE3, wherein the user hearing threshold data is obtained from at least one of a memory of the media playback device storing application data and a server communicating with the media playback device via one or more networks.
[0165] (EEE5) The method described in any of EEE1 to EEE4, wherein the population hardware data includes statistical information describing the frequency response and associated covariance of a population of listening devices and media playback devices (e.g., devices commonly used by consumers).
[0166] (EEE6) The method according to any one of EEE1 to EEE5, wherein the population hardware data includes an average frequency response vector corresponding to an average of frequency responses for a plurality of frequencies.
[0167] (EEE7) A method described in any of EEE1 to EEE6, wherein the population hearing threshold data includes statistical information describing the audiograms of people in a population (e.g., audiograms corresponding to a demographically related population of people or users of a media playback device).
[0168] (EEE8) The method according to any one of EEE1 to EEE7, wherein the population hearing threshold data comprises mean audiogram data (eg, mean vector) and corresponding covariance data (eg, covariance vector).
[0169] (EEE9) The method described in EEE8, wherein the population hearing threshold data is adjusted based on at least one of age (e.g., age or age range, age threshold, hearing age, etc.) and gender of the first user.
[0170] (EEE10) The step of acquiring population hearing threshold data includes: receiving user input corresponding to self-identifying demographic data (e.g., user input indicating age or age range, gender, location, occupation, etc.); The method of any one of EEE1 to EEE9, wherein the population hearing threshold data is based on data corresponding to the received user input.
[0171] (EEE11) The method according to any one of EEE1 to EEE10, wherein the step of determining the estimated audiogram is performed according to a Bayesian maximum a posteriori (MAP) estimation technique.
[0172] (EEE12) A method as described in any of EEE1 to EEE11, wherein the step of determining the estimated audiogram is further based on one or more of data representing the variance in playback levels (e.g., the variance between devices) and data representing the variance in user responses.
[0173] (EEE13) The method according to any one of EEE1 to EEE12, wherein the step of determining an estimated audiogram is performed by the media playback device.
[0174] (EEE14) The method of any of EEE1 to EEE13, wherein the step of determining an estimated audiogram is performed by a server device (eg, a remote server, a companion device) in communication with the media playback device.
[0175] (EEE15) receiving audio for playback at said media playback device; determining a gain set based on the estimated audiogram; generating hearing-optimized audio for playback by applying the gain set to the audio; playing the hearing-optimized audio with the media playback device; The method of any one of EEE1 to EEE14, further comprising:
[0176] (EEE16) A method according to any one of EEE1 to EEE15, further comprising providing data representing the estimated audiogram or a compensation gain associated with the estimated audiogram to an application on the media playback device (e.g., a media player app, a communication app, a game, etc.).
[0177] (EEE17) A method according to any one of EEE1 to EEE16, further comprising transmitting data representing the estimated audiogram or a compensation gain associated with the estimated audiogram to another device (e.g., a server or device different from the media playback device).
[0178] (EEE18) generating a set of personalized compensation gains based on the estimated audiogram and data representing a personalizer head-related transfer function; generating personalized hearing-compensated audio by applying a set of personalized compensation gains to the audio received for playback; The method according to any one of EEE1 to EEE17, further comprising:
[0179] (EEE19) A method of estimating device calibration of a first media reproduction device, comprising: obtaining device hearing threshold data associated with a first plurality of users and device types; obtaining population hearing threshold data corresponding to a second plurality of users (in some embodiments, the population hearing threshold data includes data characterizing the hearing thresholds of a population of people (e.g., a population corresponding to a set of demographic characteristics, a population representative of the demographics of the device's user base or a geographic region, etc.); determining estimated device calibration data based on the device hearing threshold data, the population hearing threshold data, and audiogram data representative of normal hearing (e.g., audiogram-based data indicating no hearing loss); A method comprising:
[0180] (EEE20) obtaining user hearing threshold data corresponding to a first user (in some embodiments, a set of hearing threshold measurements (e.g., thresholds expressed as digital signal levels) for one or both ears of the user corresponding to one or more frequencies); determining an estimated audiogram of the first user based on the user hearing threshold data, the estimated device calibration data, and audiogram data representative of normal hearing (e.g., audiogram-based data indicative of no hearing loss); The method of EEE1 further comprising:
[0181] (EEE21) The method according to any one of EEE19 to EEE20, wherein the device hearing threshold data comprises a plurality of hearing threshold vectors associated with device types and representing digital levels of a plurality of frequencies.
[0182] (EEE22) The method of any of EEE19 to EEE21, wherein the device type is defined in part by at least one of a device model (e.g., a phone model such as an iPhone 13), a Bluetooth identification value, a device brand, a device form factor (e.g., a mobile phone, a hearing aid, a VR / AR headset), or an application (e.g., a gaming system).
[0183] (EEE23) A method as described in any of EEE19 to EEE22, wherein the population hearing threshold data includes statistical information describing the audiograms of a population of media playback device users (e.g., audiograms corresponding to a population of demographically related users).
[0184] (EEE24) The method according to any one of EEE19 to EEE23, wherein the population hearing threshold data comprises mean audiogram data (eg, mean vectors) and corresponding covariances (eg, covariance vectors).
[0185] (EEE25) The method of EEE24, wherein the population hearing threshold data is adjusted based on at least one of age (e.g., age or age range, hearing age, etc.) and gender of the first user.
[0186] (EEE26) The method according to any one of EEE19 to EEE25, wherein the step of determining the estimated audiogram is performed by a media playback device.
[0187] (EEE27) The method of any of EEE19 to EEE26, wherein the step of determining the estimated device calibration data is performed by a server device (eg, a remote server (such as a companion device)) in communication with the media playback device.
[0188] (EEE28) The method of any one of EEE19 to EEE27, wherein the data corresponding to device type and the population hearing threshold data is received from a plurality of second media reproduction devices.
[0189] (EEE29) Further comprising receiving demographic data (e.g., age, gender, user ID, geographic region, etc.) from the plurality of second media playback devices in response to user input; The method of any of EEE19 to EEE28, wherein determining estimated device calibration data is further based on a subset of the received demographic data.
[0190] (EEE30) and below: obtaining updated user threshold data corresponding to a single frequency (e.g., thresholds generated with a different device (e.g., headphones) than that used to determine the estimated device calibration data); determining an offset between the value of the estimated device calibration data corresponding to the single frequency and updated user threshold data corresponding to the single frequency; determining updated estimator calibration data by applying the offset to data corresponding to each frequency represented in the estimator calibration data; generating updated estimated device calibration data for the second media playback device by using the estimated device calibration data.
[0191] (EEE31) The method of EEE30, further comprising the step of optimizing said offset using additional measurements based on frequencies other than said single frequency using a Bayesian MAP estimator.
[0192] (EEE32) A computing device comprising: at least one processor; a memory storing instructions that, when executed by the at least one processor, cause the computing device to perform a method as described in any one of EEE1 to EEE31; Computing equipment, including
[0193] (EEE33) A non-transitory computer-readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform a method according to any one of EEE1 to EEE31.
[0194] (EEE34) A computer program comprising instructions that, when executed by a computing device, cause the computing device to perform a method according to any one of EEE1 to EEE31.
[0195] References [1] Rasmussen, CE and Williams, CKI (2006). Gaussian processes for machine learning. The MIT Press. ISBN0-262-18253-X.
Claims
1. 1. A method for estimating an audiogram for a user of a media playback device, said method comprising: obtaining user hearing threshold data for the user, the user hearing threshold data indicating hearing thresholds for one or more frequencies and one or both ears; obtaining sample hearing threshold data, the sample hearing threshold data indicative of hearing thresholds of a sample set of individuals; obtaining at least one of sample calibration data and sample noise data, wherein the sample calibration data is indicative of a frequency response of a media playback device sample set, and the sample noise data is indicative of variations in playback level and / or variations in user response in the process of measuring hearing thresholds; determining an estimate of an audiogram for the user based on the user hearing threshold data, the sample hearing threshold data, and at least one of the sample calibration data and the sample noise data; A method comprising:
2. The method of claim 1 , wherein determining the audiogram estimate is further based on normal hearing data indicating expected hearing thresholds in the absence of hearing loss.
3. 3. The method of claim 1, wherein determining the audiogram estimate comprises applying relative weights to the user hearing threshold data and the sample hearing threshold data based on at least one of the sample calibration data and the sample noise data.
4. The method according to any one of claims 1 to 3, wherein the step of determining an estimate of the audiogram is based on a Bayesian Maximum A Posteriori (MAP) estimation technique.
5. The method of any one of claims 1 to 4, wherein the user hearing threshold data indicates hearing thresholds for the left and right ear at multiple frequencies.
6. The method according to any one of claims 1 to 5, wherein the hearing thresholds of the user hearing threshold data are represented by digital signal levels of the playback device.
7. The step of obtaining user hearing threshold data includes: outputting, by the media playback device, a plurality of audio signals at different frequencies; receiving a user input responsive to the output audio signal; generating the user hearing threshold data based on the received user input; The method according to any one of claims 1 to 6, comprising:
8. The method of any one of claims 1 to 7, wherein the sample hearing threshold data is indicative of audiogram information for the individual sample set.
9. The method of any one of claims 1 to 8, wherein the sample hearing threshold data represents the mean and covariance of the audiograms for the sample set of individuals.
10. A method according to any preceding claim, wherein the sample hearing threshold data indicates the mean and covariance of hearing thresholds for each frequency and ear for the sample set of individuals.
11. The method of any preceding claim, wherein the sample calibration data indicates the mean and covariance of frequency responses for the media playback device sample set.
12. The audiogram estimate y is given by: [Equation 1] where μ h is arbitrary, and M=Σ a (Σ a +E) -1 and E is Σ h , σ r 2 I, and σ p 2 J, and Σ a is the covariance of the vector representation of the sample hearing threshold data, and μ a → is the mean of the vector representation of the sample hearing threshold data, and Σ h is the covariance of the vector representation of the sample calibration data, and μ h → is the mean of the vector representation of the sample calibration data, and σ r 2 represents the variability of the user response in the process of measuring the hearing threshold, and σ p 2 12. A method according to any one of claims 1 to 11, wherein ℓ represents the variation in playback level, I is an identity matrix, and J is a matrix of ones, and the vector representation comprises elements for each pair of one of the left and right ears and a frequency in a predetermined set of frequencies.
13. The method of any one of claims 1 to 12, wherein the personal sample set is selected based on at least one user attribute of the user.
14. The method according to any one of claims 1 to 13, wherein the step of determining an audiogram estimate is performed in the media playback device or in a server device in communication with the media playback device.
15. receiving audio data for playback at the media playback device; determining a compensation gain set based on the determined audiogram estimate; generating hearing-optimized audio data by applying the determined compensation gain set to the audio data; rendering the hearing-optimized audio data for playback; The method of any one of claims 1 to 14, further comprising:
16. The method of claim 15 , wherein determining the compensation gain set is further based on the received audio data.
17. 1. A method of estimating calibration data for a media playback device, the calibration data indicative of a frequency response of the media playback device, the method comprising: obtaining first sample hearing threshold data, the first sample hearing threshold data indicative of hearing thresholds of a first personal sample set and associated with a given device type; obtaining second sample hearing threshold data, the second sample hearing threshold data representing hearing thresholds of a second sample set of individuals different from the first sample set of individuals; determining an estimate of the calibration data based on the first sample hearing threshold data, the second sample hearing threshold data, and normal hearing data indicative of hearing thresholds expected in the absence of hearing loss; A method comprising:
18. obtaining user hearing threshold data for a user of the media playback device; determining an estimate of the user's audiogram based on the user hearing threshold data, the estimate of the calibration data, and the normal hearing data; 20. The method of claim 17 further comprising:
19. 20. The method of claim 18, wherein the user hearing threshold data, the first sample hearing threshold data, and the second sample hearing threshold data each represent hearing thresholds for one or more frequencies and one or both ears.
20. The method of any one of claims 18 to 19, wherein the user hearing threshold data indicates hearing thresholds for the left and right ear at multiple frequencies.
21. The method according to any one of claims 18 to 20, wherein the hearing thresholds of the user hearing threshold data are represented by digital signal levels of the playback device.
22. The method of any one of claims 17 to 21, wherein the second sample hearing threshold data is indicative of audiogram information relating to the second personal sample set.
23. A method according to any one of claims 17 to 22, wherein the second sample hearing threshold data represents an average audiogram for the second sample set of individuals.
24. A method according to any one of claims 17 to 23, wherein the second sample hearing threshold data represents an average of the hearing thresholds for each frequency and ear for the second sample set of individuals.
25. The method of claim 18 or any one of claims 19 to 24 when dependent on claim 18, wherein the second personal sample set is selected based on at least one user attribute of the user.
26. A method according to claim 18 or any one of claims 19 to 25 when dependent on claim 18, wherein the step of determining an audiogram estimate is performed in the media playback device or in a server device in communication with the media playback device.
27. obtaining updated user hearing threshold data for the user for a second media playback device different from the media playback device, the updated user hearing threshold data indicating an updated hearing threshold for a given frequency; determining an offset between a user hearing threshold at the given frequency indicated by the user hearing threshold data and the updated hearing threshold; determining a second calibration data estimate for the second media playback device based on the calibration data estimate and the determined offset; The method according to claim 18 or any one of claims 19 to 26 dependent on claim 18, further comprising:
28. obtaining updated user hearing threshold data for the user for a second media playback device different from the media playback device, the updated user hearing threshold data indicating updated hearing thresholds for a plurality of given frequencies; determining an estimate of an offset between the calibration data and second calibration data for the second media playback device based on the user hearing thresholds at the plurality of given frequencies indicated by the user hearing threshold data, the updated hearing thresholds, the second sample hearing threshold data, and sample noise data, the sample noise data indicating variability in user response in a process of measuring a hearing threshold; determining an estimate of the second calibration data based on the calibration data estimate and the determined estimate of the offset; The method according to claim 18 or any one of claims 19 to 26 dependent on claim 18, further comprising:
29. 30. The method of claim 28, wherein determining the estimate of the second calibration data is based on a Bayesian maximum a posteriori (MAP) estimation technique.
30. 1. A computing device comprising: at least one processor; a memory storing instructions which, when executed by said at least one processor, cause said computing device to perform the method of any one of claims 1 to 29; Computing equipment, including
31. A computer program comprising instructions which, when executed by a computing device, cause the computing device to carry out a method according to any one of claims 1 to 29.
32. 32. A non-transitory computer-readable storage medium storing the computer program of claim 31.
Citation Information
Patent Citations
Audibility test system and hearing aid selection system using the same
JP2005137879A