Audio processing using hearing loss data

An integrated audio processing approach optimizes multimedia content for individuals with hearing impairments by calculating signal gains based on hearing loss and ambient noise, addressing the limitations of traditional fitting rules and enhancing multimedia experience.

JP2026517966APending Publication Date: 2026-06-02DOLBY LABORATORIES LICENSING CORP +1

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-05-15
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing hearing aid fitting rules fail to adequately address the complex interactions between hearing loss, media playback level, and ambient noise, leading to suboptimal preservation of loudness, timbre, and speech intelligibility in multimedia content, particularly in noisy environments.

Method used

An integrated approach that utilizes hearing loss data, ambient sound data, and media playback level data to calculate frequency- and time-dependent signal gains, optimizing audio processing for individuals with hearing impairments, considering both environmental and media masking effects.

Benefits of technology

Enhances the perception of multimedia content by preserving artistic intent, improving speech intelligibility, and maintaining spectral naturalness and spatial audio quality, even in noisy conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517966000001_ABST
    Figure 2026517966000001_ABST
Patent Text Reader

Abstract

A method for processing audio, comprising: receiving audio data for playback on a playback device, the audio data including a target audio component; receiving masking audio data; obtaining hearing loss data associated with a listener; obtaining playback level data; calculating a target excitation pattern based on the target audio component and playback level data; calculating a masking excitation pattern based on the masking audio data; calculating a signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data; and applying the signal gain to the target audio component to produce a compensated target audio component.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Cross-reference of related applications] This application claims priority to U.S. Provisional Application No. 63 / 502,391, filed on 15 May 2023, and U.S. Provisional Application No. 63 / 588,909, filed on 9 October 2023, the contents of which are incorporated herein by reference in their entirety.

[0002] [Technical field to which the review belongs] This disclosure relates to signal processing in general, and more specifically to techniques, devices, and computer executable instructions for processing signals related to multimedia content. In particular, this disclosure relates to compensation for hearing impairment. [Background technology]

[0003] Recent statistics show that approximately 15% of American adults report hearing impairment, and age is the strongest predictor of hearing loss. Men are nearly twice as likely to have hearing loss as women. About 2% of the adult population aged 45–54 have physical hearing loss, increasing to 8.5% and 25% in the 55–64 age range, and then to 65–74%. Furthermore, it is estimated that approximately 28 million American adults could benefit from using hearing aids.

[0004] To provide solutions for people with hearing loss, the hearing aid industry has developed hearing aids and associated fitting rules for decades. The goal of these fitting rules is to describe an auditory optimization algorithm (or hearing aid) in terms of which gain should be applied at which frequencies to the signal captured by the microphone on the hearing aid and reproduced to the listener. In the early days, around the 1950s, the half-gain rule was a common method for determining frequency-dependent gain. In the absence of dynamic range processing capabilities, the idea was to apply a gain corresponding to half of the hearing loss in decibels (dB). Thus, this fitting rule provided a single gain target for any particular type of hearing loss, which was the same for all input levels.

[0005] In the 1990s, compression-based fitting rules emerged, with popular examples including DSL (Desired Sensation Level) and NAL-NL1 (National Acoustics Labs, Non-Linear, version 1). The idea behind the NAL fitting rule is that it attempts to equalize, rather than normalize, loudness relationships across speech frequencies. Specifically, it assumes that speech intelligibility is maximized when all frequencies of speech are amplified to sound equally loud. In other words, the NAL method does not preserve loudness relationships between speech frequencies; instead, it aims to maximize the information provided to the listener regarding speech, assuming a certain amount of hearing loss. In NAL terminology, this is described as maximizing "effective audibility" rather than "audibility" itself or the perceived level. [Overview of the project] [Problems that the invention aims to solve]

[0006] Our perception of loudness, timbre, spectral naturalness, externalization of spatial audio, and speech intelligibility depends on various contextual factors such as the amount of hearing loss, auditory masking by media components (such as music and effects), media playback level, and ambient sounds such as background noise. These factors do not simply add up; instead, they interact in complex ways. The preservation of the artistic intent, naturalness, immersion, sense of presence, realism, and speech intelligibility of media reproduced on hearables (e.g., headphones, earphones, hearing aids, etc.) can benefit greatly from algorithms that optimize the combined adverse effects of hearing loss, the impact of playback level, and the amount of masking from other media components and ambient sounds. This is relevant to content types other than speech, including traditional linear media content such as music and movies, but also to non-linearly generated or interactive media content such as AR, VR, and games.

Means for Solving the Problems

[0007] In this specification, techniques are described for providing the preservation of the artistic intent of media (e.g., loudness, timbre, spectral naturalness, externalization of spatial audio and speech intelligibility, envelope shaping, sense of presence, realism, etc.) reproduced on playback devices (e.g., headphones, earphones, hearing aids, AR / VR devices, head-mounted displays, etc.) for individuals with hearing impairments. The techniques described are compatible with devices and systems that support MPEG-I.

[0008] A first aspect relates to a method for processing audio for a listener of a playback device, comprising: receiving audio data for playback on the playback device, wherein the audio data includes a target audio component; receiving masking audio data; obtaining hearing loss data associated with the listener; obtaining playback level data; calculating a target excitation pattern based on the target audio component and the playback level data; calculating a masking excitation pattern based on the masking audio data; calculating a signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data; and applying the signal gain to the target audio component to produce a compensated target audio component.

[0009] Playback level data describes the (frequency-dependent) conversion of a digital input signal to a sound pressure level that reaches the listener's ear. Playback level data may, for example, be the actual playback level data received from the playback device; or, in some cases, it may be estimated based on knowledge of the type of playback device; or default playback level data may be used.

[0010] The masking audio data may include background audio components contained in the audio data, and the masking excitation pattern may then include a background excitation pattern based on the background audio components and playback level data.

[0011] The masking audio data may also include ambient sounds picked up by a microphone placed close to the listener, and the masking excitation pattern may then include an ambient excitation pattern based on the ambient sounds. The ambient excitation pattern may also be based on “playback device data” that describes the influence of the playback device on the ambient sounds. Such playback device data may include, for example, the transfer function of the active filter of a set of hearables.

[0012] A further embodiment relates to a method for audio rendering, the method comprising: receiving an audio bitstream including a target audio component and a background audio component; receiving audio scene information including the current head posture and / or gaze of a listener; processing the audio scene information to obtain an audio scene state; rendering audio data based on the scene state to obtain a rendered target audio component and a rendered background audio component; receiving hearing loss data associated with a listener; receiving playback level data; calculating a target excitation pattern based on the target audio component and the playback level data; calculating a masking excitation pattern based on the background audio component and the playback level data; calculating a signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data; and applying the signal gain to the rendered target audio component to produce a compensated rendered target audio component.

[0013] Another aspect relates to an audio renderer, the audio renderer comprising: a decoder for receiving and decoding an audio bitstream to provide input audio including a target audio component and a background audio component; a scene controller configured to receive audio scene information and provide an audio scene state based on the audio scene information; and a rendering pipeline configured to render audio data based on the scene state to obtain a rendered target audio component and a rendered background audio component, receive hearing loss data associated with a listener, receive playback level data, calculate a target excitation pattern based on the target audio component and playback level data, calculate a masking excitation pattern based on the background audio component and playback level data, calculate a signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data, and apply the signal gain to the rendered target audio component to produce a compensated rendered target audio component.

[0014] Such audio rendering is relevant to virtual reality (augmented reality, extended reality) applications with accessibility interfaces. For users with hearing impairments, the calculation and application of signal gain as disclosed herein may be advantageous.

[0015] Such audio rendering may involve receiving descriptive data indicating whether a particular audio stream of audio data is a target audio component or a background audio component. Such descriptive data may be present in existing or future VR / AR / XR audio rendering devices, such as rendering devices compliant with the MPEG-I standard.

[0016] The proposed approach enables combined optimization by utilizing one or more of the following:

[0017] • Knowledge of ambient sounds, for example, via microphones attached to earphones or other wearable devices.

[0018] • Knowledge of audio media playback levels, such as knowing the media digital signal levels and their corresponding sound pressure levels, given the current volume control settings on the playback device.

[0019] • For speech intelligibility, media audio can be separated from media background audio (where applicable).

[0020] • Knowledge about the listener's hearing loss (e.g., age + sex to derive the listener's audiogram or the population mean audiogram) • To understand the frequency-dependent attenuation of the acoustic environment when a listener is wearing a hearable device, and in addition, to know which mode (e.g., passive, pass-through, or acoustic noise cancellation mode) these hearable devices are operating in to determine the perceived ambient sound level.

[0021] Given some or all of the above information, the perceptual loudness model can predict the optimal media playback gain as a function of frequency and time, in order to optimize it together with the aforementioned context, dialogue, and personalization factors.

[0022] The executable instructions for performing these functions are optionally included in a non-temporary computer-readable storage medium or in other computer program products configured to be executed by one or more processors.

[0023] Such techniques for performing these functions can complement or replace other methods for performing similar functions. [Brief explanation of the drawing]

[0024] To better understand the various embodiments described, the following description should be referred to in conjunction with the following drawings, and throughout the drawings, similar descriptions refer to corresponding parts. [Figure 1] Several embodiments demonstrate spectral naturalness and timbre distortion due to hearing impairment. [Figure 2] This is a schematic block diagram of a system according to one embodiment of the present invention. [Figure 3] Figure 2 is a more detailed block diagram of the media audio processing. [Figure 4] An MPEG-I audio renderer that implements an embodiment of the present invention is shown. [Figure 5] This shows the system architecture of the MPEG-I renderer. [Figure 6] This shows the processing in an MPEG-I renderer. [Figure 7] Figure 6 shows the rendering stages in the rendering pipeline. [Figure 8] This flowchart illustrates methods for improving the user experience for people with hearing impairments, based on several embodiments. [Modes for carrying out the invention]

[0025] The following description includes exemplary methods, parameters, devices, etc. However, it should be noted that such description is not intended to limit the scope of this disclosure and is instead provided as a description of exemplary embodiments.

[0026] There is a need for electronic devices that provide an improved user experience for people with hearing impairments. For example, there is a need for electronic devices that deliver audio to users that preserves artistic intent (e.g., loudness, timbre, spectral naturalness, externalization of spatial audio, and speech clarity, envelope, immersion, realism, etc.).

[0027] In the following description, terms such as “first,” “second,” etc., are used to describe various elements, but these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, without departing from the scope of the various embodiments described, first data may be called second data, and similarly, second data may be called first data. First data and second data are both data, but they are not identical data.

[0028] The terms used in the description of the various embodiments described herein are intended solely to describe specific embodiments and are not intended to limit them. As used in the description of the various embodiments described and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context explicitly indicates otherwise. The terms “and / or” as used herein will also be understood to refer to and encompass one or more any and all possible combinations of the enumerated items relating to the description. The terms “includes,” “including,” “comprises,” and / or “comprising,” as used herein, specify the presence of the described features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0029] The term “case” is optionally interpreted, depending on the context, as to mean “when,” “if,” “in response to a determination,” or “in response to detection.” Similarly, the phrases “if determined” or “if [the stated condition or event] is detected” are optionally interpreted, depending on the context, as to mean “when determined,” “in response to a determination,” “when [the stated condition or event] is detected,” or “in response to detection of [the stated condition or event].”

[0030] We now turn our attention to embodiments of the disclosed techniques that may be implemented on electronic devices. In some embodiments, the electronic device includes memory (optionally including one or more computer-readable storage media) and one or more processing units (such as a CPU, GPU, DSP, or ASIC). In some embodiments, the electronic device is a device type frequently used to provide AR / VR content. For example, in some embodiments, the electronic device is a smartphone, tablet, or other computing system (e.g., a desktop or laptop) that includes audio and video capture capabilities (natively or via a hardware capture device communicably coupled to the device).

[0031] In some embodiments, the electronic device is implemented as part of a distributed system that includes one or more components located in different locations. For example, in some embodiments, the electronic device includes a local hardware component and one or more remote hardware components (e.g., cloud-based components hosted on public and / or private cloud infrastructure, SaaS, PaaS, etc.). In some embodiments, the electronic device is or includes a smartphone, tablet, or other computing system (e.g., desktop, laptop) that includes audio and video capture capabilities (natively or via a hardware capture device communicatively coupled to the device).

[0032] The devices described above are merely examples of electronic devices, and it should be understood that devices may optionally have more or fewer components than those described, may optionally combine two or more components, or may optionally have different configurations or arrangements of components. The various components of the devices described above and throughout this document are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0033] Now, let us consider, for example, embodiments implemented on electronic devices as described above.

[0034] Hearing impairment is, in most cases, frequency and level dependent, and therefore loudness loss (or increase) is also frequency and level dependent. Consequently, hearing impairment distorts the naturalness of the spectrum and timbre. Figure 1 shows an example of the input audio signal spectrum perceived by a listener with normal hearing (solid line 11) and the audio signal spectrum perceived by a listener with high-frequency hearing impairment (dotted line 12). Region 13 identifies loudness loss and perceived timbre changes. The auditory optimization algorithm aims to reduce loudness loss and timbre changes by amplifying the affected signal frequencies and creating a perceived input signal spectrum that is closer to the input signal spectrum perceived by a listener without sensorineural hearing impairment (dashed line 14).

[0035] It should be noted that the actual signal level is not the only factor. Frequency-dependent playback level data, which can include the playback frequency response of the hearable, also plays an important role. This playback level data enables the conversion from media digital signal levels to acoustic sound pressure levels. Therefore, if such data is available, it is recommended to adapt it when estimating auditory-optimized gains. It should be noted that these gains are applied to the input signal (input audio stream) to obtain the final output signal.

[0036] The result of combining hearing impairment and noise masking when listening to media. Hearing impairment and noise masking (from both environmental and media backgrounds) can have significant impacts on how we perceive media, including the following:

[0037] • Difficulty understanding dialogue. Due to more pronounced masking of media background or environmental background noise, elevated auditory thresholds, and modified perception of loudness as a function of signal levels, important details of speech may be lost or reduced, leading to decreased dialogue intelligibility.

[0038] • Perceiving distorted timbre or a lack of spectral naturalness. Since both hearing loss and auditory masking are frequency and level-dependent processes, the amount of information lost due to hearing impairment or environmental noise masking is also frequency-dependent, reducing the perceived timbre and spectral naturalness. For example, in the case of low-frequency traffic noise, a listener may not be able to hear the bass in a music track, while other instruments and vocals may already be loud enough. In such cases, increasing the playback volume will not eliminate the timbre distortion.

[0039] • Loss of detail in instrument or audio objects, and loss of background sound. Quiet elements of instruments (such as high-frequency harmonics, subtle effects of string instruments, a vocalist's breathing, and quiet background sounds) may become inaudible or have reduced loudness due to hearing loss and noise masking.

[0040] • Changes in the perception of dynamics and dynamic range. It is well known that hearing impairment alters the perception of dynamics and the relationship between signal level and perceived loudness.

[0041] • Reduction of spatial and externalization for binaural audio. Externalization and spatiality are governed by subtle cues such as minimal acoustic environment simulation. Noise masking and hearing loss can cause these crucial cues to be lost or reduced in perceived level, negatively impacting the spatial audio experience.

[0042] Challenges in applying hearing aid fitting rules to media playback Given the numerous studies on optimizing speech intelligibility in noisy environments for hearing aids, and considering the effects of ambient noise and hearing loss on media consumption as outlined in the previous section, it is conceivable to apply hearing aid fitting rules and algorithms to media playback in noisy environments. However, in practice, this often results in suboptimal or insufficient performance for various reasons. For example, Signal components that are not considered to be within the speech frequency range are not considered important according to conventional fitting rules or hearing aid algorithms. As a result, signal components below several hundred hertz are inadequately amplified compared to the amount of hearing loss or masking at those frequencies, and similar problems occur at higher frequencies. Consequently, conventional hearing aid fitting rules cause significant loss of low and high frequency perception in media types where these frequency ranges are important, such as music and movies.

[0043] Optimizing speech intelligibility by equalizing the loudness of audio components can cause significant distortion of the timbre and spectral naturalness of audio content, and more importantly, of the timbre of non-audio content such as music or movies.

[0044] Traditional hearing aid fitting rules often amplify the audio component to exceed the estimated ambient noise masking threshold. In noisy environments, this results in a very small dynamic range for the processed audio, caused by aggressive processing that is unsuitable for music or other media content such as movies.

[0045] Hearing aids have very few means of inter-device communication (e.g., communication between the left and right hearing aids, and communication between the hearing aid and a media playback device), and they mostly operate independently. Therefore, source isolation for stereo media content, for example, to separate audio from other media content, is notoriously difficult in hearing aids. However, in media playback use cases, the media is strictly isolated from ambient noise, and therefore no source isolation algorithm is required to separate audio from acoustic ambient noise.

[0046] Hearing aids do not have the concept of two independent masking sources: signals from the environment and background sounds in the medium. It is important to consider both, as both can have a significant masking effect, and / or some of these content components may be considered important depending on the use case and may require optimization.

[0047] The precise media playback level may be difficult to determine without further information regarding volume settings, audible sensitivity, and frequency response, and such parameters are essential for compensating for ambient noise during media playback.

[0048] Typical implementation This disclosure proposes an integrated approach / technology aimed at enabling listeners to hear content as intended by the content creator, which means listening to media content as follows:

[0049] • There was no significant hearing loss, and / or, • No environmental noise and / or media background masking, and / or • The playback level was close to the reference playback level, or at least provided a tone close to that of playback at the reference playback level.

[0050] The technologies described below include a perceptual modeling framework for media listening and optimization in noisy environments, and a system for simultaneously optimizing audio for environmental masking and hearing loss, which improve upon the aforementioned shortcomings of existing fitting and compensation techniques (e.g., those described in the introduction section above).

[0051] Figure 2 shows high-level system diagrams of several embodiments of the proposed approach / technology. The media playback device 20 here includes a media player 21 and a set of hearables 22 (e.g., a single transducer or a pair of transducers). The media player 21 may be any type of audio signal generator, such as a smartphone, home entertainment system, or virtual reality system. The hearables 22 may be a head-mounted display, headset, or a pair of earbuds worn by the listener. The hearables 22 include at least one earphone (two in the illustrated case), each having a transducer (not shown), most typically a loudspeaker. Each earpiece may also include at least one microphone 24 configured to receive sound from the environment. In some examples, the hearables 22 may consist of a single earpiece. Alternatively, the media playback device may be a single device including both the media player 21 and the hearable devices 22, such as a head-mounted display.

[0052] The media player 21 provides an input media signal 15 that is transmitted to the hearable device 22 via the media audio processing module 23. The media player 21 may be connected to the audio processing module 23 and / or the hearable device 22 by a wired or wireless connection. The media audio processing module 23 may be integrated with the media player 21 or the hearable device 22, or it may be a separate entity configured to intercept signals from the media player 21, process them, and transmit them to the hearable device 22. Such a separate entity may be part of a cloud processing solution. The sound captured by the hearable 22 may be communicated to the media player 21 and / or the media audio processing module 23.

[0053] The media audio processing module 23 is configured to receive ambient audio data 25 acquired from one or more microphones 24 on the hearable device 22, or by any other sound detection device.

[0054] The media audio processing module 23 is further configured to receive acoustic isolation data associated with the hearable 22, also known as playback device data 26. The playback device data 26 may include a passive portion relating to the occlusion effect that the hearable device 22 has with respect to external sounds (e.g., data describing sound attenuation caused by the physical occlusion of the ear canal by the hearable device). This passive portion is associated with a particular type and design of the hearable 22 and may be obtained directly from the hearable (e.g., during the setup procedure). The playback device data 26 may further include an active portion relating to the transfer function of an active filter in the hearable device 22. Such a filter may have attenuation or amplification effects on sounds of different frequencies. The active portion of the acoustic isolation data 26 may further depend on the selected mode of the hearable 22, e.g., active noise cancellation (ANC) mode, pass-through mode. Thus, the active portion may be dynamic according to user input. The current operating mode of the hearable 22 may be an explicit portion of the playback device data 26.

[0055] The media audio processing module 23 is further configured to acquire playback level data 27 associated with the hearable 22. This data describes the (frequency-dependent) conversion of the digital input signal 15 to the actual sound pressure level – generated by the loudspeaker of the hearable 22. This data depends on the specific type and design of the hearable device 22 and includes the frequency characteristics of the loudspeaker within the hearable device 22. In some embodiments, the actual playback data is received from the playback device; however, if this is not possible, the playback level data may be estimated, for example, based on the type of playback device used, or a default may be assumed.

[0056] The media audio processing module 23 is further configured to acquire hearing loss data 28 that describes the listener's hearing impairment. The hearing loss data 28 may be in the form of an audiogram or a set of parameters that enable the estimation of an audiogram. For example, an audiogram may be acquired from a database 29 based on demographic information about the listener (e.g., age or year of birth). Alternatively, an audiogram may be acquired by performing a listening test or received from a hearing loss data service provider.

[0057] The media audio processing module 23 may be further configured to receive loudness adjustment data 30, for example, volume control user settings.

[0058] The media audio processing module 23 is configured to process the input media 15 in response to ambient audio data 25, acoustic separation data / playback device data 26, playback level data 27, and hearing loss data 29 to generate processed media 16. Finally, the processed media 16 is played back on the hearable 22.

[0059] Embodiments using a perceptual loudness model The general approach outlined above involves the transfer functions H of the outer and middle ear. OME , and auditory filter transfer function H ERB Further details are provided regarding specific implementations based on perceptual loudness models, including modeling of the human ear.

[0060] Calculation of specific loudness levels The perceived loudness of a sound is discussed by excitation and specific loudness models, for example, as described at https: / / www.nidcd.nih.gov / health / statistics / quick-statistics-hearing. Such models typically begin with the calculation of excitation levels in a set of frequency bands, from which the loudness per band is then calculated. This loudness per band is called specific loudness. For simplicity, without loss of generality, assume that in practice, there is a range of spectral bands that function more or less independently and have their own set of parameters, and explain the process for a single frequency band.

[0061] The excitation level E, represented in the frequency domain t for the media target input signal X t (f) is given by the following.

[0062] [Number]

[0063] H ome (f) is the transfer function of the outer and middle ear, H erb (f,f c ) is the auditory filter transfer function centered on frequency f c , f is the frequency in Hz, P(f) represents the reproduction level data 27 that enables the conversion from the media digital signal level to the acoustic sound pressure level that can include the reproduction frequency response of the auditory device. The reproduction level data P(f) required to convert the digital signal level to the acoustic sound pressure level may depend on various factors such as the volume control settings of the playback device. The functions H ome (f) and H erb (f,f c ) may be given by an appropriate model of the auditory system, and examples may be found in various textbooks on this topic.

[0064] The specific loudness N t'(for example, associated with the signal X(f) in the target frequency band) is calculated as follows:

[0065]

number

[0066] The model parameters C and A may depend on the amount of hearing loss. The model constant α is a number between 0 and 1, preferably between 0.1 and 0.4, and is assumed in this disclosure to be 0.25 without any preconception.

[0067] Target input signal X t (f) may be equal to or a part of the input media 15. In the latter case, some processing is required to separate the target signal from the input signal. This is briefly explained below.

[0068] Loudness model calibration The model parameters C and A0 (representing 0 dB of hearing loss) can be calibrated as follows: By definition, a 1 kHz tone at 40 dB SPL (40 phons) has a loudness of 1 son for a listener with normal hearing.

[0069]

number

[0070] H OME 2 This represents the squared amplitude transfer function of the outer and middle ear. The second data point for calibration is a 1 kHz tone at 2 dB SPL (2 phons) with a loudness of 0.003 sones for a listener with normal hearing (e.g., hearing loss (HL) = 0 dB). Thus, for a frequency of 1 kHz, the following equation is obtained:

[0071]

number

[0072] Assumption A0≪H OME 2 10 40 / 10 In this case, C can be approximated as follows:

[0073]

number

[0074] Similarly, A0≫H OME 2 10 2 / 10 Assuming this, A0 can be approximated as follows: 0.003 ≈ C{H OME 2 (1kHz)10 2 / 10 αA0 α-1} therefore,

[0075]

number

[0076] A typical value at 1kHz is H OME 2 =0.79 and α=0.25, therefore C≈0.1 and A0≈22.8.

[0077] To obtain an A0 value that reflects differences in sensitivity and / or hearing loss across frequencies, set A to a different value. ΔT By adjusting this, it may reflect the change in signal level ΔT required to achieve the same loudness as a 2-phone signal (0.003 sones):

[0078]

number

[0079] Calculation of partially specified loudness Excitation pattern E tThe loudness of a target signal can be significantly affected by the presence of other masking signals. The loudness of a target signal in such a context is called partial loudness. Two types of masking signals can be distinguished, and excitation patterns can be calculated for both of these masking signal types.

[0080] The first type is an acoustic environment where listeners are present. Acoustic environment X e (f) Corresponding ambient sound excitation pattern E e It can be calculated according to the following:

[0081]

number

[0082] Here H hearable (f) represents the passive or active attenuation frequency response (e.g., acoustic isolation data / playback device data 26) caused by the hearable 22 located above the ear or inside the ear canal. In the case of an active frequency response caused by an active filter, this may depend on the operating mode of the hearable 22. For example, if acoustic noise cancellation (ANC) is active, more attenuation will occur at lower frequencies as a result of the ANC operation.

[0083] The second type relates to components present in the input media signal 15 that are thought to have masking contributions, such as background music and effects. These components X m (f) Corresponding excitation pattern E m It can be calculated according to the following:

[0084]

number

[0085] In one embodiment, for example, an artificial intelligence-based algorithm is used to convert the input signal 15 to the media target signal X t Media background signal X m Source separation techniques are used to separate the sound sources.

[0086] Masking signal X e (f) and X m (f) and each excitation pattern E e and E m Target signal X when present t (f) Loudness N' (t,m,e) Furthermore, the hearing sensitivity (or hearing loss) adjustment ΔT is expressed as follows:

[0087]

number

[0088] optimization One potential optimization criterion for masking and compensating for hearing loss is to adjust the time and frequency-dependent gain of the target signal X. t Applying this to (f), the predicted partial loudness is equal to the partial loudness in the absence of hearing loss and masking effects. If the signal gain is G t (f) For example, the partial specific loudness in a context (e.g., with hearing loss and masking effects) is:

[0089]

number

[0090] Without context, and without contributions such as masking and / or hearing loss, the predicted loudness is determined by equation 2, where A = A0.

[0091]

number

[0092] G t The value of is N' t =N' (t,m,e) It can be calculated by finding:

[0093]

number

[0094] In some implementations, a user-controllable loudness adjustment coefficient λ may be included in the loudness adjustment data 30. In that case, the predicted loudness is multiplied by the coefficient λ.

[0095]

number

[0096] In this case, G t The values ​​will be as follows:

[0097]

number

[0098] It should be noted that signal gain can be calculated and applied in two or more frequency bands. Furthermore, signal gain can be calculated and applied in a time-varying manner.

[0099] Figure 3 provides a schematic overview of the above approach in block diagram form, and shows the processing in the media audio processing module 23 in more detail. As described above, the media audio processing module 23 receives acoustic separation data / playback device data 26 (including hearable operating mode information), playback level data 27, ambient audio data 25, hearing loss data 28, optional loudness adjustment data 30, and input media data 15, and generates a processed media signal 16 as an output.

[0100] The playback device data 26 shows the attenuation response H hearable(f) is determined by the attenuation response calculation block 31, which receives the result.

[0101] The hearing loss data 28 is parameter A ΔT This is received by the parameter calculation block 32, which calculates the value.

[0102] In the illustrated example, the media input audio 15 is received by the source isolation block 33, and the media target signal X t (f) Media background signal X m (f) (if desired). The input audio may include a set of audio components (streams or objects), in which case the media target signal includes a set of target components and the media background signal includes a set of background components. The source separation block may be configured to perform separation using a well-trained neural network or based on the characteristics of the audio components / objects, e.g., position, orientation. Alternatively, block 33 may be connected to receive some type of descriptive data indicating, for example, whether a particular audio component is a target component or a background component, or whether an audio component is related to speech content or other content. In yet another implementation, information about head orientation and / or gaze used may be used to distinguish between target audio components and background audio components.

[0103] Excitation calculation blocks 34, 35, and 36 use ambient audio 25 and an optional media background signal X. m (f) Media target signal X t Each receives (f) and the respective excitation pattern E e , E m and E tThe following is calculated. Each optimized gain is derived from various calculated excitation patterns and arbitrarily selected loudness adjustment data and applied to the media target signal. If a background media signal is available, it is mixed with the (amplified) target signal to produce the output processed media signal.

[0104] Excitation patterns Ee, Em, Et and parameter A ΔT The signal gain G t Gain G is received by gain calculation block 37, which calculates it according to the above formula. t The gain application block 38 applies the target signal X t (f) applies.

[0105] In the illustrated example, the input audio is separated into target and background, with background signal X m (f) is added to the gain-adjusted target signal at the addition point 39 in order to provide the final output audio 16.

[0106] Anticipated use cases Hearing loss compensation in quiet conditions In this use case, the goal is to apply hearing loss compensation, assuming the listener is in a relatively quiet environment. One example might be a listener listening to music or a podcast that is played in a personalized manner by taking into account the listener's hearing loss. The masking effect of ambient audio is assumed to be small or can be disabled entirely to reduce processing power, thus improving battery life if part of the processing is performed on a battery-powered device. The processing device may provide means for the listener to input loudness adjustment data so that the perceived loudness can be adjusted in a more natural way than simply adjusting the media playback level.

[0107] Personalized media optimization in noisy environments In noisy environments, media playback can be improved and personalized by compensating for the listener's hearing loss and the masking effect of the acoustic environment. Assuming the listener is wearing a hearable device, for example, in the form factor of headphones or earphones, the resulting acoustic separation or attenuation from the device (e.g., acoustic separation data) is taken into consideration in the optimization process. Furthermore, if the hearable device has a so-called ANC mode or pass-through mode, the acoustic attenuation response is appropriately adjusted to represent the combined effect of placing the device in or around the ear and its effect on passing ambient audio into the ear canal. Finally, the user can adjust the playback loudness by inputting loudness adjustment data.

[0108] Improvements to personalized, context-aware dialogues Dialogue present within media content can be embedded in media background noise and / or partially masked by ambient audio. To overcome this problem, the proposed solution separates the dialogue (media target audio signal) from the background (media background audio). The media target audio signal is then optimized to compensate for one or more of the following factors: (1) the listener's hearing loss, (2) the masking effect of media background audio, and (3) ambient audio (if any). Thus, when hearing loss is taken into account, the dialogue enhancement function becomes personalized and context-aware.

[0109] Personalized Context-Aware Augmented Reality One of the challenges of augmented reality is loudness management; the perceived loudness of virtual (augmented) content needs to be loud enough to be heard in noisy environments, but not so loud as to be unpleasant. By associating the augmented reality signal with a target signal and taking into account ambient audio and hearing loss, the loudness of augmented reality components can be managed entirely automatically without requiring user interaction.

[0110] Directional, interactive, personalized, context-aware object enhancement In some use cases, it is desirable to improve the audibility of audio components that have an intended direction or position within the listener's field of view. In such use cases, separation of target media signals from background signals can be applied by determining whether or not the target media signal is within the field of view. One embodiment of such a process is to use object-based audio coordinates in combination with head tracking to determine which objects are within the field of view. Field-of-view signal components are optimized based on the assumption that all objects outside the field of view are considered background audio components. With listener head tracking, the listener can indicate which objects need to be optimized simply by looking towards those objects. Thus, this enhancement is directional, interactive, context-aware, and personalized.

[0111] In some use cases, in contrast to the previous example, objects outside the field of view may be amplified so that they become clearly audible when masking occurs due to background media and / or ambient audio. Such amplified objects are particularly relevant to warning signals occurring outside the field of view.

[0112] Implementation in VR applications Specific use cases relate to VR (AR, XR, etc.) applications where audio is rendered using head-tracking data to simulate the effect of movement within an audio landscape. Such applications may feature an "accessibility user interface" that allows users to provide input intended to eliminate, or at least reduce, the effects of various disabilities, such as hearing impairment.

[0113] The accessibility user interface can provide acoustic scene adjustments such as early reflection, reverberation and distance acoustic effects, as well as directional focus effects intended to attenuate distracting sounds from outside the spatial area of ​​interest. All of these effects are generated in response to the user's 6 DoF (6 degrees of freedom) posture and / or gaze in the VR / AR / XR world.

[0114] The drawback of this approach is that users must independently and selectively adjust each effect to achieve the desired overall effect. Some knowledge of acoustic effects may also be required to properly adjust the parameters associated with specific effects. This may be undesirable for most users with hearing impairments. By implementing the techniques described above, it is possible to improve the user experience for all users, regardless of whether or not they have received a proper hearing loss diagnosis.

[0115] The auditory optimization gain (e.g., Gt as described above) can be estimated based on, for example, the audiogram, signal level information, and playback level information. In virtual reality applications, when a user explores a virtual space with 6 degrees of freedom (6 DoF) movement, acoustic effects such as distance attenuation, occlusion, early reflection, reverberation, and diffraction are introduced into the input signal. This can affect the overall signal level.

[0116] The foreground / target signal content (Xt) can be defined as the original input signal content or the combined input content and the resulting sound effects. In some cases, it may be necessary to retain the input signal content as the foreground / target signal, but the sound effects may be reduced to improve clarity. Alternatively, the sound effects themselves can be classified as background signals.

[0117] Background signals (Xm), such as background music, may also be provided as part of the content creator's intent, as an input audio stream. Several foreground / target and background signals may be specified. Classifying audio streams as foreground or background signals can potentially improve the calculation of auditory optimization gains. Depending on the signal levels of both the background and target signals, the gain is adjusted to ensure that the target signal is audible to the hearing impaired when the background is present.

[0118] Alternatively, the audio stream may already be the foreground / target signal itself, mixed with the background signal. In augmented reality applications, this could be an audio signal coming from a real sound source mixed with background noise or the environment. In this case, source separation or denoising may need to be performed beforehand. An additional audio stream containing the extracted background noise can be added as input to an audio renderer, such as an MPEG-I audio renderer. This stream may have the same pause as its mixed audio stream. In some cases, it may be necessary to provide the “denoised” audio stream itself, or the renderer can perform the denoising operation if “denoising” filter coefficients, such as frequency-domain Wiener filtering, are provided.

[0119] Figure 4 shows an audio renderer 41 (such as an MPEG-I audio renderer) that takes a set of audio streams 42 and a 6 DoF head pose and / or gaze 43 as input and provides rendered audio 44 as output. The audio renderer 41 here comprises an auditory optimization rendering stage 45, which takes an audiogram 46 and a playback level P(f) 47 as input to apply signal gain according to this disclosure. The auditory optimization rendering stage may implement a media audio processing module 23, for example, as shown in Figure 3.

[0120] Based on the above example, the following parameters may be useful in a system implementing the techniques disclosed herein (for example, an MPEG-I audio renderer).

[0121] The frequency-dependent absolute playback level P(f) in dB SPL (e.g., default: 70 dB SPL) or relative playback level in dB (e.g., default: 90 dB) is expressed as a scalar or array. This is labeled 27 in Figure 3.

[0122] An audiogram is an array of frequency-dependent hearing loss levels in dB (e.g., derived from demographic data such as age and sex). This is labeled 28 in Figure 3.

[0123] A flag indicating whether the target signal audio stream is considered a foreground / target [true] or background signal [false] (e.g., default: true). This is used as descriptive data in block 33 of Figure 3.

[0124] It should be noted that the playback level P(f) may come from the playback device (such as the current volume setting combined with sensitivity information like the frequency response of headphones), or from a separate bitstream (e.g., an MPEG-I bitstream) if provided as a suggested or required playback level. If playback level data is not available, the system may use common or appropriate default playback level data. For example, the system may assume that a digital signal level with a linear amplitude of 1.0 re digital full scale corresponds to 90 dB SPL in the acoustic domain.

[0125] An interesting aspect of applying this disclosure to a renderer (e.g., audio renderer 41) that has access to individual audio sources / streams 42 is that it allows content creators to classify certain audio sources as, for example, foreground sources or background sources. This can be done, for example, by utilizing parameters that indicate which audio sources should be preferred (e.g., not subject to "culling").

[0126] For example, in future versions of MPEG-I, it is expected that the "noCulling" parameter, which is appended to each audio stream, will be specified as one of the MPEG-I parameters. If this is set to "1" or "true", it indicates that the audio stream will never be culled, and this may be used to specify the target signal flag. Setting the parameter to "1" may be interpreted as indicating that the audio stream is the foreground or target signal.

[0127] If there is no information to identify an audio stream as "target" or "background," input audio can be separated into these types based on the current head posture and / or gaze. For example, if a user is looking in a particular direction, this is likely to be where the user's focus is, and audio from this direction is likely to be "target" audio.

[0128] Future versions of MPEG-I are also expected to include a loudness user interface as an interface for applying loudness-based audio culling. Its purpose is to deliver existing MPEG-D / H loudness metadata to the MPEG-I renderer, enabling audio culling based on absolute and relative loudness criteria. Such loudness information may be specified by the content creator and may be taken as additional input to the rendering stage labeled 48 in Figure 4.

[0129] In addition, additional processing blocks (e.g., “auditory-optimized rendering stages”) are provided within a system implementing the techniques disclosed herein. For example, an implementation of MPEG-I as shown in Figure 4 includes a processing block that enables optimized rendering. This rendering stage is responsible for real-time, frame-by-frame adjustment of time-frequency-dependent auditory-optimized gains based on the input parameters described above.

[0130] Future versions of MPEG-I are expected to include user-specified "accessibility equalization" spectral compensation. Such spectral compensation aims to synthesize user-specified filters (in renderer framework local configuration parameters) and then perform filtering on the binauralized audio output within the "binaural spatializer" stage. This can handle both diplopia and binocular separation spectral compensation. Such processing provides an opportunity to perform the auditory optimization rendering stage 45.

[0131] Figure 5 shows an overview of the rendering process, including an example of an audio renderer 41 implemented here as an MPEG-I renderer. The MPEG-I renderer 41 typically operates at a global sampling frequency of 48 kHz. Figure 5 shows how the renderer 41 in this example is connected to an audio element bitstream 51 encoded via an MPEG-H 3 DA decoder 52.

[0132] All audio elements input to the renderer (e.g., channels, objects, HOAs) have corresponding elements in the MPEG-I immersive audio standard, i.e., so-called source types.

[0133] —Object sources are provided with VR / AR-specific properties.

[0134] —The channel source is played back in the virtual world via a virtual loudspeaker setup.

[0135] HOA sources can be rendered into the virtual world in two different ways: individually with 3 degrees of freedom (user orientation), or in groups of 6 degrees of freedom.

[0136] For all three paradigms, the encoded waveforms can be directly carried over from MPEG-H 3D audio to MPEG-I immersive audio without the need for re-encoding and associated quality loss.

[0137] The decoded audio is rendered along with an MPEG-I bitstream 53. The MPEG-I bitstream 53 carries the audio scene description and other metadata used by the renderer 41. The renderer 41 also has interfaces for accessing consumption environment information (e.g., LSDF) 54, scene updates during playback 55, and user position and interaction information 56. Inputs 54, 55, and 56 are sometimes referred to as audio scene information.

[0138] Renderer 41 enables real-time audibility of complex 6 DoF audio scenes, allowing users to directly interact with entities within the scene. To achieve this, a multi-threaded software architecture is divided into several workflows and components. A block diagram with all renderer components is shown in Figure 6. Renderer 41 supports rendering of VR and AR scenes. For VR and AR scenes, rendering metadata and audio scene information are obtained from bitstream 53. For AR scenes, listening space information is obtained as an LSDF (Listener Space Description Format) file 55 during playback. The components in the diagram are briefly described below.

[0139] The control workflow 61 is the entry point for the renderer 41 and is responsible for interface with external systems and components. Its main functions are incorporated into the scene controller component 62, which coordinates the state of all entities in the 6 DoF scene and implements the interactive interface of the renderer 41. The scene controller 62 supports external updates of mutable properties of scene objects, as well as reading and parsing the LSDF file 55 to complete the information in the bitstream. The scene controller 62 also tracks timeor position-dependent properties of scene objects (e.g., interpolated position or listener proximity conditions).

[0140] Scene State 63 always reflects the current state of all scene objects, including audio elements, transforms / anchors, and geometry. Other components of the renderer can subscribe to changes in Scene State 63. Before rendering begins, all objects in the entire scene are created, and their metadata is updated to reflect the desired scene configuration at the start of playback.

[0141] The stream manager 64 provides a unified interface for renderer components to access the audio stream 42 associated with the audio element of the scene state 63. The audio stream 42 is a (decoded) PCM float sample. The source of the audio stream 42 may be, for example, a decoded MPEG-H audio stream 51 or locally captured audio.

[0142] Clock 65 provides an interface for the renderer component to obtain the current scene time in seconds. The clock input may be, for example, a synchronization signal from another subsystem or the renderer's internal wall clock. The clock input to the scene is not related to audio synchronization.

[0143] The rendering workflow 66 generates a PCM float audio output signal 44. This is isolated from the control workflow 61, and only the scene state 63 (for communicating any changes in the 6 DoF scene) and the stream manager 64 (for providing the input audio stream) are accessible from the rendering workflow 66 for communication between the two workflows.

[0144] The renderer pipeline 67 makes the input audio stream 42, provided by the stream manager, audible based on the current scene state. Rendering is organized in a sequential pipeline so that each rendering stage implements an independent perceptual effect and utilizes the processing of preceding and succeeding stages.

[0145] The spatializer 68 terminates the renderer pipeline 67 and makes the output of the renderer stage audible into a single output audio stream suitable for the desired playback method (e.g., binaural or adaptive loudspeaker rendering). Finally, the limiter 69 provides clipping protection to the audible output signal.

[0146] Figure 7 shows the renderer pipeline 67, where each box represents a separate rendering stage. Rendering stages are instantiated during renderer initialization. The rendering stages are computed in the sequence shown in the figure.

[0147] The auditory optimization rendering stage 45 (see Figure 4) benefits from knowing the classification of individual sound sources (e.g., from the noCulling flag), so the auditory optimization rendering stage 45 may be implemented within the renderer pipeline 67, immediately before the spatializer 68 (binaural or loudspeaker), for example, at the end of the renderer pipeline after the MP-HOA rendering stage in Figure 7. Note that at this stage, all 6 DoF acoustic effects in VR / AR / XR applications (occlusion, diffraction, reflection, reverberation, etc.) are available in the form of rendering items to be fed to the spatializer 68.

[0148] Figure 8 is a flowchart illustrating a method for optimizing the user experience on an electronic device for individuals with hearing impairments, according to several embodiments. Method 300 is performed on an electronic device (for example, an electronic device as described herein).

[0149] The electronic device receives audio data (e.g., 302). In some embodiments, the audio data includes one or more audio streams. In some embodiments, the audio data includes loudness data. The electronic device obtains hearing loss data associated with a listener (e.g., 304). In some embodiments, the hearing loss data is audiogram data associated with the listener. In some embodiments, the audiogram data is derived from a hearing test associated with the listener. In some embodiments, the audiogram data is derived from demographic information associated with the listener.

[0150] The electronic device receives context data including playback level data (e.g., 306) and calculates gain data (e.g., time-frequency-dependent hearing optimization gain) based on the audio data, hearing loss data, and context data (e.g., 308). In some embodiments, the playback level data is a scalar or array of frequency-dependent absolute playback levels in dB SPL units; in some embodiments, relative playback levels in dB units.

[0151] The electronic device processes at least a portion of the audio data using gain data to generate compensated audio data (e.g., 310).

[0152] In some embodiments, context data includes signal description data. In some embodiments, signal description data indicates whether each audio data corresponds to a foreground signal or a background signal. In some embodiments, signal description data indicates whether the foreground signal is an audio signal or a non-audio signal.

[0153] In some embodiments including signal description data, the electronic device determines that the value associated with the signal description data is either a first value or a second value (e.g., true or false), and according to the determination that the value associated with the signal description data is a first value (e.g., true), the gain data is a first set of gain data, and according to the determination that the value associated with the signal description data is a second value (e.g., false), the gain is a second set of gain data different from the first set of gain data.

[0154] The above description has been made with reference to specific embodiments for illustrative purposes. However, the above exemplary description is not intended to be exhaustive or to limit the invention to the exact form disclosed. In view of the above teachings, many modifications and variations are possible. The embodiments have been selected and described to best illustrate the principles of the art and their practical applications. Those skilled in the art will be able to best utilize the art and its various embodiments by making various modifications suitable for the specific use intended.

[0155] While this disclosure and examples are adequately described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. Such changes and modifications should be understood to fall within the scope of this disclosure and examples as defined by the claims.

[0156] Various aspects and implementations of this disclosure can also be understood from the following listed exemplary embodiments (EEEs) that are not part of the claims.

[0157] EEE1. A method for processing audio for a listener of a playback device, Receiving audio data and Receiving hearing loss data associated with the aforementioned listener, Receiving context data that includes one or more of the following: ambient sound data, sound separation data, and playback level data, Calculating audio level data based on at least a portion of the aforementioned audio data, Calculating gain data based on the aforementioned audio level data and the aforementioned context data, A method comprising processing at least a portion of the audio data using the gain data to generate compensated audio data.

[0158] EEE2. The audio data includes background and target components, and the method of EEE1 includes processing at least a portion of the audio signal, which includes processing the target component with gain data and refraining from processing the background component with gain data.

[0159] EEE3. Before receiving the audio data, determine the target component of the audio data, at least partially based on the direction or position of one or more objects or sources represented within the audio data. The method described in EEE2, further including the method described in EEE2.

[0160] EEE4. Further includes calculating one or more perceptual excitation levels, Calculating audio level data is partially based on at least one of the one or more perceptual excitation levels. The method described in any one of the items EEE1 to EEE3.

[0161] EEE5. The method according to EEE4, wherein calculating the one or more perceptual excitation levels includes calculating an excitation loss function based on the hearing loss data.

[0162] EEE6. Any method from EEE4 to EEE5, wherein the context data includes ambient sound data and the calculation of one or more perceptual excitation levels includes calculating the excitation levels associated with the ambient sound data.

[0163] EEE7. The method according to any one of EEE4 to EEE6, wherein the context data includes acoustic separation data, and the excitation level associated with the ambient sound data is calculated based in part on the acoustic separation data.

[0164] When subject to EEE6.EEE2, the calculation of the one or more perceptual excitation levels is To calculate the excitation level related to the background component of the audio data, The method according to EEE4 to EEE7, comprising calculating the excitation level associated with the target component of the audio data.

[0165] When subject to EEE7.EEE5, calculating the one or more sensory excitation levels is: Applying the excitation loss function to the excitation level associated with the background component, The method according to EEE6, comprising applying the excitation loss function to the excitation level associated with the target component of the audio data.

[0166] EEE8. Receiving loudness adjustment data, The gain is updated based on the loudness adjustment data, Any of the aforementioned methods of the EEE, including the above.

[0167] EEE9. Gain data is calculated and applied in two or more frequency bands using any of the methods of the aforementioned EEE.

[0168] EEE10. One of the aforementioned EEE methods, in which gain data is calculated and applied in a time-varying manner.

[0169] EEE11. Any of the above methods of EEE, including calculating gain data and optimizing the perceived loudness constraint.

[0170] EEE12. Any of the aforementioned EEE methods, wherein the context data includes ambient sound data, and the ambient sound data is derived from one or more microphones on a device in close proximity to the listener.

[0171] EEE13. Hearing loss data is derived from the listener's age or sex at birth, using one of the aforementioned EEE methods.

[0172] EEE14. Input loudness adjustment data is received by the playback device via user input from the listener, using one of the aforementioned EEE methods.

[0173] EEE15. The playback device is a hearable device, A method of any of the aforementioned EEEs, wherein one or more of the following are performed on a hearable device: calculating audio level data, calculating gain data, and processing audio data using the gain data.

[0174] EEE16. Context data includes ambient sound data, and ambient audio data is captured from one or more microphones of a hearable device communicating with a mobile device, in any of the aforementioned EEE methods.

[0175] EEE17. A non-temporary computer-readable storage medium that, when executed by a computing device, stores instructions causing the computing device to perform the actions described in any one of the items EEE1 to EEE16.

[0176] EEE18. Computing device, At least one processor, A computing device comprising: a memory that, when executed by at least one processor, stores instructions causing the computing device to perform the method described in any one of EEE1 to EEE16.

[0177] EEE19. Methods for processing audio, Receiving audio data and To obtain hearing loss data related to listeners, Receiving context data including playback level data, Calculating gain data based on the audio data, the hearing loss data, and the context data, A method comprising processing at least a portion of the audio data using the gain data to generate compensated audio data.

[0178] EEE20. The context data includes signal description data. Determining that the value associated with the signal description data is either a first value or a second value, In accordance with the determination that the value associated with the signal description data is a first value, the gain data is a first set of gain data, The method according to EEE19, wherein, according to the determination that the value associated with the signal description data is a second value, the gain is a second set of gain data different from the first set of gain data.

[0179] EEE21. A non-temporary computer-readable storage medium that, when executed by a computing device, stores instructions that cause the computing device to perform either method EEE19 or EEE20.

[0180] EEE22. Computing device, At least one processor, A computing device, which, when executed by at least one processor, includes memory for storing instructions that cause the computing device to perform either EEE19 or EEE20.

Claims

1. A method for processing audio for a listener on a playback device, Receiving audio data for playback on the aforementioned playback device, wherein the audio data includes a target audio component. Receiving masked audio data, To obtain hearing loss data associated with the aforementioned listener, To obtain playback level data, Calculating a target excitation pattern based on the target audio component and the playback level data, The process involves calculating a masking excitation pattern based on the aforementioned masking audio data, Calculating the signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data, A method comprising applying the signal gain to the target audio component to generate a compensated target audio component.

2. The method according to claim 1, wherein the masking audio data includes a background audio component included in the audio data, and the masking excitation pattern includes a background excitation pattern based on the background audio component and the playback level data.

3. The method of claim 2, further comprising receiving descriptive data indicating whether a particular audio stream of the audio data is a target audio component or a background audio component.

4. The method according to claim 2, further comprising separating the audio data into the target audio component and the background audio component.

5. The method according to claim 4, wherein separating the audio data into the target audio component and the background audio component includes determining the target audio component of the audio data based on the orientation or position of one or more audio objects or audio sources represented in the audio data.

6. The method according to claim 1, wherein the masking audio data includes ambient sound captured by a microphone placed close to the listener, and the masking excitation pattern includes an ambient excitation pattern based on the ambient sound.

7. The method according to claim 6, wherein the playback device includes a set of hearables, and the ambient sound is captured from the microphones of the hearables.

8. Receiving playback device data describing the effect of the playback device on ambient sound, wherein the ambient excitation pattern is also based on the playback device data. The method according to claim 6.

9. The method according to claim 8, wherein the playback device includes a set of hearables having at least one earphone, each earphone including a microphone, a loudspeaker, and an active filter, and the playback device data includes the transfer function of the active filter.

10. The method according to claim 1, wherein the signal gain is calculated and applied in two or more frequency bands.

11. The method according to claim 1, wherein the signal gain is calculated and applied in a time-varying manner.

12. Calculating the aforementioned signal gain is, To calculate the predicted specific loudness of the target excitation pattern without masking and without hearing loss, The partially specified loudness of the target excitation pattern, which is masked by the masking excitation pattern, is calculated based on the hearing loss data. The method according to claim 1, comprising determining the signal gain as the gain of the target audio component that produces a predicted specific loudness approximately equal to the partial specific loudness.

13. The aforementioned signal gain is calculated as the predicted specific loudness as follows: [Math 1] E t This is the target excitation pattern, E m This is a background excitation pattern, E e This is an environmental excitation pattern, A 0 is a model parameter representing a hearing loss of 0 dB, and A ΔT The method according to claim 12, wherein is a hearing loss-dependent model parameter, and α is a constant between 0.1 and 0.

4.

14. Receiving the loudness adjustment coefficient, The signal gain is also calculated based on the aforementioned loudness adjustment coefficient, The method according to claim 1, further comprising:

15. The method according to claim 14, wherein the loudness adjustment coefficient is received by the playback device via user input from the listener.

16. The aforementioned signal gain is calculated as the predicted specific loudness as follows: [Math 2] λ is the loudness adjustment coefficient, and E t is the target excitation pattern, and E m is the background excitation pattern, and E e is the environmental excitation pattern, and A 0 is the model parameter representing a hearing loss of 0 dB, and A ΔT is the hearing loss-dependent model parameter, and α is a constant between 0.1 and 0.

4. The method according to claim 14.

17. The method according to claim 1, wherein the hearing loss data is derived from the age or sex at birth of the listener.

18. The playback device includes a set of hearables, The method according to claim 1, wherein one or more of the following are performed on the hearable: calculating a signal gain and applying the signal gain.

19. A method for audio rendering, Receiving an audio bitstream that includes the target audio component and the background audio component, Receiving audio scene information including the listener's current head posture and / or gaze, Processing the audio scene information in order to obtain the audio scene state, In order to obtain the rendered target audio component and the rendered background audio component, audio data is rendered based on the scene state, To obtain hearing loss data associated with the aforementioned listener, To obtain playback level data, Calculating a target excitation pattern based on the target audio component and the playback level data, The process involves calculating a masking excitation pattern based on the background audio component and the playback level data, Calculating the signal gain based on the target excitation pattern, the masking excitation pattern, and the hearing loss data, A method comprising applying the signal gain to the rendered target audio component to produce a compensated rendered target audio component.

20. The method according to claim 19, further comprising receiving descriptive data indicating whether a particular audio stream of the audio data is a target audio component or a background audio component.

21. The method according to claim 19, wherein at least some of the target audio components and / or background audio components are identified based on the current head posture and / or line of sight.

22. A non-temporary computer-readable storage medium that, when executed by a computing device, stores instructions causing the computing device to perform the method according to any one of claims 1 to 21.

23. A computing device, At least one processor, A computing device comprising: a memory that, when executed by the at least one processor, stores instructions causing the computing device to perform the method according to any one of claims 1 to 21.

24. It is an audio renderer, A decoder for receiving and decoding an audio bitstream and providing input audio including a target audio component and a background audio component, A scene controller configured to receive audio scene information and provide an audio scene state based on the audio scene information, It is a rendering pipeline, To obtain the rendered target audio component and the rendered background audio component, audio data is rendered based on the scene state, We obtain hearing loss data associated with the listener, Obtain playback level data, Based on the target audio component and the playback level data, the target excitation pattern is calculated. Based on the background audio component and the playback level data, a masking excitation pattern is calculated. Based on the target excitation pattern, the masking excitation pattern, and the hearing loss data, the signal gain is calculated. An audio renderer including a pipeline configured to apply the signal gain to the rendered target audio component to produce a compensated rendered target audio component.

25. The audio renderer according to claim 24, wherein the rendering pipeline is further configured to receive descriptive data indicating whether a particular audio stream of the audio data is a target audio component or a background audio component.

26. The audio renderer according to claim 24 or 25, wherein the audio renderer is compatible with the MPEG-I standard.