Personalized ambient sound playback
By generating digital representations of specific external sounds and adjusting processing parameters based on feedback data, the problem of insufficient personalization of audio devices in ambient sound processing is solved, achieving a more natural user experience and more efficient resource utilization.
Patent Information
- Application Number
- CN202480010363.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-01
- Filing Date
- 2024-01-31
- Publication Date
- 2025-09-12
AI Technical Summary
Existing audio devices lack personalized adjustments when processing ambient sounds, resulting in an unnatural user experience, especially inconsistent perception of external sounds due to differences in anatomical details.
By generating a digital representation of specific external sounds, obtaining feedback data, and adjusting frequency-related processing parameters, personalized calibration data is formed to optimize the ambient sound playback of audio devices.
It improves the user's natural perception of ambient sound, enhances the personalized adaptability of audio equipment, and reduces resource consumption and current consumption.
Smart Images

Figure CN120642346A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to audio processing, and more particularly to audio processing of ambient sounds in audio devices. Background Art
[0002] With the likes of With the introduction of portable electronic devices such as the iPhone and later mobile phones, access to various forms of audio has increased dramatically. Audiobooks, favorite songs or interesting podcasts are always within reach. This has led to several innovations in the field of audio devices and sound control. One of the revolutionary innovations is the ability to provide personalized sound for the user of the playback device. This includes adjusting the played sound to compensate for any hearing deviations of the user. Further innovations include controlling the ambient sound so that external noise from, for example, fans or vehicles can be reduced or even eliminated. This is often referred to as active noise cancellation or active noise control (ANC). In addition, some devices also implement advanced equalizers (EQ) that are configured and controlled according to the ambient sound. These equalizers are generally called adaptive equalizers.
[0003] These innovations have improved sound quality and the user listening experience, regardless of the environment they are listening to audio in. However, we can do more, and there is room to further improve the user listening experience. Summary of the Invention
[0004] It is in view of the above considerations and other considerations that the various embodiments of the present disclosure have been made. Therefore, the present disclosure recognizes that there is a need for alternatives (e.g., improvements) to the above-mentioned prior art. The purpose of some embodiments is to solve, alleviate, mitigate or eliminate at least some of the above-mentioned or other shortcomings. The purpose of the present disclosure is to achieve a new type of ambient sound processing, or in other words, to improve the existing technology and eliminate or at least alleviate the shortcomings discussed above. More specifically, the purpose of the present invention is to provide a calibration method for personalized ambient sound. These objects are achieved by the technology set forth in the attached independent claims and the preferred embodiments defined in the dependent claims related thereto.
[0005] In a first aspect, a method for providing personalized ambient sound playback (ASP) calibration data associated with an audio device and a specific user is proposed. The method includes remotely generating a first specific external sound relative to the audio device, and obtaining a digital representation of the first specific external sound by the audio device. The method also includes processing the digital representation of the first specific external sound based on an initial set of frequency-related processing parameters, generating a first internal sound based on the processed digital representation of the first specific external sound by the audio device when the specific user wears the audio device, and obtaining first feedback data indicating the similarity between the first internal sound and the first specific external sound. In addition, the method includes adjusting the initial set of frequency-related processing parameters based on the first feedback data to obtain a personalized set of frequency-related processing parameters, and providing the personalized set of frequency-related processing parameters as ASP calibration data to the audio device when the specific user wears the audio device.
[0006] In one variant, the first specific external sound is an external sound within a first frequency band, and the initial set of frequency-dependent processing parameters is further adjusted based on the first frequency band. This is beneficial because it improves the quality of the ASP calibration data and reduces resource-intensive processing.
[0007] In one variation, the method is repeated for a second specific external sound within a second frequency band, thereby obtaining second feedback data. The personalized set of frequency-dependent processing parameters is further adjusted based on the second feedback data associated with the second specific external sound and the second frequency band. This is beneficial because it improves the quality of the ASP calibration data.
[0008] In one variation, the first specific sound includes frequency content that is also in the second frequency band, and the second specific sound includes frequency content that is also in the first frequency band. This is beneficial because it improves the accuracy of the feedback data.
[0009] In one variation, the first specific sound includes frequency content substantially entirely within a first frequency band, and the second specific sound includes frequency content substantially entirely within a second frequency band.
[0010] In one variation, the first frequency band and the second frequency band are selected from a set of frequency bands including at least two of a sub-bass region, a bass region, a mid-bass region, a mid-mid range region, a mid-treble region, a presence region, and a detail region.
[0011] In one variation, obtaining the first feedback data includes obtaining feedback data from a specific user, which is beneficial because subjective perception constitutes part of the feedback.
[0012] In one variation, feedback data from a specific user is obtained by the specific user indicating feedback data in a two-dimensional space, wherein at least one dimension includes an emotion indicator. This is beneficial because the user can easily provide accurate data, thereby improving the accuracy of the feedback data.
[0013] In one variation, the emotion indicator is configured based on a frequency band associated with an external sound related to the feedback data. This is beneficial because the user can easily provide accurate data, thereby improving the accuracy of the feedback data.
[0014] In one variant, obtaining the first feedback data comprises obtaining feedback data from a feedback microphone circuit of the audio device.This is advantageous because the method or parts of the method can be performed without specific user interaction.
[0015] In one variant, the first feedback data comprises amplitude feedback data indicating the similarity of the sound pressure level SPL between the first internal sound and the first specific external sound. This is beneficial because the volume of the ASP is correct.
[0016] In one variant, the initial set of frequency-dependent processing parameters is also adjusted based on the first feedback data based on one or more equal loudness contours. Equal loudness is known from the work of Fletcher and Munson and ensures that the perceived loudness is correct relative to the set playback level.
[0017] In one variation, the method further includes processing the digital representation of the first external sound based on the personalized set of frequency-related processing parameters; generating, by the audio device, a personalized first internal sound based on the personalized processed digital representation of the first specific external sound when the audio device is worn by a specific user; obtaining updated first feedback data, the updated first feedback data indicating a similarity between the personalized first internal sound and the first specific external sound; and adjusting the personalized set of frequency-related processing parameters based on the updated first feedback data.
[0018] In one variation, the initial set of frequency-dependent processing parameters is based on one or more calibration input parameters, wherein the calibration input parameters are one or more of: a wearing state of the audio device and / or a relative position of a sound generator configured to generate a first specific external sound.
[0019] In one variant, the personalized frequency-dependent processing parameters and the ASP calibration data are configured with a limited bandwidth, preferably corresponding to the human hearing bandwidth. This improves processing efficiency and reduces, for example, current consumption.
[0020] In one variant, the personalized frequency-dependent processing parameters and ASP calibration data are set to unity, so that no processing is performed at frequencies below 20 Hz, preferably no processing is performed at frequencies below 50 Hz, and most preferably no processing is performed at frequencies below 70 Hz. This improves processing efficiency and reduces, for example, current consumption.
[0021] In one variant, the personalized frequency-dependent processing parameters and ASP calibration data are set to 1 so that no processing is performed at frequencies above 20 kHz, preferably no processing is performed at frequencies above 15 kHz, and most preferably no processing is performed at frequencies above 12 kHz. This improves processing efficiency and reduces, for example, current consumption.
[0022] In one embodiment, the first specific external sound is a predefined sound selected from a set of sounds comprising a plurality of sounds, wherein at least one of the sounds is suitable for determining ASP calibration data associated with at least one frequency band selected from the following regions: a bass region, a mid-bass region, a mid-mid-range region, a mid-treble region, a presence region and / or a detail region.
[0023] In a second aspect, an ASP calibration system is provided. The ASP calibration system includes an audio device, a sound generator, a feedback pre-matching circuit, and at least one processor circuit configured to pre-match ASP calibration data for a specific user and the audio device according to the method of the first aspect. The audio device includes a feedforward microphone circuit configured to acquire a digital representation of a specific external sound generated by the sound generator; a transducer circuit; an input circuit configured to acquire audio data; and a processor circuit configured to process the digital representation of the specific external sound based on the ASP calibration data and play the processed specific external sound and audio data through the transducer circuit.
[0024] In a third aspect, an audio device is proposed. The audio device is configured to form part of the ASP calibration system of the second aspect, thereby obtaining ASP calibration data for a specific user and the audio device according to the method of the first aspect. The audio device includes: a feedforward microphone circuit, which is configured to obtain a digital representation of external sound; a transducer circuit; and a processor circuit. In one embodiment, the processor device is configured to process the digital representation of the external sound based on the ASP calibration data and broadcast the processed external sound through the transducer circuit. Preferably, the processor circuit is also configured to process the external sound based on the hearing profile of the specific user. In one embodiment, the audio device also includes an input circuit configured to obtain audio data through an audio interface. The processor circuit is configured to broadcast the audio data through the transducer circuit. Preferably, the processor circuit is also configured to process the audio data based on the hearing profile of the specific user. In a fourth aspect, a computer-readable storage medium is proposed. The computer-readable storage medium includes program instructions, which, when executed by the processor circuit, cause the processor circuit to perform the method according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Embodiments of the invention will be described hereinafter; reference is made to the accompanying schematic drawings which show non-limiting examples of how the inventive concept may be put into practice.
[0026] Figure 1 is a schematic diagram of an audio device according to some embodiments of the present disclosure;
[0027] Figures 2a to 2c is a cross-sectional view of an ear of a user wearing an audio device according to some embodiments of the present disclosure;
[0028] Figure 3 is a schematic diagram of an audio device according to some embodiments of the present disclosure;
[0029] Figure 4 is a schematic diagram of an ASP calibration system according to some embodiments of the present disclosure;
[0030] Figures 5a to 5d is a schematic diagram of an ASP calibration system according to some embodiments of the present disclosure;
[0031] Figure 6 is a simplified signaling diagram according to some embodiments of the present disclosure;
[0032] Figure 7 is a schematic diagram of a method for providing ASP calibration data according to some embodiments of the present disclosure;
[0033] Figure 8is a schematic diagram of a method for providing ASP calibration data according to some embodiments of the present disclosure;
[0034] Figure 9 is a schematic flow structure for providing ASP calibration data according to some embodiments of the present disclosure;
[0035] Figure 10 is a schematic diagram of an audio device according to some embodiments of the present disclosure;
[0036] Figure 11 is a schematic diagram of an audio device according to some embodiments of the present disclosure;
[0037] Figures 12a to 12c is a graph showing frequency content and frequency bands for specific external sounds according to some embodiments of the present disclosure;
[0038] Figure 13 is a view of a two-dimensional space for providing feedback data according to some embodiments of the present disclosure;
[0039] Figure 14 is a schematic diagram of a computer program and a computer-readable storage medium according to some embodiments of the present disclosure; and
[0040] Figure 15 is a schematic diagram of the loadability of a computer-readable storage medium according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0041] Certain embodiments will be described more fully below with reference to the accompanying drawings. However, the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of the invention (as defined in the appended claims) to those skilled in the art.
[0042] The term "coupled" is defined as connected, although not necessarily directly connected, nor necessarily mechanically connected. Similarly, the term "connected" or "operably connected" is defined as connected, although not necessarily directly connected, nor necessarily mechanically connected. Two or more "coupled" or "connected" items can become one with each other. Unless otherwise clearly required by the present disclosure, the terms "a" and "an" are defined as one or more. The terms "substantially," "roughly," and "approximately" are defined to be largely, but not necessarily fully specified, as understood by those of ordinary skill in the art. The terms "comprise" (and any form thereof), "have" (and any form thereof), "include" (and any form thereof), and "contain" (and any form thereof) are open linking verbs. Therefore, a method that "comprises," "has," "includes," or "contains" one or more steps has the one or more steps, but is not limited to having only the one or more steps.
[0043] Figure 1 1 shows a simplified view of an embodiment of the audio device 10 as worn by a user 40. Figure 1 , the audio device 10 is shown as a pair of ear-hook headphones. The audio device 10 of this embodiment is worn on the outer ear 41 of the user 40. As will be seen in other parts of this disclosure, this is merely an example and the teachings of this disclosure are applicable to many forms of audio devices 10, such as but not limited to on-ear, surround or in-ear types. For the present disclosure, the audio device 10 generally represents any device that can be configured to generate sound from a sound generator 220 and attract at least one ear of the user, and thereby at least partially block or change to some extent the user's perception of ambient sound (the latter will be described in more detail in later sections). The sound generator 220 can be any suitable sound generator 220 that is operably or directly connected to the audio device 10. In some embodiments, the sound generator 220 can be included in the audio device 10 (or vice versa). Figure 1 The sound generator 220 in the embodiment of the present invention is shown as an electronic device in the form of a mobile phone 20, but it can be any suitable device, such as, but not limited to, a home audio system, a portable media storage (e.g., an iPod, a portable MP3 player), a car audio system, a portable speaker device, etc. The sound generator 220 can be connected to the audio device 10 via any suitable audio interface 30. In some embodiments, the audio interface 30 is a wired interface, such as a cable connected to the audio device 10, and can be connected to the sound generator 220 via a 3.5mm or 6.6mm phono plug. Advantageously, the audio interface 30 is a wireless interface, such as, for example, Bluetooth, WIFI, a 3GPP specified interface, or a suitable proprietary ISM interface.
[0044] As mentioned above, Figure 1The audio device 10 is an ear-hook device. Figure 2a As shown in the cross-sectional view of the user 40 and the earhook audio device, a cavity C is formed between the eardrum 43 of the user 40 and the audio device 10. Depending on the type of the audio device 10 and the fit of the audio device 10, the size and acoustic characteristics (e.g., open / closed) of the cavity C may vary. Figure 2a In the case where the audio device is an ear-hook device, cavity C is relatively large and includes the outer ear 41 and ear canal 42 of the user 40. Depending on the size of the audio device 10 relative to the outer ear 41, cavity C will be open or closed. If cavity C is open, it is in fluid communication with the exterior O of the audio device 10. It should be noted that, additionally or alternatively, cavity C can be in fluid communication with the exterior O through the audio device 10, thereby forming an open cavity C regardless of the fit of the audio device 10.
[0045] exist Figure 2b , another exemplary embodiment of an audio device 10 is shown. In this embodiment, the audio device 10 is an in-ear audio device that is arranged in the outer ear 41 of the user 40. This type of audio device 10 can be called an earplug and generally specifically rests on the concha, that is, the opening of the outer ear 41, which is connected to the ear canal 42. In this embodiment, the cavity C formed between the audio device 10 and the eardrum 43 of the user 40 only includes a portion of the outer ear 41 (a portion of the concha) and the entire ear canal 42 between the outer ear 41 and the eardrum 43. Figure 2b The cavity C is smaller than Figure 2a The cavity C is formed by the audio device 10. Generally, the cavity C is considered open because it is challenging to achieve a tight fit of the audio device 10 at the concha, and an air gap is generally formed between the audio device 10 and the concha.
[0046] exist Figure 2c , yet another exemplary embodiment of an audio device 10 is shown. In this embodiment, the audio device 10 is an in-ear audio device that is disposed within the ear canal 43 of a user 40. This type of audio device may be referred to as an earphone and is typically squeezed into the ear canal 43, thereby forming a tight fit between the audio device 10 and the ear canal 43. Thus, the cavity C formed by the audio device only includes a portion of the ear canal 43 and is smaller than the reference ear canal. Figure 2a and Figure 2b Due to the tight fit, the cavity C provided by the audio device 10 is generally considered to be a closed cavity, although a breathing valve or the like is generally introduced to provide greater comfort to the user 40.
[0047] refer to Figures 2a to 2cThe embodiments of the audio device 10 presented are non-exhaustive examples of audio devices 10 to which the teachings of the present disclosure may be applicable. As described above, each audio device 10 will form a specific cavity C at the ear of the user 40. Therefore, sounds originating from the exterior O of the audio device 10, i.e., ambient or external sounds that are generally not generated by the audio device 10, will be affected by the shielding provided by the audio device 10 before reaching the eardrum 43 of the user 40. When the user 40 wears the audio device 10, the external sounds may be suppressed, blocked, distorted, or otherwise affected. To alleviate this situation, many audio devices 10 are configured with hear-through, ambient sound functions, or ambient sound playback (ASP), by which the audio device 10 is configured to actively transmit sound from the exterior O to the cavity C, i.e., to the eardrum 43 of the user 40. However, as previously described, the size and form of the cavity C will affect the sound at the cavity, and as mentioned, the size and form of the cavity C will depend on, for example, the fit of the audio device 10. This is a problem that the inventors of the present disclosure have identified, and the teachings presented herein will enable users to specifically adapt and personalize hear-through, ambient sound functionality, or ASP.
[0048] For efficiency, the sound at or originating from the exterior O of the audio device 10 may be referred to as external sound Se, and the sound at or originating from the cavity C may be referred to as internal sound Si.
[0049] As is generally known, and as Figure 3 As schematically shown, the audio device 10 includes one or more transducer circuits 12 operably connected to a processor circuit 100. The transducer circuit 12 of the audio device 10 is configured to generate sound that is propagated into a cavity C formed between the audio device 10 and the eardrum 43 of the user 40. The processor circuit 100 can be implemented as any circuit between an impedance matching circuit and an advanced DSP-based circuit that is configured to control the audio provided to the transducer circuit 12. For the purposes of this disclosure, it is assumed that the audio device 10 includes or is operably connected to a processor circuit 100 that is configured to perform or cause the performance of the teachings set forth herein. The processor circuit 100 may also include or be operably connected to an input circuit 110 that is configured to interface with, for example, a sound generator 220 via an audio interface 30. Specifically, the input circuit 110 is configured to obtain audio data 112 from an audio source (see Figure 10 ). The audio source will not be explained further and those skilled in the art understand that the audio source may depend on the audio interface 30 and may span sources such as a Walkman and streaming content from, for example, Spotify or YouTube. It should be mentioned that although Figure 3The processor circuit 100 shown includes an input circuit 110 , but this is only a non-limiting example, and the processor circuit 100 and the input circuit 110 may also be independent circuits.
[0050] Figure 3 The audio device 10 also includes at least one microphone circuit 14, 16. The microphone circuits 14, 16 can be any form of audio / sound sensing circuitry. At least one microphone circuit 14, 16 is a feedforward microphone circuit 14. The feedforward microphone circuit 14 is advantageously configured to acquire, measure, or otherwise obtain an indication of sound at an exterior O of the audio device 10. The feedforward microphone circuit 14 is typically provided in an audio device 10 that is configured for use with, for example, a mobile phone 20, because the microphone 14 is configured to acquire speech from a user during, for example, hands-free operation of the mobile phone 20. In addition to the above, as an optional feature, the audio device 10 may include a feedback microphone circuit 16. The feedback microphone circuit 16 is advantageously configured to acquire, measure, or otherwise obtain an indication of sound at a cavity C between the audio device 10 and the eardrum 43 of the user 40. The feedback microphone circuit 16 is generally provided in an audio device 10 configured to perform active noise cancellation / control (ANC) to provide feedback of the amount of noise remaining in a cavity C formed between the audio device 10 and the eardrum 43 of the user 40 .
[0051] It should be understood by those skilled in the art that Figure 3 The schematic diagrams of the audio devices presented herein may not be complete, and further hardware and / or software components, modules, circuits, or devices may be required to provide a fully operational audio device. For simplicity of disclosure, these features, such as analog-to-digital converters, digital-to-analog converters, amplifiers, transceivers, etc., are not described in further detail in this disclosure as they are well known to those skilled in the art.
[0052] It should be mentioned that although in the foregoing, the microphone circuits 14, 16 are shown and described as being included in the audio device 10, one or more or all of the microphone circuits 14, 16 may be separate from the audio device 10 and operatively connected to the audio device, for example via the audio interface 30.
[0053] Generally, when the audio device 10 (e.g., Figure 3 When the audio device 10) operates in ASP mode, the feedforward microphone circuit 14 can be configured to be external O( Figure 3The external sound Se is acquired (recorded, measured, sensed) at the user 40 (not shown). The acquired external sound Se is then broadcast through the transducer circuit 12 to provide the internal sound Si at the cavity C formed by the audio device 10 at the ear of the user 40. The acquired external sound Se can be processed by, for example, the processor circuit 100 to compensate for, for example, damping and / or occlusion effects of the audio device 10.
[0054] However, due to the wide variations in how sound is perceived and transmitted from the exterior O of the audio device 10 to the cavity C (e.g., the interior of the audio device 10), obtaining appropriate processing parameters for processing the external sound Se is very challenging. As previously mentioned, the transmission of the external sound Se is affected by various factors, such as the fit of the audio device 10, the type of external sound Se, etc. The inventors of the present disclosure have recognized that a personalized process is required when adjusting the ASP.
[0055] As previously described, for a particular user 40, the perception of external sound Se depends on the anatomical details (size, geometry, etc.) of the individual's upper torso and ears (external ear 41 and ear canal 42). Specifically, these anatomical details may depend on the size of the user's 40 head and shoulders, the geometry of the external ear, the diameter and depth of the ear canal, etc. In some examples, the ASP may be provided with a default configuration (non-anthropomorphic) and tuned using an acoustic device that models humans using anatomical details established by averaging over a large group of humans. Therefore, such a non-anthropomorphic ASP may not sound very natural because the configuration is not suitable for a particular user 40, for example, the particular user 40 has a particular ear and upper torso size and geometry, headphone placement and / or fit that is different than those of the acoustic device.
[0056] The inventors have realised that personalisation of the ASP will result in a specific user 40 perceiving a more natural external sound Se, ie ambient sound.
[0057] For purposes of explanation and throughout this disclosure, a natural-sounding ASP is an ASP in which the difference between the perception of external sound Se perceived when wearing an audio device 10 with an anthropomorphic ASP and the perception of the same external sound Se perceived when not wearing any audio device 10 is relatively small. That is, comparing an audio device 10 without personalized ASP activation (hear-through) to an audio device 10 with ASP activation, the difference in perception of external sound Se when wearing the audio device 10 and when not wearing the audio device 10 is reduced when the personalized ASP is activated. To provide this, the ASP needs to be personalized, and ASP calibration data is required for each specific user 40. One method to provide this is to place the user 40 in an anechoic chamber with a microphone inserted into the ear canal 42. The user 40 may then be exposed to a variety of sounds from a sound source, which can be detected by the microphone in the ear canal. For example, such sounds can be pure-tone sinusoidal test signals. In such an environment, the user's 40's open-ear frequency transfer function, known as the open-ear frequency transfer function, can be accurately acquired. This constitutes a reference frequency response for the open ear. However, this implementation has many technical implications, such as the depth of the microphone in the ear canal 43, the characteristics of the sound source (a diffuse sound field of noise or a point source with chirping), source impact cancellation, microphone frequency response, etc. In addition, this approach is cumbersome and technically challenging. In either case, the second measurement as described above must be performed, in which the user 40 wears the ASP-activated audio device 10. This results in the ear frequency response being blocked. In order to personalize the ASP processing, the ASP can be adjusted so that the frequency response of the blocked ear is equal to the reference (open ear) frequency response. However, this is only partially correct because there are other characteristics that may affect the sound quality. These characteristics may be based on, for example, the delay of the processed audio. Preferably, the delay of the processed audio should not be too long, because in some cases, the extended delay may be perceived as an echo of any leakage signal transmitted from the outside O to the cavity C, that is, the external sound Se leaking into the ear canal 42. Further characteristics relate to, for example, any differences in processing between the left audio device 10 and the right audio device 10, or between the left earphone and the right earphone of the stereo audio device 10. If there are significant differences in the processing, this will affect the user's ability to perceive binaural sound (binaural amplitude, delay, and coherence) and the ability to determine where the external sound Se originates. In addition, listening to just a sinusoidal signal (pure tone) is cumbersome, and multiple iterations using different frequencies are usually required to fully evaluate the open and occluded frequency responses. On the other hand, listening to an audio signal with full bandwidth and sub-band detail characteristics is very difficult. In other words, rendering a music track and requiring the user 40 to adjust the ASP processing using a multi-band equalizer is only suitable for experienced audio engineers, not for the average audio content consumer.Another obvious disadvantage is that allowing each user 40 to be evaluated in an anechoic chamber would be very cumbersome and expensive.
[0058] The inventors of the present disclosure have recognized that there is a need for a personalization process and have found the above-mentioned problems therein. There is a need to provide an efficient and flexible personalization process. Advantageously, such a process can be configured to propose and / or recommend adjustments to characteristics in situations where, for example, the user 40 is unable to decide or conclude on a path forward. The inventors have further recognized that this can be achieved by utilizing specific external sounds Se when determining the ASP calibration data. In doing so, the ASP calibration data 303 can be provided at any suitable location where such specific sounds Se can be reliably generated (see Figure 9 For this purpose, reference will be made to Figure 4 The ASP calibration system 200 is introduced. The ASP calibration system 200 includes an audio device 10, which can be any suitable audio device 10 presented in the present disclosure. The audio device 10 includes at least one transducer circuit 12 and at least one feedforward microphone circuit 14. Preferably, the audio device 10 includes at least one processing circuit 100, but in some embodiments, the audio device 10 can be operatively connected to a suitable processing circuit 100. The ASP calibration system 200 also includes at least one sound generator 220 remote from the audio device 10. The sound generator 220 can be configured to generate a specific external sound Se. The ASP calibration system 200 can optionally include an ASP calibration processing circuit 210 and / or a feedback pre-matching circuit 215. As Figure 4 As shown in the example of , the ASP calibration processing circuit 210 is configured to communicate with the sound generator 220 via the first interface 201. The sound generator 220 can be configured to transmit or provide a specific external sound Se to the audio device 10 via the second interface 202. The audio device 10 can be configured to communicate with the ASP calibration processing circuit 210 via the third interface 203. The first interface 201 and the third interface 203 can be any suitable interface, such as a wired interface or a wireless interface, for example a Bluetooth interface. The second interface 202 is preferably a direct air interface that transmits the sound (i.e., air pressure changes) generated by the sound generator 220.
[0059] Figure 4 The schematic diagram of the ASP calibration system 200 shown is an example. The ASP calibration system 200 can be configured and / or formed in a variety of different ways, all of which are fully within the scope of the present disclosure. Figure 5a, an exemplary block diagram of an ASP calibration system 200 according to various configurations is shown. In this configuration, the ASP calibration system 200 includes an ASP calibration apparatus 205, which is shown to include an ASP calibration processing circuit 210, a feedback provisioning circuit 215, and a sound generator 220. The audio device 10 includes a feedback microphone circuit 14 and a transducer circuit 12.
[0060] exist Figure 5b , another exemplary block diagram of an ASP calibration system 200 is shown. In this configuration, the ASP calibration system 200 includes a mobile phone 20, which is shown to include an ASP calibration processing circuit 210, a feedback pre-matching circuit 215, and a sound generator 220. It should be noted that the functions of the ASP calibration processing circuit 210, the feedback pre-matching circuit 215, and the sound generator 220 (which will be described in detail in subsequent sections) can be performed by circuits included in a general mobile phone 20. For example, the ASP calibration processing circuit 210 can be a processor circuit of the mobile phone 20, the feedback pre-matching circuit 215 can be a user interface including a touch interface of the mobile phone 20, and the sound generator 220 can be a speaker of the mobile phone 20. The audio device 10 includes a feedback microphone circuit 14 and a transducer circuit 12.
[0061] exist Figure 5c , another exemplary block diagram of an ASP calibration system 200 according to a configuration is shown. In this configuration, the ASP calibration system 200 includes a mobile phone 20, which is shown as including an ASP calibration processing circuit 210. In this embodiment, a feedback pre-matching circuit 215 is included in the audio device 10 together with the feedback microphone circuit 14 and the transducer circuit 12. The feedback pre-matching circuit 215 can be implemented by, for example, an input button / sensor (e.g., a volume button / sensor) at the audio device 10. In this exemplary embodiment, the sound generator 220 is a separate device that is operably connected to the mobile phone 20 via the first interface 201. The sound generator 220 can be, for example, a portable Bluetooth speaker or one or more network speakers, such as, but not limited to, a device that supports Google Audio or Sonos.
[0062] exist Figure 5d , another exemplary block diagram of an ASP calibration system 200 according to a configuration is shown. In this configuration, the ASP calibration system 200 includes a mobile phone 20, which includes a feedback provisioning circuit 215 and a sound generator 220. The audio device 10 includes a processor circuit 100, a feedback microphone circuit 14, and a transducer circuit 12. This means that the functionality of the ASP calibration processing circuit 210 is performed by the processor circuit 100 of the audio device 10.
[0063] from Figure 4 and Figures 5a to 5d It can be clearly seen that there are many different ways to compose and arrange the different devices of the ASP calibration system 200. In addition to what has been shown, it should be mentioned that, for example, the functionality of the ASP calibration processing circuit 210 (which will be described in detail in the following section) can be distributed across multiple processing circuits 100, 210 and devices 10, 20, 205.
[0064] With this in mind, reference will be made to Figure 6 A method 300 for providing personalized ambient sound playback is presented (see Figure 7 ) is an exemplary signaling diagram. The calibration method 300 is performed for a specific user 40 and a specific audio device 10. The calibration method 300 can be initiated by configuring the sound generator 220 through the ASP calibration processing circuit 210 to generate a first specific external sound Se. The first specific external sound Se will be further explained in the following section. The first specific external sound Se is provided to the processor circuit 100 of the audio device 10, for example, through the feedforward microphone circuit 14. This process may involve filtering, analog to digital conversion, etc., but these are all within the knowledge of those skilled in the art. The processor circuit 100 will process the first specific external sound Se and provide the processed internal sound Se' to the transducer circuit 12 of the audio device 10. The transducer circuit 12 will generate a first internal sound Si that can be heard by the specific user 40, that is, at this stage, the user 40 is preferably wearing the audio device 10. That is, the internal sound Si propagates in the cavity C. Based on the first internal sound Si, the ASP calibration processing circuit 210 obtains first feedback data 301 indicating the similarity between the first internal sound Si and the first specific external sound Se. The first feedback data 301 may be provided by a user via, for example, the feedback provisioning circuit 215. Alternatively or additionally, in an embodiment where the audio device includes a feedback microphone circuit 16, the feedback microphone circuit 16 may provide all or part of the first feedback data 301. If the feedback data 301 is provided by a particular user 40, it may be provided as a subjective indication obtained by comparing the sound perceived when listening to the first specific external sound Se when not wearing the audio device 10, with the sound perceived when wearing the audio device and listening to the first internal sound Si. In this exemplary embodiment, in which the ASP calibration processing circuit 210 is shown as being separate from the processing circuit 100 (although they may be separate software functions or modules executed by the same physical device), the first feedback data 301 is provided to the processing circuit 100. This enables the processing circuit 100 to adjust the initial set 123 of frequency-dependent processing parameters based on the first feedback data 301, see Figure 9 This provides a personalized set of frequency-dependent processing parameters 125, see Figure 9The processing circuit 100 may then provide the personalized set 125 of frequency-dependent processing parameters as ASP calibration data 303 for future ASP processing. The ASP calibration data 303 will be specific to a particular user 40 when the audio device 10 is worn by the particular user 40 .
[0065] refer to Figure 7 , a method 300 for providing personalized ASP calibration data 303 will be outlined in more detail. The method 300 may be referred to as a personalization process, ASP personalization, etc. Note that the different tasks described with reference to the method 300 are not necessarily performed by the same device. The tasks may be performed by any suitable one or more devices mentioned in this disclosure. Figure 7 The features of the method described in detail are exemplary features, and the method may also include any other suitable features presented herein. One step of the method 300 includes generating 310 a first specific external sound Se. The first specific external sound Se is generated remotely relative to the audio device 10. The first specific external sound Se is advantageously generated by the sound generator 220. In some embodiments, the first specific external sound Se is an external sound within a first frequency band (sometimes referred to as a frequency region), and the initial set 123 of frequency-related processing parameters is adjusted based on the first feedback data 301 and the first frequency band. Thus, the initial set 123 of frequency-related parameters is personalized based on the feedback 301 from the specific user 40, and a personalized set 125 of frequency-related processing parameters is formed.
[0066] It should have been mentioned that the method 300 can be repeated, preferably completely, for a second specific external sound Se, wherein the second specific external sound Se can be within the second frequency band. It can be seen that, in addition to the first feedback data 301 and the first frequency band as explained above, the personalized set 125 of frequency-dependent processing parameters is adjusted based on the second feedback data 301 associated with the second specific external sound Se and the second frequency band. As will be explained, there can be several specific external sounds Se suitable for each frequency region, and the method 300 can be repeated for the same frequency region but using different specific external sounds Se.
[0067] Another step of method 300 includes acquiring 320 a first specific external sound Se by audio device 10. As previously described, this is preferably accomplished via feedforward microphone 14 of audio device 10. Typically, microphone 14 converts sound into an analog electrical representation of the sensed sound, in this case, the first specific external sound Se. Because any further processing may be performed in the digital domain, acquiring 320 typically involves converting the analog electrical signal into a digital representation of the first specific external sound Se.
[0068] The method 300 also includes processing the first specific external sound Se obtained by 330. Since the processing 330 is advantageously performed in the digital domain, what is processed is a digital representation of the first specific external sound Se. The processing 330 is performed based on the initial set 123 of frequency-related processing parameters. The initial set 123 of frequency-related processing parameters can be a set 123 of factory-preset frequency-related processing parameters set by the audio device 10. In some embodiments, the initial set 123 of frequency-related processing parameters can include, for example, filter parameters, which include gain parameters for multiple frequencies. In another embodiment, the gain of the filter parameters of the initial set 123 of frequency-related processing parameters can be set to 1 (unity), that is, no gain is added.
[0069] In some embodiments, the initial set 123 of frequency-dependent processing parameters is based on one or more calibration input parameters 121 (see Figure 9 ), the calibration input parameters 121 are advantageously provided, for example, during the setup of method 300. The calibration input parameters 121 may include user-specific data, such as the user's age, etc. The calibration input parameters 121 may additionally or alternatively be based on one or more wearing states of the audio device 10. The wearing state may describe whether the audio device 10 is in an in-ear, earbud, or over-ear configuration, for example. The calibration input parameters 121 may additionally or alternatively include an indication of the relative position of the sound generator 220, such as the distance and / or direction from the audio device 10. A specific user 40 may provide one or more calibration input parameters 121, such as the wearing state, user-specific data, the relative position of the sound generator 220, etc., via, for example, the feedback provisioning circuit 215 or any other suitable input device. In some embodiments, one or more calibration input parameters 121 may be obtained from a remote data storage, such as a cloud server or the like. Those skilled in the art will appreciate that there may be more methods for obtaining these parameters, for example, depending on the type of audio device 10 and sound generator 220. For example, if the sound generator 220 is a Bluetooth-enabled sound generator 220 that communicates with the audio device 10 via a Bluetooth interface, data from that communication (e.g., beam direction, etc.) can be used to determine the relative position of the sound generator 220. The audio device 10 can also be provided with one or more sensors or switches configured to indicate, sense, or detect the state in which the audio device 10 is operating. This is particularly beneficial for reconfigurable audio devices 10 so that they can operate as, for example, in-ear headphones or earbuds, depending on the configuration.
[0070] The method also includes generating 340 a first internal sound Si based on the processed digital representation of the first specific external sound Se. This is performed by the audio device 10 when worn by the specific user 40. In simple terms, the audio device 10 plays a processed version of the first external sound Se so that the specific user 40 can perceive it. In other words, the transducer circuit 12 of the audio device is configured to play a processed version of the first external sound Se.
[0071] In order to personalize the initial set 123 of frequency-dependent processing parameters, feedback data 301 related to the generated first internal sound Si is advantageous. For this purpose, the method 300 also includes obtaining 350 first feedback data 301 related to the first internal sound Si. The first feedback data 301 preferably indicates the similarity between the first internal sound Si and the first specific external sound Se. As previously indicated, the first feedback data 301 can be provided by a specific user 40 and / or by the feedback microphone 16 of the audio device 10 (if present). Other parts of this disclosure will provide some specific examples of feedback data 301.
[0072] The method 300 further comprises adjusting 360 the initial set 123 of frequency-dependent processing parameters based on the first feedback data 301. The adjusted set 123 of frequency-dependent processing parameters may be described as a personalized set 125 of frequency-dependent processing parameters. To give a very simple example, if the specific user 40 indicates that the volume of the first internal sound Si is low compared to the first specific external sound Se, the adjustment may comprise increasing the gain of the personalized frequency-dependent processing parameters 125 compared to the gain provided by the initial set 123 of frequency-dependent processing parameters.
[0073] The method 300 may further include providing 370 the personalized set 125 of frequency-dependent processing parameters as ASP calibration data 303 for subsequent ASP processing. The ASP calibration data 303 will be specific to the specific user 40 and audio device 10.
[0074] As already indicated, part of the method 300 or the entire method 300 may be iterated multiple times so that further feedback data 301 related to additional external sounds Se may be acquired and the ASP calibration data 303 may be updated accordingly. In some embodiments, or, if applicable, in an iteration of the method 300, the method 300 may further comprise iterating the method 400, see Figure 8 The iterative method 400 is beneficial because it allows providing feedback on the internal sound Se through the personalized set 125 of applied frequency-dependent processing parameters.
[0075] To this end, the iterative method 400 includes processing 410 the digital representation of the first external sound Se based on the personalized set 125 of frequency-dependent processing parameters. This can be performed in a manner similar to, for example, the processing 330 of the first specific external sound Se obtained above. It should be noted that if more than one external sound Se was used when providing the ASP calibration data 303, this can also be applied to the additional external sounds Se. The method 400 also includes generating 420 a personalized first internal sound Si' based on the personalized processed digital representation of the first specific external sound Se. This can be performed in a manner similar to, for example, the generation 340 of the first internal sound Si described above. Furthermore, the iterative method 400 includes obtaining 430 updated first feedback data 301'. This can be performed in a manner similar to the obtaining 350 of the first feedback data described above. The updated first feedback data 301' advantageously indicates a similarity between the personalized first internal sound Si' and the first specific external sound Se. Furthermore, the iterative method 400 can include adjusting 440 the personalized set 125 of frequency-dependent processing parameters based on the updated first feedback data 301. Optionally, in some embodiments, the iterative method 400 may include providing 450 the personalized set 125 of frequency-dependent processing parameters as ASP calibration data 303 for subsequent ASP processing.
[0076] The iterative method 400 can be run one or more times. For example, in embodiments where the audio device 10 includes a feedback microphone 16, the personalization process or portions of the personalization process can be performed without the particular user 40 actively providing feedback data 301. This allows the method 300 and / or iterative method 300 to run autonomously without requiring interaction from the particular user 40. It may be advantageous to have a particular user initiate and / or set up the calibration, but otherwise, the methods 300, 400 can be performed autonomously.
[0077] In some embodiments, when providing a personalized set 125 of frequency-dependent processing parameters, iterations of the calibration method 300 of the iterative method 400 may include averaging functions and / or control functions, such as a product part, an integral part, and / or a derivative part.
[0078] Figure 9A simplified diagram illustrating how ASP calibration data 303 may be provided based on the teachings of the present disclosure. As indicated above, the initial set 123 of frequency-dependent processing parameters may be a predetermined set of parameters. Optionally, the initial set 123 of frequency-dependent processing parameters may additionally or alternatively be based on one or more calibration input parameters 121. The initial set 123 of frequency-dependent processing parameters is provided to the method 300 for providing personalized ASP calibration data 303 and, optionally, also to the iterative method 400. From the method 300 (or iterative method 400) for providing personalized ASP calibration data 303, personalized ASP calibration data 303 is provided.
[0079] Additional technical features, examples, and implementations are described below, which can be combined and used with any suitable device or method disclosed herein.
[0080] Based on the teachings presented in this article, reference will be made to Figure 10 To present advantageous embodiments of the audio device 10. The audio device 10 may be any audio device 10 presented herein, such as Figure 3 The audio device 10 is configured to form a reference Figure 4 and Figures 5a to 5d ASP calibration system 200 is presented. To this end, depending on the specific implementation, the audio device 10 includes the features required of the audio device 10 to form part of different examples of the ASP calibration system 200. Specifically, the audio device 10 includes a feedforward microphone circuit 14 for obtaining a digital representation of the external sound, a transducer circuit 12 for providing the internal sound Si, and a processor circuit 100 for performing appropriate processing. When forming part of the ASP calibration system 200, the audio device 10 can obtain ASP calibration data 303 associated with a specific user 40 (and itself).
[0081] Alternatively, the audio device 10 can be configured to process a digital representation of the external sound Se based on the ASP calibration data 303 and broadcast the processed external sound Se through the transducer circuit 12. This allows the audio device 10 to operate in a personalized ASP mode, in which the ASP is processed based on the acquired ASP calibration data 303. In this mode, the specific user 40 is less affected by any negative effects of the audio device 10 on the external sound Se than when used in a non-personalized ASP mode. In addition to providing comfort to the specific user 40, this also increases the safety of the specific user 40, for example, because the risk of not hearing or misinterpreting traffic sounds is reduced.
[0082] The audio device 10 may also include an input circuit 110. As previously explained, the input circuit 110 is configured to acquire audio data 112 via the audio interface 30. The processor circuit 100 is generally configured to broadcast the audio data 112 via the transducer circuit 12, which would constitute normal operation of the audio device 10. However, the present audio device 10 combines the audio data 112 with processed external sounds Se, so that a particular user 40 experiences both the ambient sounds and the audio data 112 (a favorite song or audiobook) simultaneously.
[0083] It should be noted that the processor circuit can also be configured to process the audio data 112 and / or external sounds Se based on the hearing profile of a specific user 40. Processing audio streams using hearing profiles is known in the art and well described. It should be noted that, in addition to or in lieu of various audiograms describing the hearing of a specific user 40, the hearing profile can also include further details and preferences related to the specific user 40. Such preferences can include, but are not limited to, specific equalization settings associated with the specific user 40 (e.g., should bass be increased). A different hearing profile can be applied to the external sounds Se than to the audio data 112.
[0084] As previously indicated, the methods 300, 400 and features described herein can be stereo or mono. Stereo processing can be performed serially or advantageously in parallel on multiple (two or more) channels and output to two or more separate transducer circuits 12. Mono processing can be processing performed on one channel and output to one or more transducer circuits 12. As a guideline, over-ear headphones are generally stereo, while in-ear headphones and earbuds are mono, i.e., one channel per ear. However, in some embodiments, one earbud / in-ear headphone in a pair of earbuds / in-ear heads is configured to also perform processing on the other earbud / in-ear headphone and send the processed data to the other earbud / in-ear headphone. All of these variations are within the scope of the present disclosure. Reference Figure 11, shows a modular view of an audio device 10 and associated processor circuit 100. The modular view of the audio device 10 is an exemplary, non-limiting view, which is provided to illustrate where a personalized ASP can be provided. As before, the audio device 10 includes a feedforward microphone circuit 14 and a transducer circuit 12, wherein the processor circuit 100 is configured to process the signal from the feedforward microphone circuit 14 before providing the signal to the transducer circuit 12. The first module 101 can be a noise reduction module 101, the second module 102 can be a filter module 102, the third module 103 can be a dynamic amplification module 103, and the fourth module 104 can be a personalized filter 104. Modules 101, 102, 103, 104 are preferably implemented in software, but in some embodiments, can be a combination of software and hardware. It should be mentioned that modules 101, 102, 103, 104 can be arranged in any suitable order, and some can be arranged in parallel. The methods 300, 400 for personalization, the initial set 123 of frequency-dependent processing parameters, the personalized set 125 of frequency-dependent processing parameters, and the ASP calibration data 303 described herein can be associated with one or more of the modules 101, 102, 103, 104. Generally, any of the modules 101, 102, 103, 104 can be personalized, but the filter module 102 is typically configured by the vendor of the audio device 10 and is considered the factory-default filter module 102. Therefore, in some embodiments, the filter module 102 is not personalized through the teachings of the present disclosure, but rather retains its default settings, i.e., the initial set 123 of frequency-dependent processing parameters for that module 102 is not personalized through the teachings of the present disclosure. The dynamic amplification module 103 and the noise reduction module 101 can be configured partially through the default factory settings and partially through the personalization described herein. That is, some of the initial set 123 of frequency-dependent processing parameters of these modules 103, 101 can be personalized, and other parameters in the initial set 123 of frequency-dependent processing parameters can be retained at default settings (factory settings, predetermined settings, etc.). The personalized filter 104 is generally personalized and configured based on the teachings presented herein, that is, the initial set 123 of frequency-dependent processing parameters associated with the personalized filter 104 can be personalized in whole or in part based on the teachings presented herein.
[0085] The personalized filter 104 can be described as comprising two parts, a first part, a personalized filter denoted H(f), which is configured, constructed, and / or updated (iteratively) according to the teachings of the present disclosure. The second part of the personalized filter 104 can be a temporary filter denoted T(f), which can be updated during various parts of the personalization process and, for example, reset at the beginning of each part.
[0086] Initially, i.e., before any personalization is performed, the noise reduction module 101, the filter module 102, and the dynamic amplification module 103 are configured with an initial set 123 of frequency-dependent processing parameters, which may include a factory default configuration provided by the audio equipment vendor. The initial set 123 of frequency-dependent processing parameters may be provided by an acoustic measurement device, such as a head and torso simulator (HATS) with a conventional ear simulator. Such a device may have a measurement bandwidth of, for example, approximately 20 Hz to 10 kHz, but larger bandwidths are also common, and bandwidths of approximately 20 Hz to 20 kHz or even higher are contemplated. There are various methods for configuring the ASP using the HATS so that a set of KPI measurements are approximately equivalent when compared after measurements are taken using open and occluded ears. This may include, for example, performing directional free-field measurements of a set of point source locations and averaging all sub-results weighted into a final open and occluded ear frequency response. Other methods may include performing diffuse-field measurements using open and occluded ears. An exemplary method is disclosed in US Pat. No. 10,951,990 B2. These methods are suitable for providing a factory default configuration, such as the initial set 123 of frequency-dependent processing parameters presented herein. However, this factory default setting is only valid for the audio device 10 and not for a particular user 40. The teachings of the present disclosure address this problem.
[0087] It should be mentioned that an initial set 123 of frequency-dependent processing parameters is advantageous, since otherwise the particular user 40 would be forced to start the personalization process from scratch. This is certainly possible, but can prove tedious and even difficult to accomplish. In this respect, the personalization process, i.e. the method 300 proposed herein, can be viewed as an individualization of the factory default settings, so that only minor adjustments need to be made - without a complete characterization of the personal hearing capabilities of the particular user 40 and / or the audio device 10. The personalized filter 104 can be configured with a bandwidth that corresponds to the auditory range of human hearing, approximately 20 Hz to 20 kHz. As indicated before, for a particular user who has to endure a pure tone hearing test, performing, for example, a pure tone hearing test over this bandwidth would be very tedious and time consuming. For this purpose, the bandwidth of the personalized filter 104 can be divided into frequency bands. This Figures 12a to 12c , where the bandwidth is divided into eight frequency bands B1 to B8. This is beneficial for the duration of the personalization process (i.e., the duration of method 300) and also for the computational complexity when determining the personalized set 125 of frequency-dependent processing parameters. Dividing (splitting) the frequency into frequency bands B1 to B8 can create a trade-off between complexity, accuracy, and duration of the personalization process.
[0088] It should be mentioned that the eight frequency bands B1 to B8 are an example and any suitable number of frequency bands may be used.Many frequency divisions are available, such as octave band divisions, 1 / 3 octave band divisions, combinations of these divisions or other divisions for different bandwidth ranges.
[0089] In this example, for completeness, a first frequency band B1 is defined between a lower frequency f0 (e.g., 20 Hz) and a first frequency f1. A second frequency band B2 is defined between the first frequency f1 and the second frequency f2. A third frequency band B3 is defined between the second frequency f2 and the third frequency f3. A fourth frequency band B4 is defined between the third frequency f3 and the fourth frequency f4. A fifth frequency band B5 is defined between the fourth frequency f4 and the fifth frequency f5. A sixth frequency band B6 is defined between the fifth frequency f5 and the sixth frequency f6. A seventh frequency band B7 is defined between the sixth frequency f6 and the seventh frequency f7. An eighth frequency band B8 is defined between the seventh frequency f7 and a higher frequency (not shown) (e.g., 20 kHz). Bands B1 to B8 may all have the same bandwidth, some bands may have the same bandwidth, while other (or all) bands may have separate bandwidths.
[0090] The inventors have recognized that advantageous zoning can be provided by the divisions outlined below, which allows a particular user 40 to remain focused and active during the personalization process while still producing acceptable accuracy within a reasonable personalization process duration. A first frequency band B1 can define a sub-bass region. For example, the first frequency f1 can be approximately 70 Hz. A second frequency band B2 can define a bass region. For example, the second frequency f2 can be approximately 250 Hz. A third frequency band B2 can define a mid-low region. For example, the third frequency f3 can be approximately 500 Hz. A fourth frequency band B4 can define a mid-mid region. For example, the fourth frequency f4 can be approximately 2 kHz. A fifth frequency band B5 can define a mid-high region. For example, the fifth frequency f5 can be approximately 4 kHz. A sixth frequency band B6 can define a presence region. For example, the sixth frequency f6 can be approximately 6 kHz.
[0091] The seventh frequency band B7 may define a detail area. For example, the seventh frequency f7 may be approximately 12 kHz.
[0092] The eighth frequency band B8 may define a bright area.
[0093] It should be mentioned that the frequency ranges presented above, sub-bass (approximately 20 Hz to 60 Hz), bass (approximately 60 Hz to 250 Hz), mid-bass (approximately 250 Hz to 500 Hz), mid-midrange (approximately 0.5 kHz to 2 kHz), mid-treble (approximately 2 kHz to 4 kHz), presence (approximately 4 kHz to 6 kHz), detail and brightness (approximately 6 kHz to 20 kHz), are well known to those skilled in the art. The frequency bands B1 to B8 do not necessarily match these frequency ranges, but these ranges are general definitions that can be used for illustrative purposes. Furthermore, there are a large number of sounds to choose from within each frequency range, and the examples given in this disclosure are not exhaustive. Further non-limiting examples of suitable specific external sounds Se, Se1 to Se8 for the frequency ranges include:
[0094] ● Bass: drum beats from bass drums, bass guitar tunings, etc.
[0095] ● Mid-bass: acoustic guitar tuning, male vocals, etc.
[0096] ● Alto: male or female vocals, electric guitar chords, birdsong, etc.
[0097] ● Mid-high pitch: male or female vocals, electric guitar chords, etc.
[0098] ● On-the-spot: hi-hat rhythm, cymbal rhythm, tenor songs, etc.
[0099] ● Details: birdsong, soprano songs, piano chords, sound effects, etc.
[0100] exist Figure 12a In , one external sound Se1 to Se8 is provided for each frequency band B1 to B8. Figure 12a , these specific external sounds Se1 to Se8 are shown as single-frequency specific external sounds Se1 to Se8, which is the case if, for example, the teachings of the present disclosure can be well performed with one or more specific external sounds Se1 to Se8 as single-frequency sounds.
[0101] The inventors have further realized that it is advantageous to assign a specific external sound Se1 to Se8 to each frequency band B1 to B8. Figure 12b. The inventors have also recognized that the choice of external sound Se (e.g., audio file) plays an important role in the quality of the ASP calibration data 303 provided by the calibration method 300. This is particularly true in embodiments where a specific user is asked to provide feedback 301, which enables the initial set 123 of frequency-dependent processing parameters to be adjusted so as to reduce the difference between the sound of the open ear and the occluded ear. It can be seen that any instructions prompting a specific user 40 to provide feedback 301 are advantageously clear and appropriate, so that the user can easily understand what is being requested and how to complete the request (what feedback 301 is expected).
[0102] Those skilled in the art are well aware that the audio industry's nomenclature contains many complex terms, many of which are unknown to non-technical personnel, such as "warm sound," "wet sound," "high frequency," etc. Furthermore, the inventors have recognized that if the specific external sound Se, Se1 to Se8 is a sound that the specific user 40 can cognitively relate to, that is, if the specific user 40 is, for example, able to recognize, is familiar with, and / or has prior knowledge of the specific external sound Se, Se1 to Se8, the quality and accuracy of the feedback 301 provided by the specific user 40 will be improved.
[0103] To this end, the inventors have developed an embodiment in which several audio files are created for each part of the personalization process, such as the external sound Se, and the frequency division is described in detail later. Figure 12b It is schematically shown in FIG, in which a specific external sound Se1 to Se8 is provided for each frequency band B1 to B8. The specific external sounds Se1 to Se8 are not single-frequency sounds, but are configured with frequency content that matches the associated frequency band B1 to B8 and are contained within the relevant frequency band B1 to B8. By way of example, for the lower range, such as the first frequency band B1 to the third frequency band B2, the specific external sound Se includes, for example, a combination of suitable drum beats, bass repeats and / or suitable low-frequency signals. For the middle range, such as the third frequency band B3 to the sixth frequency band B6, the specific external sound Se includes, for example, a combination of suitable guitars, voices and / or suitable intermediate-frequency signals. For the high range, such as the sixth frequency band B6 to the eighth frequency band B8, a specific external sound Se including, for example, bright instruments with high-frequency harmonics (such as hi-hat drums, piano notes) and / or a suitable combination of high-frequency signals is suitable.
[0104] The inventors have also realized that the audio signal (ie the specific external sound Se) is not necessarily band-limited according to, for example, the frequency bands B1 to B8 of the personalized filter 104. Figure 12c, where, for example, the fourth specific external sound Se4 has a bandwidth starting between the lower frequency f0 and the first frequency f1 and ending between the fifth frequency f5 and the sixth frequency f6. This is beneficial because if recognizable sounds are limited to the specific frequency bands B1 to B8, the user may have a negative perception of the recognizable sounds because they appear band-limited or distorted.
[0105] It is advantageous to select specific external sounds Se, Se1 to Se8 (i.e. audio signals) based on non-exclusive characteristics, because it is preferable to have multiple different specific external sounds Se, Se1 to Se8, the characteristics of which are only slightly different, but all relate to the tested frequency bands B1 to B8.
[0106] In some embodiments, it may be advantageous to reduce any processing at low frequencies (e.g., below 70 Hz to 100 Hz) to save resources. Isolating and reproducing low frequencies can be challenging, which means that the processing for low frequencies can be reduced (or not processed at all) without significantly changing the ASP quality. Similarly, at high frequencies, e.g., above 12 kHz, leakage between the audio device 10 and cavity C may increase, and reduced processing or no processing can be achieved without significantly adversely affecting the ASP quality. Furthermore, generally, above 12 kHz, there is little information that will enhance user perception, and for simplicity, these frequencies can be removed by, for example, low-pass filtering.
[0107] As a non-limiting detailed implementation example of the method 300 , a personalization process according to the present disclosure will be described below.
[0108] The specific user is positioned in front of the sound generator 220, wears the audio device 10, and runs a software application on the audio device 10, mobile phone, and / or calibration processing circuit 210. The software application may be configured to indicate to the specific user 40 that she should stand still while the recording process is being conducted.
[0109] The calibration process can begin by checking the connection with the sound generator 220. This can be performed using a process known in the art, wherein the appropriate control device 10, 20, 100, 210 configures the sound generator 220 to emit a well-known, uniquely identifiable signal, such as a chirp or a pseudo-random sequence signal, and causes the audio device 10 to actively record on the feedforward microphone circuit 14. The audio device 10 can be configured to analyze the recorded signal, for example, to find a correlation with the source signal (which can be stored in the audio device 10), or the audio device 10 can relay the recorded data (compressed or uncompressed, or in an analyzed form) to the control device 10, 20, 100, 210 for analysis and / or detection and identification. Generally, this is not a time-critical stage, and one would not need to consider power consumption at this stage. If positive detection and identification are achieved, then the sound generator 220 is determined to be sufficiently close to the audio device 10. If not, the specific user 40 can be instructed to follow a set of predefined actions to obtain positive detection and identification. Such predefined actions may include checking connections between system components, moving closer to the sound generator 200, etc. Furthermore, if the signal recorded (by the feedforward microphone circuit 14 on the audio device 10) has significant noise or interfering signal components, specific measures may be required to mitigate these sources.
[0110] Once the sound generator 220 is identified, the process can continue to check the signal quality. Since the acoustic environment (space, interior of the space, etc.) and relative position of the specific user 40 and the sound generator 220 are unknown, it is beneficial to ensure that the sound presented by the sound generator 220 is correctly received by the feedforward microphone circuit 14 of the audio device 10. For this purpose, the control device 10, 20, 100, 210 can be configured to cause a known signal to be presented on the sound generator 220 and notify the audio device 10 to use the feedforward microphone circuit 14 to record the environmental sound. The known signal preferably covers the bandwidth of operation (70Hz to 12kHz in this example), such as white noise, pink noise or pseudo-random sequence signal. Each frequency band B1 to B8 can define a set of KPIs, such as spectral flatness, minimum energy and maximum energy, peak value, valley maximum amplitude, etc. These KPIs are preferably characteristics related to the recorded sound rather than the source sound. If these KPIs are not met, a set of predefined actions similar to the actions given above about the connection with the sound generator 220 can be used to attract specific users. For example, if noise and / or interference is detected, the specific user 40 can be asked to eliminate or reduce the source of the noise and / or interference. If the signal level is low, the specific user 40 can be asked to move closer to the sound generator 220 or increase the volume of the sound generator 220. If relatively deep peaks or valleys are detected in the recorded spectrum, the specific user 40 can be asked to rearrange the settings, for example, if the sound generator 220 is located near a wall, move it and / or position it in a different location in the area. When the KPIs deemed necessary are met, for example, above a predefined set of thresholds (one for each KPI or a weighted common threshold), the quality of the sound environment is determined to be good, and the specific user 40 can be instructed to turn slightly to the side and the process restarted. The number of times the specific user 40 is instructed to turn can be configured as a trade-off between sound quality, measurement accuracy, and the duration of the personalization process. Advantageously, at least one position with frontal incident sound (sound generator 200 located in front of a particular user 40) is provided, preferably in combination with two to four additional different relative positions between the particular user 40 and the sound generator 200, such as 45 and 90 degrees rotated to the left and right, respectively. A higher number of positions is advantageous in order to provide an appropriate average of the quality of the ambient sound recording while maintaining a relatively low sensitivity to sound direction, so that no single sound direction is too prominent. While averaging across multiple directions is advantageous, a set of three to five measurements, including those for the front, left, and right directions, has proven sufficient. At each position, the control device 10, 20, 100, 210 can be configured to process the time between interacting with and instructing the particular user 40, presenting the audio signal on the transducer, notifying the audio device 10 to record, and receiving the recorded sound from the audio device 10 for analysis, or directly receiving the analyzed data.
[0111] During the signal quality check, the system 200 can be configured to store an estimate of the acoustic transmission response, as described below. The received microphone signal (frequency domain representation) is S m (f), and the source signal is S(f), then:
[0112] S m (f)=M(f)HRTF(f)R(f)T R (f)S(f)
[0113] where M(f) is the microphone frequency response, HRTF(f) is the head-related transfer function for the specific user in the current setup, R(f) is the frequency transfer of the space (the frequency representation of the spatial impulse response), and T R (f) is the frequency response of the transducer. The microphone frequency response M(f) is generally a stable characteristic with a certain degree of uncertainty depending on, for example, hardware tolerances, and is typically stored in the audio device 10. Thus, the frequency response of sound propagation from and including the transducer to the earphone can be described as:
[0114]
[0115] Where k is the number of positions in which the specific user 40 is required to stand. As will be appreciated by those skilled in the art, since a typical set of in-ear headphones occupies a larger space in the ear than a typical measurement microphone, the HRTF(f) is not exactly the specific user's 40 true HRFT, but rather a good approximation. Once all positions are approved according to the KPIs, the specific user 40 may be prompted to continue the personalization process 300. The specific user 40 may be asked to specify some information about her ears. These calibration input parameters 121 may be provided to determine an initial set 123 of frequency-dependent processing parameters. Typically, this can be done by, for example, taking a picture and having the control device 10, 20, 100, 210 identify which type of HRFT is appropriate, providing more detailed information. The specific user 40 may be asked to specify which cannula is used on the audio device 10, typically defined as small, medium, or large, with one of the three being the factory default (installed at the factory). Based on the cannula used, the system 200 can simulate the ear canal 42 of the specific user 40, as will be described in detail in the following sections.
[0116] A specific user 40 may be prompted to begin a personalization process. At the start of the process, the specific user 40 wears the audio device 10 with the ASP of the audio device 10 configured to be activated. The specific user 40 may be instructed to position herself in front of the sound generator 220, but may additionally or alternatively be instructed to rotate in a manner similar to that used when checking signal quality to average across several directions. At any suitable point in time, the specific user 40 may remove one or both of the audio devices 10 to listen to the sound generator 220 (specific external sound Se) with an open (unobstructed) ear and perceive its sound without obstruction. A portion of the process may begin with the first specific external sound Se presented by the sound generator 220. As described above, each portion may have several specific external sounds to select from. A graphical user interface (GUI) may be updated to display controls with which the specific user 40 can interact to modify the ASP processing, i.e., to update / change the personalized set 125 of frequency-dependent processing parameters. The specific user 40 may be required to interact with the GUI controls and listen to the specific internal sound Si presented by the audio device 10. The particular user 40 can adjust the controls of the GUI, thereby changing the processing of the presented audio signal, until she is satisfied. The particular user 40 can also be prompted to rate the perceived quality of the particular internal sound Si and proceed to the next section or end the process. Once the particular user 40 completes a section, the personalized set 125 of frequency-dependent processing parameters can be weighted together with any previous personalized set 125 of frequency-dependent processing parameters or the initial set 123 of frequency-dependent processing parameters, and the ASP calibration data 303 is updated accordingly. After personalization is complete, the particular user 40 can be asked to adjust the level / volume so that the overall sound level is correct. The ASP calibration data 303 is applied to the audio device 10 and is advantageously also stored locally and / or in the cloud.
[0117] The inventors have also recognized that, in embodiments where a particular user 40 provides at least some of the feedback data 301, it would be advantageous if at least some of the feedback data 301 could be provided in an intuitive manner that does not require the skill and experience of an audio engineer. To this end, the particular user 40 could provide the feedback data 301 as two-dimensional feedback data.
[0118] exist Figure 13, an exemplary two-dimensional space 500 for representing feedback 301 is shown. The two-dimensional space 500 includes a first dimension 510 and a second dimension 520. The feedback data 301 can be described as, for example, a point in the two-dimensional space 500, or a vector 535 in the two-dimensional space 500. The first dimension 510 and the second dimension 520 can represent different parameters that can be used to describe the internal sound Si, that is, the similarity between the internal sound Si and the external sound Se, or the similarity between the sounds perceived with and without the audio device 10 (occluded and unoccluded ears). The parameters represented by each dimension 510, 520 can be different depending on the stage in the personalization process (e.g., method 300). The first dimension 510 describes a first parameter between a first parameter first value 513 and a first parameter second value 517. The second dimension 520 describes a second parameter between a second parameter first value 523 and a second parameter second value 527.
[0119] Advantageously, one of the dimensions 510, 520 represents amplitude feedback data, which is configured to indicate the similarity in sound pressure level (SPL) between the perception of the internal sound Si and the external sound Se. However, the particular user 40 may not be comfortable providing feedback in terms of SPL. Therefore, the similarity in SPL can be provided in terms that are easier to manage, with which the particular user 40 may feel more comfortable. Assuming that the SPL is described by the first dimension 510, in order to simplify the task of providing feedback data 301, the first parameter first value 513 can indicate that the perceived internal sound Si is "weaker" compared to the external sound Se. Similarly, the first parameter second value 517 can indicate that the perceived internal sound Si is "louder" than the external sound Se. This means that the particular user 04 will provide feedback between two subjective extremes, such as "weak" and "loud".
[0120] Similar to SPL, the other dimensions 510, 520 can be configured to provide an indication of another parameter of the internal sound Si. To this end, the first value 513, 523 and the second value 517, 527 of the other dimensions 510, 520 can be defined as mood indicators, i.e., subjective indicators such as "bright" / "dark", "deep" / "light", etc. The choice of indicator to be used can depend on the current frequency band B1 to B8.
[0121] Feedback data 301 may also include a quality indicator provided by a particular user 40. The quality indicator may be an indicator configured to indicate the overall sound similarity compared to an open ear (preferred sound). The quality indicator may be an indicator that spans subjective terms, such as "very different" to "same." The subjective quality indicator may be a mapper of numerical values, such as [0, ..., 1], where 0.1 or less indicates poor, 0.5 indicates acceptable, and 0.9 or above indicates good similarity. The quality indicator may be represented as a slider on a user interface that the particular user 40 can manipulate.
[0122] Advantageously, the two-dimensional space 500 is presented to the specific user 40 as a user interface, such as on a touch display. This allows the specific user 40 to directly select a point describing the feedback data 301. Furthermore, in some embodiments, the specific user 40 can drag, move, or otherwise change the feedback data 301 substantially continuously during the personalization process, for example, as described in the iterative method 400. This allows the specific user 40 to manipulate the interface such that the initial set 123 of frequency-dependent processing parameters or the personalized set 125 of frequency-dependent processing parameters can be updated substantially in real time to reflect this. Furthermore, a personalized inner sound Si' can be provided based on the personalized set 125 of frequency-dependent processing parameters or the updated personalized set 125 of frequency-dependent processing parameters in substantially real time, thereby allowing the specific user to immediately perceive changes in the personalized inner sound Si'.
[0123] A further specific non-limiting example of how the method 300 may be performed, and in particular providing substantially real-time updating of the personalized interior sound Si′ based on the feedback data 30 , is given below.
[0124] A specific user 40 can change the temporary filter T(f) and thereby the ASP processing by changing the feedback data 301, advantageously via a (graphical) user interface. In some examples, the temporary ASP processing is not taken into account, i.e. provided as a personalized set 125 of frequency-dependent processing parameters or ASP calibration data 303, until the quality indicator is set. It should be mentioned that there are various ways of processing the feedback data 301, which may depend on the type of feedback data 301, the implementation, etc. Consider Figure 13Consider a two-dimensional space, assuming that the first dimension 510 (y-axis) represents the amplification of specific frequency bands B1 to B8. The shape of the temporary filter T(f) can be adjusted by operating along the second dimension 520 (x-axis). For example, if the x-values lean toward "brighter" mood indicators, this may result in an increase in high-frequency content relative to low-frequency content. Conversely, a shift toward "darker" mood indicators may result in amplification of lower frequencies rather than higher frequencies. A personalized filter 104 or temporary filter T(f) can be provided for each frequency band B1 to D8. The filter for each frequency band B1 to D8 can be configured with a gain that is linear with frequency, where the slope of the gain can be controlled by the brighter / darker mood indicator. The filtered frequency range can be determined based on the portion of the personalization process currently being performed. This is advantageous because each portion can target a specific frequency band B1 to B8 of the ambient sound. Each portion of the spectrum can have its own gain. However, to avoid saturation, it is advantageous to achieve a common amplification at the end of the personalization process. The final scaling (gain) can be determined based on the amplification (gain) of each frequency band B1 to B8 and the indication from the specific user 40 indicating the overall level adjustment. The average amplification level can be transmitted to the dynamic amplification module 103, optionally including a limiter function. This is a common process for avoiding audio distortion due to saturation inside the filter. The filter is generally normalized at 0 dB by, for example, eliminating the average level. The amplification performed by the dynamic gain controller with / without a limiter function (a component that can handle amplification while avoiding saturation) can include the average level. Therefore, most of the amplification will be provided by the dynamic amplification module 103, for example including a limiter, while the relative spectral differences are obtained through the filtering process.
[0125] Assume that a particular user 40 has rated a number of specific external sounds Se at various parts of the personalization process. The temporary ASP process can then be stored as an H for each specific external sound Se, Se1 to Se8k. k (f), where k is the number of specific external sounds Se, Se1 to Se8. The system 200 may be configured to store (e.g., on the ASP calibration processing circuit 210, the audio device 10, etc.) each iteration H k (f) The frequency response adjustment applied at time t and the quality indicator w set by the specific user 40 k The frequency response of the personalized filter 104 is determined as:
[0126]
[0127] H k (f) can be discretely defined as H(f n), where there are a total of N frequency points (N can be set to be equal to the number of frequency bands B1 to B8). The personalized filter 104 can be updated with H(f). As an optional embodiment, personalization can also include background ASP configuration. The background ASP configuration can be performed by any suitable processing circuit of the audio device 10 or the calibration system 200. During the personalization process of the specific user 40, if the specific user 40 is not satisfied with their selection, the background ASP configuration circuit can be configured to calculate a recommended adjustment of the temporary ASP filter T(f) as an alternative. If the specific user 40 completes a certain section and sets a poor quality indicator, the specific user 40 can be presented with the option of listening and comparing the temporary ASP filter T(f) determined by the background ASP configuration with the temporary ASP filter T(f) configured by the specific user 40 (frequency-dependent processing parameters 125). The specific user 40 can choose to keep its temporary ASP filter T(f) or switch to the temporary ASP filter T(f) determined by the background ASP configuration. If the specific user 40 changes the frequency-dependent processing parameters 125, the specific user 40 preferably provides a new quality indicator. Advantageously, when a specific user configures the temporary ASP filter T(f), the proposal determined by the background ASP configuration is calculated in the background. This calculation is advantageously based on the spectrum of the current specific external sound Se (the sound that the user is listening to), the user information about the sleeve used in the audio device 10, and the sound propagation frequency function H sp (f). The background ASP configuration can determine the recommended temporary filter T BASP (f) in order to:
[0128] Among them E c (f) and E0(f) are the frequency transfer functions of the occluded ear and the open ear, respectively. L(f) is the frequency transfer function of the leakage, M(f) is the frequency transfer function of the feedforward microphone circuit 14, and T t (f) is the frequency response of the transducer circuit 12. c (f) and E0(f) can generally be modeled with good accuracy as a cylindrical waveguide with either two closed ends or one open end and one closed end. The diameter and length of the waveguide can be set by associating the sleeve size with a set of numbers that specify the diameter and length. The leakage transfer function can be approximated by a fixed function of frequency obtained from measurements on, for example, a HATS. For simplicity, if it is determined that the audio device 10 provides a good fit / seal, then L(f) = 0 can be estimated. Furthermore, if nothing is known about the frequency transfer function of the leakage, then L(f) = 0 is a good approximation because leakage (if the earbud is properly inserted) is not a major contributor to the sound signature that alerts a particular user 40. The solution to the equation provided above is called a (regularized) least squares optimization problem.
[0129] For the noise reduction module 101, the total gain applied to the recorded microphone signal may sometimes be relatively high. Depending on the signal-to-noise ratio (SNR) of the microphones 14, 16, the self-noise of the microphones 14, 16 may be heard by a specific user and annoy the specific user. The specific user 40 can be prompted to move to a quiet location. In a quiet location, a slider can be presented to the specific user 40, which can adjust the amount of noise reduction so that any low self-noise is reduced to a tolerable level. Generally, this can be provided by a slider showing the degree of noise reduction. Generally, 0dB to 10dB or up to 15dB of noise reduction can be applied without causing any substantial adverse effects.
[0130] This disclosure proposes various methods, examples, implementations and features related to personalization of ASP, etc. This teaching can be used in whole or in part by Figure 14 The computer program 600 is implemented as shown. The computer program 600 includes program instructions 610, which, when executed by a suitable control device 10, 20, 100, 210 or processor circuit 100, 210, causes the device 10, 20, 100, 210 or processor circuit 100, 210 to perform any feature, method, example or embodiment presented herein. Specifically, the program instructions 610 cause the processor circuit 100 or ASP processor circuit 210 to perform the reference Figure 7 and Figure 8 At least a portion of one or both of the methods 300, 400 presented. Figure 14 As further shown, the computer program 600 may be stored on a computer readable storage medium 700. Preferably, the computer readable storage medium 700 is a non-volatile computer readable storage medium 700 such as, but not limited to, a flash-based memory device, a CD-ROM, or the like.
[0131] like Figure 15 As shown, the computer program 600 can be implemented by Figure 15 The computer-readable storage medium 700 is shown loaded onto the processor circuit 100 , 210 or transmitted via a computer network. Figure 15 6 is shown as being loaded onto the ASP calibration system 200. This means that the computer program 610 can be loaded onto any suitable device of the ASP calibration system 200.
[0132] Those skilled in the art, having the benefit of the teachings presented in the foregoing description and the associated drawings, will appreciate that modifications and other variations of the described embodiments will occur to them. It should be understood, therefore, that the embodiments are not limited to the specific example embodiments described in this disclosure, and that modifications and other variations are intended to be included within the scope of this disclosure. For example, while embodiments of the present invention have been described with reference to a portable audio device 10, those skilled in the art will understand that embodiments of the present invention may be equally applicable to other audio playback devices, such as for transmitting sound to a vehicle or ear protection devices. Furthermore, although specific terms may be used herein, they are used in a general and descriptive sense only and not for limiting purposes. Therefore, those skilled in the art will recognize that many variations of the described embodiments still fall within the scope of the appended claims. Furthermore, although individual features may be included in different claims (or embodiments), such features may be advantageously combined, and the inclusion of different claims (or embodiments) does not mean that a combination of features is not feasible and / or advantageous. Furthermore, reference to the singular does not exclude the plural. Finally, reference numerals in the claims are provided merely as clarifying examples and should not be construed as limiting the scope of the claims in any way.
Claims
1. A method (300) for providing personalized ambient sound playback (ASP) calibration data (303) associated with an audio device (10) and a specific user (40), the method (300) comprising: generating (310) a first specific external sound (Se) remotely relative to the audio device (10); acquiring (320) by the audio device (10) a digital representation of the first specific external sound (Se); processing (330) the digital representation of the first specific external sound (Se) based on the initial set (123) of frequency-dependent processing parameters; When the specific user (40) wears the audio device (10), the audio device (10) generates (340) a first internal sound (Si) based on the processed digital representation of the first specific external sound (Se); obtaining (350) first feedback data (301) indicating similarity between the first internal sound (Si) and the first specific external sound (Se); adjusting (360) an initial set (123) of frequency-dependent processing parameters based on the first feedback data (301) to obtain a personalized set (125) of frequency-dependent processing parameters; as well as When the specific user (40) wears the audio device (10), the personalized set (125) of frequency-dependent processing parameters is provided (370) as ASP calibration data (303) for the audio device (10).
2. The method (300) of claim 1, wherein: The first specific external sound (Se) is an external sound within a first frequency band (B1), and the initial set (123) of frequency-dependent processing parameters is also adjusted (360) based on the first frequency band (B1).
3. The method (300) of claim 2, wherein: The method (300) is repeated for a second specific external sound (Se2) within a second frequency band (B2) to obtain second feedback data (301), wherein the personalized set (125) of frequency-related processing parameters is further adjusted (360) based on the second feedback data (301) associated with the second specific external sound (Se2) and the second frequency band (B2).
4. The method (300) according to claim 3, wherein: The first specific sound (Se1) includes frequency content also in the second frequency band (B2), and the second specific sound (Se2) includes frequency content also in the first frequency band (B1).
5. The method (300) of claim 3, wherein: The first specific sound (Se1) includes frequency content substantially entirely within the first frequency band (B1), and the second specific sound (Se2) includes frequency content substantially entirely within the second frequency band (B2).
6. The method according to any one of claims 3 to 5, wherein The first frequency band (B1) and the second frequency band (B2) are selected from a frequency band set including at least two of a sub-bass region, a bass region, a mid-bass region, a mid-mid-range region, a mid-treble region, a presence region and a detail region.
7. The method (300) according to any one of the preceding claims, wherein: Obtaining (350) the first feedback data (301) includes obtaining feedback data (301) from the specific user (40).
8. The method (300) of claim 7, wherein: The feedback data (301) from the specific user (40) is obtained by the specific user (40) indicating feedback data in a two-dimensional space (500), wherein at least one dimension (510, 520) includes an emotion indicator.
9. The method (300) according to claim 8 and any one of claims 2 to 6, wherein: The emotion indicator is configured based on a frequency band (B1, B2) associated with an external sound (Se, Se1, Se2) related to the feedback data (301).
10. The method (300) according to any one of the preceding claims, wherein Obtaining (350) the first feedback data (301) includes obtaining feedback data (301) from a feedback microphone circuit (16) of the audio device (10).
11. The method (300) according to any one of the preceding claims, wherein: The first feedback data (301) includes amplitude feedback data indicating similarity in sound pressure level SPL between the first internal sound (Si) and the first specific external sound (Se).
12. The method according to claim 11, wherein The initial set (123) of frequency-dependent processing parameters is also adjusted (360) based on the first feedback data (301) based on one or more equal loudness curves.
13. The method (300) according to any one of the preceding claims, further comprising: processing (410) the digital representation of the first external sound (Se) based on the personalized set (125) of frequency-dependent processing parameters; When the specific user wears the audio device (10), the audio device (10) generates (420) a personalized first internal sound (Si') based on the personalized processed digital representation of the first specific external sound (Se); obtaining (430) updated first feedback data (301'), said updated first feedback data (301') indicating similarity between said personalized first internal sound (Si') and said first specific external sound (Se); as well as The personalized set (125) of frequency-dependent processing parameters is adjusted (440) based on the updated first feedback data (301).
14. The method (300) according to any one of the preceding claims, wherein: The initial set (123) of frequency-dependent processing parameters is based on one or more calibration input parameters (121), wherein the calibration input parameters (121) are one or more of the following: the wearing state of the audio device (10) and / or the relative position of a sound generator (220) configured to generate the first specific external sound (Se).
15. The method (300) according to any one of the preceding claims, wherein: The personalized frequency-dependent processing parameters (125) and the ASP calibration data (303) are configured with a limited bandwidth, preferably corresponding to the human hearing bandwidth.
16. The method (300) of claim 15, wherein: The personalized frequency-dependent processing parameters (125) and the ASP calibration data (303) are set to 1 so that no processing is performed at frequencies below 20 Hz, preferably no processing is performed at frequencies below 50 Hz, and most preferably no processing is performed at frequencies below 70 Hz.
17. The method (300) according to claim 15 or 16, wherein: The personalized frequency-dependent processing parameters (125) and the ASP calibration data (303) are set to 1 so that no processing is performed at frequencies above 20 kHz, preferably no processing is performed at frequencies above 15 kHz, and most preferably no processing is performed at frequencies above 12 kHz.
18. The method (300) according to any one of the preceding claims, wherein: The first specific external sound (Se) is a predefined sound selected from a sound set comprising a plurality of sounds, wherein at least one of the sounds is suitable for determining ASP calibration data (303) associated with at least one frequency band (B1-B8) selected from the following areas: bass area, mid-bass area, mid-mid range area, mid-treble area, presence area and / or detail area.
19. An ASP calibration system (200), comprising an audio device (10), a sound generator (220), a feedback pre-configuration circuit (215), and at least one processor circuit (100, 210), the at least one processor circuit (100, 210) being configured to pre-configure ASP calibration data (303) for a specific user (40) and the audio device (10) according to the method (300) of any one of claims 1 to 18, wherein: The audio device (10) comprises: a feedforward microphone circuit (14) configured to acquire a digital representation of a specific external sound (Se) generated by the sound generator (220); transducer circuit (12); an input circuit (110) configured to acquire audio data (112); and A processor circuit (100) is configured to process a digital representation of the specific external sound (Se) based on the ASP calibration data (303), and to play the processed specific external sound (Se) and the audio data (112) through the transducer circuit (12).
20. An audio device (10) configured to form part of an ASP calibration system (200) according to claim 19, and thereby to obtain ASP calibration data (303) of a specific user (40) and the audio device (10) according to a method (300) according to any one of claims 1 to 18, wherein: The audio device (10) comprises: a feedforward microphone circuit (14) configured to acquire a digital representation of an external sound (Se); transducer circuit (12); and Processor circuit (100).
21. The audio device (10) of claim 20, wherein The processor device (100) is configured to: process the digital representation of the external sound (Se) based on the ASP calibration data (303) and broadcast the processed external sound (Se) through the transducer circuit (12). Preferably, the processor circuit (100) is also configured to process the external sound (Se) based on the hearing profile of the specific user (40).
22. The audio device (10) according to claim 20 or 21, further comprising an input circuit (110), the input circuit (110) being configured to acquire audio data (112) via the audio interface (30), wherein: The processor circuit (100) is configured to broadcast the audio data (112) through the transducer circuit (12). Preferably, the processor circuit (100) is further configured to process the audio data (112) based on the hearing profile of the specific user (40).
23. A computer-readable storage medium (700) comprising program instructions (610) which, when executed by a processor circuit (100, 210), cause the processor circuit (100, 210) to perform the method (300) according to any one of claims 1 to 18.
Citation Information
Patent Citations
Spatial headphone transparency
US10951990B2