Media compensation pass-through and mode switching

By adjusting the levels of media and microphone input data in the audio device, and using inertial sensors and microphone data to process the audio output, the problem of not being able to hear external sounds when wearing the audio device is solved, and the effect of hearing external favorable sounds when wearing the audio device is achieved.

CN114286248BActive Publication Date: 2025-08-12DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111589336.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-06-30
Filing Date
2017-06-14
Publication Date
2025-08-12
Estimated Expiration
2037-06-14

AI Technical Summary

Technical Problem

When worn by existing audio devices, the user cannot hear external favorable sounds, such as close car sounds or friend's voices, resulting in sound blockage problems.

Method used

By receiving media stream and microphone input audio data, adjusting the levels of each frequency band to mix the output audio data, enhancing the loudness of the microphone output audio data, suppressing or pausing the media input audio data, determining the sound source direction and movement using inertial sensors and microphone data, and adjusting the audio processing process based on the mode switching indication.

Benefits of technology

It is realized that when wearing the audio device, it is possible to hear external favorable sounds while maintaining audio quality, enhancing the communication ability between the user and others.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114286248B_ABST
    Figure CN114286248B_ABST
Patent Text Reader

Abstract

The present application relates to media compensation and mode switching. Media input audio data corresponding to a media stream and microphone input audio data from at least one microphone can be received. A first level of at least one of a plurality of frequency bands of the media input audio data and a second level of at least one of a plurality of frequency bands of the microphone input audio data can be determined. Media output audio data and microphone output audio data can be generated by adjusting the level of one or more of the first and second plurality of frequency bands based on the perceived loudness of the microphone input audio data, the perceived loudness of the microphone output audio data, the perceived loudness of the media output audio data, and the perceived loudness of the media input audio data. One or more processes can be modified after receiving a mode switch indication.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information about divisional applications

[0002] This application is a divisional application. The parent application is an invention patent application filed on June 14, 2017, with application number 201780036541.1 and the title of the invention being “Media Compensation Pass and Mode Switching.” Technical Field

[0003] The present invention relates to processing video data. In particular, the present invention relates to processing media input audio data corresponding to a media stream and microphone input audio data from at least one microphone. Background Art

[0004] The use of audio devices such as headphones and earplugs has become extremely common. Such audio devices can at least partially block sounds from the outside world. Some headphones can form a substantially closed system between the headphone speaker and the eardrum, in which the sounds from the outside world are greatly reduced. There are various potential advantages to reducing the sounds from the outside world via headphones or other such audio devices, such as eliminating distortion, providing smooth equalization, etc. However, when wearing such audio devices, the user may not be able to hear sounds from the outside world that would be beneficial, such as the sound of an approaching car, the sound of a friend's voice, etc. Summary of the Invention

[0005] Some methods disclosed herein may involve receiving media input audio data corresponding to a media stream and receiving microphone input audio data from at least one microphone. As used herein, the terms "media stream," "media signal," and "media input audio data" may be used to refer to audio data corresponding to music, podcasts, movie soundtracks, and the like. However, the terms are not limited to such examples. Alternatively, the terms "media stream," "media signal," and "media input audio data" may be used to refer to audio data corresponding to other sounds received for playback, such as, for example, a portion of a telephone conversation. Some methods may involve determining a first level for at least one of a plurality of frequency bands of the media input audio data and determining a second level for at least one of a plurality of frequency bands of the microphone input audio data. Some such methods may involve generating media output audio data and microphone output audio data by adjusting the levels of one or more of the first and second plurality of frequency bands. For example, some methods may involve adjusting the levels so that a first difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone input audio data in the presence of the media input audio data. Some such methods may involve mixing media output audio data and microphone output audio data to produce mixed audio data.Some such examples may involve providing the mixed audio data to a speaker of an audio device (e.g., headphones or earbuds).

[0006] In some embodiments, the adjustment may involve applying a microphone gain and a media gain to one or more of the first and second pluralities of frequency bands. At least one of the microphone gain and the media gain may be calculated as a function of the microphone and media input levels. The function may have at least one of the following characteristics over a range of desired microphone input levels: for a fixed microphone input level, the microphone gain increases with increasing media input level; or for a fixed media input level, the microphone gain decreases with increasing microphone input level.

[0007] In some embodiments, regulation may involve only raising the level of one or more of the multiple frequency bands of the microphone input audio data. However, in some instances, regulation may involve both raising the level of one or more of the multiple frequency bands of the microphone input audio data and attenuating the level of one or more of the multiple frequency bands of the media input audio data. In some instances, the perceived loudness of the microphone output audio data in the presence of the media output audio data may be substantially equal to the perceived loudness of the microphone input audio data. According to some instances, the total loudness of the media and microphone output audio data may be within a range between the total loudness of the media and microphone input audio data and the total loudness of the media and microphone output audio data. However, in some instances, the total loudness of the media and microphone output audio data may be substantially equal to the total loudness of the media and microphone input audio data, or may be substantially equal to the total loudness of the media and microphone output audio data.

[0008] Some embodiments may involve receiving (or determining) a mode switch indication and modifying one or more processes based at least in part on the mode switch indication. For example, some embodiments may involve modifying at least one of receiving, determining, generating, or mixing processes based at least in part on the mode switch indication. In some examples, the modification may involve increasing the relative loudness of microphone output audio data relative to the loudness of media output audio data. According to some such examples, increasing the relative loudness of microphone output audio data may involve suppressing media input audio data or pausing a media stream.

[0009] According to some embodiments, the mode switch indication may be based at least in part on an indication of head movement and / or an indication of eye movement. In some such embodiments, the mode switch indication may be based at least in part on inertial sensor data. For example, the inertial sensor data may correspond to movement of the headset. In some examples, the indication of eye movement may include camera data and / or electroencephalogram data.

[0010] Some examples may involve determining the direction of a sound source based at least in part on microphone data from two or more microphones. Some such examples may involve determining whether the direction of the sound source corresponds to head movement and / or eye movement. Alternatively or additionally, some examples may involve receiving an indication of a selected sound source direction from a user. Some such examples may involve determining the direction of the sound source based at least in part on microphone data from two or more microphones. Some such examples may involve determining that the location of the sound source is a mode switch indication if the location of the sound source corresponds to the selected sound source direction.

[0011] Some other examples may involve determining a direction of a sound source based at least in part on microphone data from two or more microphones. Some such examples may involve determining whether a mode switch indication is present based at least in part on a direction of movement of the sound source. Some such examples may involve determining a mode switch indication based at least in part on determining that a direction of movement of the sound source may be toward at least one of the microphones.

[0012] Alternatively or additionally, some examples may involve determining a speed of the sound source.Some such examples may involve determining a mode switch indication based at least in part on determining that the speed of the sound source exceeds a threshold.

[0013] According to some embodiments, the mode switch indication may be based at least in part on identifying speech in microphone input audio data. Some such examples may involve classification of the microphone input audio data. For example, classification may involve determining whether the microphone input audio data contains the sound of a car horn, an approaching vehicle, screaming, shouting, the voice of a pre-selected individual, pre-selected keywords, and / or a public broadcast announcement. The mode switch indication may be based at least in part on the classification.

[0014] The methods disclosed herein can be implemented via hardware, firmware, software stored in one or more non-transitory media, and / or a combination thereof. For example, at least some aspects of the present invention can be implemented in a device comprising an interface system and a control system. The interface system can include a user interface and / or a network interface. In some embodiments, the device can include a memory system. The interface system can include at least one interface between the control system and the memory system.

[0015] The control system may include at least one processor, such as a general purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, and / or combinations thereof.

[0016] According to some examples, the control system may be able to receive media input audio data corresponding to a media stream and receive microphone input audio data from at least one microphone. In some embodiments, the device may include a microphone system with one or more microphones. In some instances, the microphone system may include two or more microphones. In some embodiments, the device may include a speaker system including one or more speakers. According to some such embodiments, the device may be a component of a headset or headphones. However, in other embodiments, the device may be configured to receive microphone input audio data and / or receive media input audio data corresponding to a media stream from another device.

[0017] In some instances, the control system may be able to determine a first level of at least one of a plurality of frequency bands of media input audio data and determine a second level of at least one of a plurality of frequency bands of microphone input audio data. For example, the control system may be able to generate media output audio data and microphone output audio data by adjusting the levels of one or more of the first and second plurality of frequency bands. For example, the control system may be able to adjust the levels so that a first difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone input audio data in the presence of the media input audio data. In some instances, the control system may be able to mix the media output audio data and the microphone output audio data to generate mixed audio data. According to some instances, the control system may be able to provide the mixed audio data to a speaker of an audio device (e.g., headphones or earbuds).

[0018] In some embodiments, regulation may involve only raising the level of one or more of the multiple frequency bands of the microphone input audio data. However, in some instances, regulation may involve both raising the level of one or more of the multiple frequency bands of the microphone input audio data and attenuating the level of one or more of the multiple frequency bands of the media input audio data. In some instances, the perceived loudness of the microphone output audio data in the presence of the media output audio data may be substantially equal to the perceived loudness of the microphone input audio data. According to some instances, the total loudness of the media and microphone output audio data may be within a range between the total loudness of the media and microphone input audio data and the total loudness of the media and microphone output audio data. However, in some instances, the total loudness of the media and microphone output audio data may be substantially equal to the total loudness of the media and microphone input audio data, or may be substantially equal to the total loudness of the media and microphone output audio data.

[0019] According to some examples, the control system may be capable of receiving (or determining) a mode switch indication and modifying one or more processes based at least in part on the mode switch indication. For example, the control system may be capable of modifying at least one of receiving, determining, generating, or mixing processes based at least in part on the mode switch indication. In some examples, the modification may involve increasing the relative loudness of the microphone output audio data relative to the loudness of the media output audio data. According to some such examples, increasing the relative loudness of the microphone output audio data may involve suppressing the media input audio data or pausing the media stream.

[0020] According to some embodiments, the control system may be capable of determining a mode switch indication based at least in part on an indication of head movement and / or an indication of eye movement. In some such embodiments, the device may include an inertial sensor system. According to some such embodiments, the control system may be capable of determining a mode switch indication based at least in part on inertial sensor data received from the inertial sensor system. For example, the inertial sensor data may correspond to movement of the headset.

[0021] In some examples, the device may include an eye movement detection system. According to some such embodiments, the control system may be capable of determining a mode switch indication based at least in part on data received from the eye movement detection system. In some instances, the eye movement detection system may include one or more cameras. In some instances, the eye movement detection system may include an electroencephalogram (EEG) system, which may include one or more EEG electrodes. According to some embodiments, the EEG electrodes may be configured to be placed in the user's ear canal and / or on the user's scalp. According to some such examples, the control system may be capable of detecting the user's eye movement by analyzing EEG signals received from one or more EEG electrodes of the EEG system. In some such instances, the control system may be capable of determining a mode switch indication based at least in part on an indication of eye movement. The indication of eye movement may be based on camera data and / or EEG data from the eye movement detection system.

[0022] According to some examples, the control system may be able to determine the direction of the sound source based at least in part on microphone data from two or more microphones. According to some such examples, the control system may be able to determine whether the direction of the sound source corresponds to head movement and / or eye movement. Alternatively or in addition, the control system may be able to receive an indication of a selected sound source direction from the user. In some such examples, the control system may be able to determine the direction of the sound source based at least in part on microphone data from two or more microphones. For example, if the position of the sound source corresponds to the selected sound source direction, then the control system may be able to determine that the position of the sound source is a mode switch indication.

[0023] Some other examples may involve determining the direction of a sound source based at least in part on microphone data from two or more microphones. In some such examples, the control system may be capable of determining whether a mode switch indication is present based at least in part on the direction of movement of the sound source. In some such examples, the control system may be capable of determining a mode switch indication based at least in part on determining that the direction of movement of the sound source is toward at least one microphone.

[0024] Alternatively or additionally, the control system may be capable of determining a speed of the sound source.In some such instances, the control system may be capable of determining a mode switch indication based at least in part on determining that the speed of the sound source exceeds a threshold.

[0025] According to some embodiments, the mode switch indication can be based at least in part on identifying speech in microphone input audio data. In some such instances, the control system can be capable of classifying the microphone input audio data. For example, the classification can involve determining whether the microphone input audio data contains the sound of a car horn, an approaching vehicle, screaming, shouting, the voice of a pre-selected individual, pre-selected keywords, and / or a public broadcast announcement. The mode switch indication can be based at least in part on the classification.

[0026] Some embodiments may include one or more non-transitory media on which software is stored. In some instances, the non-transitory media may include flash memory, a hard drive, and / or other memory devices. The software may include instructions for controlling at least one device to receive media input audio data corresponding to a media stream and to receive microphone input audio data from at least one microphone. The software may include instructions for determining a first level of at least one of a plurality of frequency bands of the media input audio data and a second level of at least one of a plurality of frequency bands of the microphone input audio data. The software may include instructions for generating media output audio data and microphone output audio data by adjusting the levels of one or more of the first and second plurality of frequency bands. For example, the software may include instructions for adjusting the levels so that a first difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone input audio data in the presence of the media input audio data. In some instances, the software may include instructions for mixing the media output audio data and the microphone output audio data to generate mixed audio data. Some such examples may involve providing mixed audio data to speakers of an audio device (eg, headphones or earbuds).

[0027] In some embodiments, regulation may involve only raising the level of one or more of the multiple frequency bands of the microphone input audio data. However, in some instances, regulation may involve both raising the level of one or more of the multiple frequency bands of the microphone input audio data and attenuating the level of one or more of the multiple frequency bands of the media input audio data. In some instances, the perceived loudness of the microphone output audio data in the presence of the media output audio data may be substantially equal to the perceived loudness of the microphone input audio data. According to some instances, the total loudness of the media and microphone output audio data may be within a range between the total loudness of the media and microphone input audio data and the total loudness of the media and microphone output audio data. However, in some instances, the total loudness of the media and microphone output audio data may be substantially equal to the total loudness of the media and microphone input audio data, or may be substantially equal to the total loudness of the media and microphone output audio data.

[0028] In some instances, the software may include instructions for receiving (or determining) a mode switch indication and for modifying one or more processes based, at least in part, on the mode switch indication. For example, in some embodiments, the software may include instructions for modifying at least one of a receiving, determining, generating, or mixing process based, at least in part, on the mode switch indication. In some examples, the modification may involve increasing the relative loudness of the microphone output audio data relative to the loudness of the media output audio data. According to some such instances, the software may include instructions for increasing the relative loudness of the microphone output audio data by suppressing the media input audio data or pausing the media stream.

[0029] According to some embodiments, the software may include instructions for determining a mode switch indication based at least in part on an indication of head movement and / or an indication of eye movement. In some such embodiments, the software may include instructions for determining a mode switch indication based at least in part on inertial sensor data. For example, the inertial sensor data may correspond to movement of the headset. In some examples, the indication of eye movement may include camera data and / or electroencephalogram data.

[0030] In some instances, the software may include instructions for determining the direction of the sound source based at least in part on microphone data from two or more microphones. In some such instances, the software may include instructions for determining whether the direction of the sound source corresponds to head movement and / or eye movement. Alternatively or in addition, in some instances, the software may include instructions for receiving an indication of a selected sound source direction from the user. In some such instances, the software may include instructions for determining the direction of the sound source based at least in part on microphone data from two or more microphones. In some such instances, if the position of the sound source corresponds to the selected sound source direction, the software may include instructions for determining that the position of the sound source is a mode switch indication.

[0031] Some other examples may involve determining the direction of the sound source based at least in part on microphone data from two or more microphones. According to some embodiments, the software may include instructions for determining whether a mode switch indication is present based at least in part on the direction of movement of the sound source. Some such examples may involve determining a mode switch indication based at least in part on determining that the direction of movement of the sound source may be toward at least one of the microphones.

[0032] Alternatively or additionally, in some instances the software may include instructions for determining the speed of the sound source. In some such instances, the software may include instructions for determining a mode switch indication based at least in part on determining that the speed of the sound source exceeds a threshold.

[0033] According to some embodiments, the mode switch indication can be based at least in part on identifying speech in the microphone input audio data. In some such instances, the software can include instructions for classifying the microphone input audio data. For example, the classification can involve determining whether the microphone input audio data contains the sound of a car horn, an approaching vehicle, screaming, shouting, the voice of a pre-selected individual, pre-selected keywords, and / or a public broadcast announcement. In some such instances, the software can include instructions for determining the mode switch indication based at least in part on the classification.

[0034] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1A is a block diagram illustrating an example of components of a device capable of implementing various aspects of the present invention.

[0036] Figure 1BAn example is shown in which the speaker system and the control system are in different devices.

[0037] Figure 2 is an overview which can be obtained by e.g. Figure 1A or Figure 1B Flowchart of an example of a method performed by a device shown in FIG.

[0038] Figure 3 An example of an audio device incorporating an inertial sensor system is shown.

[0039] Figure 4 An example of a microphone system comprising a pair of coincident vertically stacked directional microphones is shown.

[0040] Figure 5 Another example of a microphone system comprising a pair of coincident vertically stacked directional microphones is shown.

[0041] Figure 6 Examples of azimuth and elevation angles relative to a microphone system comprising coincident pairs of vertically stacked directional microphones are shown.

[0042] Figure 7 is a graph showing an example of a curve indicating the relationship between the azimuth angle and the ratio of intensity or level (L / R energy ratio) between right and left microphone audio signals produced by a pair of coincident vertically stacked directional microphones.

[0043] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION

[0044] The following description relates to certain embodiments for the purpose of describing some innovative aspects of the present invention and examples of situations in which these innovative aspects can be implemented. However, the teachings herein can be applied in a variety of different ways. For example, although various embodiments are described with respect to specific audio devices, the teachings herein are broadly applicable to other known audio devices, as well as audio devices that may be introduced in the future. In addition, the described embodiments can be implemented at least in part in various devices and systems as hardware, software, firmware, cloud-based systems, etc. Accordingly, the teachings of the present invention are not intended to be limited to the embodiments shown in the figures and / or described herein, but instead have broad applicability.

[0045] As mentioned above, audio devices that provide at least some degree of sound blocking offer various potential advantages, such as the ability to control improvements in audio quality. Other advantages include attenuating potentially annoying or distracting sounds from the outside world. However, users of such audio devices may not be able to hear sounds from the outside world that would be beneficial, such as the sound of approaching cars, car horns, public announcements, etc.

[0046] Accordingly, one or more types of sound blockage management will be desirable. Various embodiments described herein relate to sound blockage management during the time when a user is listening to a media stream of audio data via headphones, earbuds, or another such audio device. As used herein, the terms "media stream," "media signal," and "media input audio data" may be used to refer to audio data corresponding to music, podcasts, movie soundtracks, and the like, as well as audio data corresponding to sounds received for playback, such as part of a telephone conversation. In some embodiments, such as earbud embodiments, a user may be able to hear a large amount of sound from the outside world even when listening to audio data corresponding to a media stream. However, some audio devices (e.g., headphones) may significantly reduce sounds from the outside world. Accordingly, some embodiments may also relate to providing microphone data to the user. The microphone data may provide sounds from the outside world.

[0047] When the microphone signal corresponding to the sound outside the audio device (e.g., headphones) is mixed with the media signal and played via the speakers of the headphones, the media signal usually masks the microphone signal, making the external sound inaudible or incomprehensible to the listener. Thus, it is desirable to process both the microphone and the media signal so that when mixed, the microphone signal is auditorily higher than the media signal, and both the processed microphone and the media signal remain naturally audible. To achieve this effect, it is useful to consider models of perceived loudness and partial loudness, as disclosed herein. Some such embodiments provide one or more types of pass-through modes. In pass-through mode, the media signal can be reduced in volume, and the conversation between the user and other people (or other external sounds of interest to the user, as indicated by the microphone signal) can be mixed into the audio signal provided to the user. In some instances, the media signal can be temporarily muted.

[0048] Figure 1Ais a block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present invention. In this example, apparatus 100 includes an interface system 105 and a control system 110. Interface system 105 may include one or more network interfaces, one or more user interfaces, and / or one or more external device interfaces (e.g., one or more Universal Serial Bus (USB) interfaces). In some examples, interface system 105 may include a communication interface between control system 110 and a memory system (e.g., Figure 1A 1 ). However, the control system 110 may include a memory system. For example, the control system 110 may include a general-purpose single- or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components. In some embodiments, the control system 110 may be capable of at least partially performing the methods disclosed herein.

[0049] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Non-transitory media may reside, for example, on a Figure 1A 1 and / or reside in the control system 110. Accordingly, various innovative aspects of the subject matter described herein may be implemented in non-transitory media having software stored thereon. The software may, for example, include instructions for controlling at least one device to process audio data. The software may, for example, be executable by a control system (e.g., Figure 1A The control system 110) is executed by one or more components.

[0050] In some examples, device 100 may include: an optional microphone system 120 comprising one or more microphones; an optional speaker system 125 comprising one or more speakers; and / or an optional inertial sensor system 130 comprising one or more inertial sensors, such as Figure 1A Some examples of microphone configurations are disclosed herein. For example, an inertial sensor may include one or more accelerometers or gyroscopes.

[0051] However, in some embodiments, interface system 105 and control system 110 may be in one device, and microphone system 120, speaker system 125, and / or inertial sensor system 130 may be in one or more other devices. Figure 1BAn example is shown in which the speaker system and the control system are in different devices. In this example, the speaker system 125 comprises earbuds 150 and the control system is in a smartphone 100a attached to the user's arm. Accordingly, the smartphone is Figure 1A An example of the device 100 is shown in FIG. In alternative examples, some of which are described below, the speaker system 125 may include headphones.

[0052] Figure 2 is an overview which can be obtained by e.g. Figure 1A 1B or 1B is a flowchart of an example of a method performed by the device shown in FIG. Like other methods described herein, the blocks of method 200 are not necessarily executed in the order indicated. In addition, such methods may include more or fewer blocks than shown and / or described.

[0053] In this example, block 205 of method 200 involves receiving media input audio data corresponding to a media stream. The audio data may correspond to music, a television program soundtrack, a movie soundtrack, a podcast, or the like, for example.

[0054] Here, block 210 involves receiving microphone input audio data from at least one microphone. According to some embodiments, the microphone input audio data may be received from one or more local microphones, so that the microphone input audio data corresponds to sounds from the outside world. In some such instances, the control system block 205 of method 200 involves receiving media input audio data and microphone input audio data via an interface system.

[0055] exist Figure 2 In the example of , block 215 involves determining a first level for each of a plurality of frequency bands of media input audio data. Here, block 220 involves determining a second level for each of a plurality of frequency bands of microphone input audio data. The terms "first level" and "second level" are used herein to distinguish between the levels of a frequency band of media input audio data and the levels of a frequency band of microphone input audio data. Depending on the particular circumstances, the first level may or may not be substantially different from the second level. In some examples, blocks 215 and 220 may involve performing a transform from the time domain to the frequency domain. However, in alternative examples, the received media input audio data and / or the received microphone input audio data may have already been transformed from the time domain to the frequency domain.

[0056] In this embodiment, block 225 involves generating media output audio data and microphone output audio data by adjusting the levels of one or more of the first and second pluralities of frequency bands. According to this example, the levels are adjusted based at least in part on perceived loudness. Specifically, some examples involve adjusting the levels of one or more of the first and second pluralities of frequency bands such that a first difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone input audio data in the presence of the media input audio data. Some detailed examples are described below.

[0057] Here, block 230 involves mixing the media output audio data with the microphone output audio data to produce mixed audio data.The mixed audio data may, for example, be provided to speakers of an audio device (eg, headphones or earbuds).

[0058] In some instances, the adjustment process may involve only raising the levels of multiple frequency bands of microphone input audio data. Some such instances may involve only temporarily raising the levels of multiple frequency bands of microphone input audio data. However, in some embodiments, the adjustment may involve both raising the levels of multiple frequency bands of microphone input audio data and attenuating the levels of multiple frequency bands of media input audio data.

[0059] In some examples, the perceived loudness of the microphone output audio data in the presence of the media output audio data can be substantially equal to the perceived loudness of the microphone input audio data. According to some embodiments, the total loudness of the media and microphone output audio data can be within a range between the total loudness of the media and microphone input audio data and the total loudness of the media and microphone audio data produced by simply boosting the microphone signal. Alternatively, the total loudness of the media and microphone output audio data can be equal to the total loudness of the media and microphone input audio data, or can be equal to the total loudness of the media and microphone audio data produced by simply boosting the microphone signal.

[0060] According to some embodiments, the loudness model is defined by a specific loudness function L{·} operated on an excitation signal E. The excitation signal, which varies across both frequency and time, is intended to represent the time-varying distribution of energy induced by the audio signal of interest along the basilar membrane of the ear. In practice, the excitation is calculated via a filter bank analysis that breaks the signal into discrete frequency bands b with each band signal varying across time t. Ideally, but not necessarily, the spacing of these bands across frequencies could be commensurate with a perceptual frequency scale such as ERB (equivalent rectangular bandwidth). Representing this filter bank analysis by the function FB{·}, a multi-band version x of the input media and microphone signals is given bymed (t) and x mic (t) can be generated, for example, as shown in Equations 1a and 1b:

[0061] X med (b,t)=FB{x med (t)} (1a)

[0062] X mic (b,t)=FB{x mic (t)} (1b)

[0063] In Equation 1a, X med (b,t) represents the multi-band version of the input media signal. In Equation 1b, X mic (b, t) represents a multi-band version of the input microphone signal. In some instances, Figure 2 Block 205 may involve receiving a time domain version of an input media signal, e.g., x med (t), and block 210 may involve receiving a time domain version of the input microphone signal, e.g., x mic (t). However, in an alternative embodiment, Figure 2 Block 205 may involve receiving a multi-band version of an input media signal, e.g., X med (b, t), and block 210 may involve receiving a multi-band version of the input microphone signal, e.g., X mic (b,t).

[0064] According to some embodiments, the excitation functions of the media and microphone signals are next calculated. In some such instances, the excitations of the media and microphone signals can be calculated as the time-smoothed power of the multi-band signal with the applied frequency-varying perceptual weights W(b), for example, as shown in Equations 2a and 2b:

[0065] E med (b,t)=λE med (b,t-1)+(1-λ)W(b)|X med (b,t)| 2 (2a)

[0066] E mic (b,t)=λE mic (b,t-1)+(1-λ)W(b)|X mic (b,t)| 2 (2b)

[0067] In some embodiments, W(b) can take into account the transfer functions of the headphone, outer ear, and middle ear. In Equation 2a, E med(b, t) represents the excitation of the media signal, and in Equation 2b, E mic (b, t) represents the excitation of the microphone signal. Equations 2a and 2b involve simple single-pole smoothing functions parameterized by the smoothing coefficient λ, but other smoothing filters are possible. Equation 2a provides Figure 2 Equation 2b provides an example of the process of block 220 .

[0068] With the excitation signal generated, the specific loudness function L{·} can be applied to generate the specific loudness of the media and microphone, for example, as shown in Equations 3a and 3b:

[0069] L med (b,t)=L{E med (b,t)} (3a)

[0070] L mic (b,t)=L{E mic (b,t)} (3b)

[0071] In Equation 3a, L med (b, t) represents the specific loudness function corresponding to the media signal, and in Equation 3b, L mic (b, t) denotes the specific loudness function corresponding to the microphone signal. The specific loudness function models various nonlinearities in the human perception of loudness, and the resulting specific loudness signal describes the time-varying distribution of perceived loudness across frequency. Accordingly, the specific loudness function L mic (b,t) provides the above reference Figure 2 An example of the “perceived loudness of microphone input audio data” described in block 225 of FIG.

[0072] These specific loudness signals for media and microphone audio data represent the perceived loudness of the media stream and the sound from the microphone when each is heard in isolation. However, when the two signals are mixed, masking can occur. Specifically, if one signal is significantly louder than the other, it will mask the softer signal, thereby reducing the perceived loudness of the softer signal relative to the perceived loudness of the softer signal when heard in isolation.

[0073] This masking phenomenon can be modeled by a partial loudness function PL{·,·}, which requires two inputs. The first input is the excitation of the signal of interest, and the second input is the excitation of the competing signal. The partial loudness function returns a partial specific loudness signal PL that represents the perceived loudness of the signal of interest in the presence of the competing signal. If the excitation of the competing signal is zero, then the partial specific loudness of the signal of interest is equal to its specific loudness, PL=L. As the excitation of the interfering signal grows, PL decreases to below L due to masking. However, in order for this decrease to be significant, the level of the competing signal excitation must be close to or greater than the excitation of the signal of interest. If the excitation of the signal of interest is significantly greater than the competing signal excitation, then the partial specific loudness of the signal of interest is approximately equal to its specific loudness,

[0074] For the purpose of maintaining the audibility of the microphone signal in the presence of the media signal, we can consider the microphone as the signal of interest and the media as the competing signal. With this designation, the partial specific loudness of the microphone can be calculated from the excitations of the microphone and the media, for example, as shown in Equation 4:

[0075] PL mic (b,t)=PL{E mic (b,t),E med (b,t)} (4)

[0076] Generally speaking, the partial specific loudness PL of a microphone in the presence of media is mic (b,t) is less than its isolated specific loudness L mic (b, t). To maintain the audibility of the microphone signal when mixed with the media, the microphone and media signals may be processed so that the partial specific loudness of the processed microphone signal in the presence of the processed media signal is closer to L mic (b, t), which represents the audibility of the isolated microphone signal. Specifically, the frequency-varying and time-varying microphone and media processing gain G mic (b,t) and G med (b, t) can be calculated so that the microphone specific loudness L mic (b,t) and the specific loudness of the processed microphone part The difference between them is less than the microphone specific loudness L mic (b, t) and the specific loudness PL of the unprocessed microphone part mic The difference between (b,t):

[0077]

[0078] So that:

[0079]

[0080] The expression in Equation 5b Provide the above reference Figure 2 An example of a “first difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone output audio data in the presence of the media output audio data” as described in block 225 of FIG. 5B . Similarly, the expression L in Equation 5b is mic [b,t]-PL mic [b,t] Provide reference to the above Figure 2 This is an example of “a second difference between the perceived loudness of the microphone input audio data and the perceived loudness of the microphone input audio data in the presence of the media input audio data” as described in block 225 of FIG.

[0081] Once these gains are calculated, processed media and microphone signals can be generated by applying a synthesis filterbank or inverse transform to the corresponding gain-modified filterbank signals, for example as shown below:

[0082] y med (t) = FB -1 {G med (b,t)X med (b,t)} (6a)

[0083] y mic (t) = FB -1 {G mic (b,t)X mic (b,t)} (6b)

[0084] The expression y in Equation 6a med (t) Provide reference to the above Figure 2 6b is an example of the "media output audio data" described in block 225 of FIG. mic (t) Provide reference to the above Figure 2 An example of “microphone output audio data” is described in block 225 of FIG.

[0085] In some instances, the final output signal can be generated by mixing the processed media and the microphone signal:

[0086] y(t)=y med (t)+y mic (t) (7)

[0087] Accordingly, Equation 7 provides the above reference Figure 2 Block 230 is an example of "mixing media output audio data and microphone output audio data to generate mixed audio data."

[0088] To calculate the required microphone and media processing gains, it may be useful to define an inverse partial specific loudness function that returns the excitation of the signal of interest that corresponds to a particular partial specific loudness of the signal of interest in the presence of competing signal excitations, e.g.:

[0089] PL -1 {PL int ,E comp}=E int (8a)

[0090] Make

[0091] PL int =PL{E int ,E comp} (8b)

[0092] In Equations 8a and 8b, PL -1 represents the inverse partial specific loudness function, PL int represents the specific loudness of the signal of interest, E int represents the excitation of the signal of interest and E comp Incentives that represent competing signals.

[0093] One example of a solution that meets the general goal of the embodiment described by Equation 5 is to make the processed microphone portion specific loudness equal to the isolated microphone's specific loudness, for example, as shown below:

[0094]

[0095] Setting this condition stipulates that the loudness of the processed microphone in the presence of processed media is the same as the loudness of the original, unprocessed microphone on its own. In other words, the perceived loudness of the microphone should remain consistent regardless of the playback of the media signal. Substituting Equation 9 and Equation 3b into Equation 5a and using the definition of inverse partial specific loudness given in Equations 8a and 8b yields the processing gain G for the microphone: mic The corresponding solution of (b,t):

[0096]

[0097] Imposing the constraint that the media signal remains unprocessed means that G med (b, t) = 1, yielding a unique solution for the microphone processing gain that is computed from the known microphone and media excitation signals, as seen in (10). This particular solution may involve boosting only the microphone signal to maintain its perceived loudness, while leaving the media signal alone. Thus, this solution for the microphone gain is referred to as G boost (b,t).

[0098] Although the solution G boost (b, t) does maintain the audibility of the microphone signal above the media, but in practice the combined processed microphone and media sound may become too loud or unnatural sounding. To avoid this, it may be necessary to impose different constraints on Equation 10 to obtain a unique solution for the microphone and media gains. One such alternative is to constrain the total loudness of the mixture to be equal to some target. The total loudness of the unprocessed microphone and media mixture, L tot (b,t) can be given by the loudness function applied to the sum of the microphone and media excitation:

[0099] L tot (b,t)=L{E mic (b,t)+E med (b,t)} (11a)

[0100] The total loudness of the processed mixture of the media output audio data and the microphone output audio data can be defined in a similar way:

[0101]

[0102] The total loudness of the enhancement-only solution can be expressed as follows:

[0103]

[0104] To reduce the overall loudness of the processed mixture, the total loudness of the processed mixture can be specified to be somewhere between the total loudness of the enhancement-only solution and the total loudness of the unprocessed mixture, for example, as follows:

[0105]

[0106] Combining Equation 12 with Equation 10 specifies a unique solution for both microphone and media gains. When α = 1, the resulting solution is equivalent to the enhancement-only solution, and when α = 0, the overall loudness of the mixture remains the same as the unprocessed mixture by additionally attenuating the media signal. When α is between one and zero, the overall loudness of the mixture lies somewhere between these two extremes. Regardless, the application of Equation 10 ensures that the partial loudness of the processed microphone signal remains equal to the loudness of the microphone signal alone, thereby maintaining its audibility in the presence of the media signal.

[0107] Conventional audio devices such as headphones and earbuds typically have one operating mode (media playback mode) in which media input audio data from a laptop, computer, mobile phone, mobile audio player, or tablet computer is reproduced to the user's eardrum. In some examples, such media playback mode may use active noise cancellation technology to eliminate or at least reduce interference from ambient sound or background noise.

[0108] Some audio methods disclosed herein may relate to additional modes, for example, by mode. Some such by modes are described above. In some by mode instances, media audio signals may be reduced or muted in volume, and conversations between the user and other people (or other external sounds of interest to the user) may be captured by the microphone of an audio device (for example, a headset or earplugs), and mixed into the output audio for playback. In some such embodiments, the user may be able to participate in a conversation without having to stop media playback and / or remove the audio device from the user's ear. Accordingly, some such modes may be referred to as "conversation modes" herein. In some instances, the user may give a command, for example, via the user interface of the interface system 105 described above, so that operating mode is changed to conversation mode. Such a command is an example of "mode switching indication" as described herein.

[0109] However, other types of operating mode switching for audio devices are disclosed herein. According to some such embodiments, the mode switching may not require user input, but may instead be automatic. One or more types of audio processing may be modified after receiving a mode switch indication. According to some such examples, the above reference Figure 2 One or more of the receiving, determining, generating, or mixing processes described above can be modified based on the mode switch indication. In some examples, the modification can involve increasing the relative loudness of the microphone output audio data relative to the loudness of the media output audio data. For example, the modification can involve increasing the relative loudness of the microphone output audio data by suppressing the media input audio data or pausing the media stream.

[0110] The inventors contemplate various types of mode switch indications. In some instances, the mode switch indication may be based at least in part on an indication of head movement. Alternatively or additionally, the mode switch indication may be based at least in part on an indication of eye movement. In some instances, head movement may be detected by an inertial sensor system. Accordingly, in some embodiments, the mode switch indication may be based at least in part on inertial sensor data from the inertial sensor system. The inertial sensor data may indicate movement of the headset, for example, movement of the headset worn by the user.

[0111] Figure 3An example of an audio device that includes an inertial sensor system is shown. In this example, the audio device is a headset 305. Inertial sensor system 310 includes one or more inertial sensor devices, such as one or more gyroscopes, one or more accelerometers, etc. Inertial sensor system 310 is capable of providing inertial sensor data to a control system. In this example, at least a portion of the control system is a component of device 100b, which is an example of device 100 described elsewhere herein. Alternatively or in addition, at least a portion of the control system can be a component of an audio device, such as headset 305. The inertial sensor data can indicate movement of headset 305, and therefore can indicate movement of the user's head when the user wears headset 305.

[0112] exist Figure 3 In the example shown in , device 100b includes a camera system having at least one camera 350. In some instances, the camera system may include two or more cameras. In some embodiments, a control system (e.g., of device 100b) may be able to determine the user's eye movements and / or the direction the user is currently looking based at least in part on camera data from the camera system. Alternatively or in addition, the control system may be able to determine the user's eye movements based on electroencephalogram (EEG) data. Such EEG data may, for example, be received from an EEG system of a headset 305. In some embodiments, the headset 305 (or another audio device, e.g., earbuds) may include one or more EEG electrodes configured to be placed in the user's ear canal and / or on the user's scalp. The user's eye movements may be determined via analysis of EEG signals from the one or more EEG electrodes.

[0113] In this example, headset 305 includes headset units 325a and 325b, each of which includes one or more speakers of speaker system 125. In some examples, each of headset units 325a and 325b can include one or more EEG electrodes. According to some such examples, each of headset units 325a and 325b can include at least one EEG electrode on the front side so that the EEG electrodes can be placed near the eyes of user 370 when headset 305 is worn. Figure 3In an example of the present invention, when wearing headset 305, EEG electrode 375a of headset unit 325a can be placed near the right eye of user 370 and EEG electrode 375b of headset unit 325b can be placed near left eye 380. In some such embodiments, the potential difference between EEG electrode 375a and EEG electrode 375b can be used to detect eye movement. In this example, headset units 325a and 325b also include microphones 320a and 320b. In some examples, the control system of device 100b or headset 305 can be capable of determining the direction of the sound source based at least in part on microphone data from two or more microphones (e.g., microphones 320a and 320b). According to some such examples, the control system can be capable of determining the direction corresponding to the location of the sound source based at least in part on the intensity difference between the first microphone audio signal from microphone 320a and the second microphone audio signal from microphone 320b. In some examples, an "intensity difference" may be or may correspond to a ratio of intensities or levels between a first microphone audio signal and a second microphone audio signal.

[0114] Alternatively or additionally, the control system may be capable of determining a direction corresponding to the location of the sound source based at least in part on a time difference between a first microphone audio signal from microphone 320a and a second microphone audio signal from microphone 320b. Some examples of determining an azimuth angle corresponding to the location of the sound source and determining an elevation angle corresponding to the location of the sound source are provided below.

[0115] In some instances, the control system may be able to determine whether the direction of a sound source corresponds to a head movement or an eye movement. This type of implementation is potentially advantageous because this combination of events indicates that the user's attention has briefly shifted from the content of the media stream to an event of interest in the real world. For example, there may be some audibility of ambient sound to the user, either actively through the ambient sound via microphone input audio data, or passively due to the incomplete sound blockage provided by headset 305. In some examples, the user may be able to determine that there is activity indicated by the ambient sound, but the ambient sound may not be sufficiently intelligible to engage in a conversation without switching modes or removing headset 305. Based on this ambient sound and / or visual information, the user can generally determine whether there is an event requiring their attention. If so, the user's natural reaction would be to turn their head and / or glance in the direction of the sound source. Specifically, if an audio event from a particular direction is followed by an immediate or nearly immediate head rotation in the direction of the sound event, it is reasonable to assume that the audio event corresponds to the event of interest.

[0116] Accordingly, in some embodiments where the control system is able to determine whether the direction of the sound source corresponds to head movement or eye movement, such a determination would be an example of a mode switch indication. In some examples, the control system may modify the receiving, determining, generating, or mixing process based at least in part on the mode switch indication (see above with reference to FIG. Figure 2 For example, the control system may increase the relative loudness of the microphone output audio data relative to the loudness of the media output audio data. In some such examples, increasing the relative loudness of the microphone output audio data may involve suppressing the media input audio data or pausing the media stream.

[0117] For the sake of computational simplicity, it may be advantageous to have some correspondence between the orientation of the microphone system and the orientation of the inertial sensor system. Figure 3 , microphones 320a and 320b are aligned parallel to an axis of a coordinate system 335 of inertial sensor system 310. In this example, axis 345 passes through microphones 320a and 320b. Here, the y-axis of coordinate system 335 is aligned with headband 330 and parallel to axis 345. In this example, the z-axis of headset coordinate system 335 is aligned vertically relative to the top of headband 330 and the top of inertial sensor system 310. In this embodiment, coordinate system 335 is an x, y, z coordinate system, but other embodiments may use another coordinate system, such as a polar, spherical, or cylindrical coordinate system.

[0118] Other types of mode switching can be based at least in part on the direction of movement of the sound source. If the sound source is determined to be moving toward the user, this can be important for safety reasons. Examples include approaching car noises, footsteps, shouts from a runner, etc.

[0119] Accordingly, some embodiments may involve providing a method based at least in part on signals from two or more microphones (e.g., Figure 3 The direction of movement of the sound source can be determined based on microphone data from microphones 320a and 320b (shown in Figure 3). Such embodiments may involve determining whether a mode switch indication exists based at least in part on the direction of movement of the sound source. If the direction of movement is toward one or more of the microphones of the user's device, then this is an indication that the object generating the sound is moving toward the user. For example, the direction of movement toward the user can be determined based on a noticeable increase in volume of the sound source as the sound source approaches the microphones. Therefore, some embodiments may involve determining a mode switch indication based at least in part on determining the direction of movement of the sound source toward at least one microphone.

[0120] If the sound source is approaching the user and moving above a predetermined speed, this may be more significant in terms of potential danger to the user. Accordingly, some embodiments may involve determining the speed of the sound source and determining a mode switch indication based at least in part on determining that the speed of the sound source exceeds a threshold. For example, the speed of an approaching sound source (e.g., a car) can be determined by measuring the change in volume of the car noise and comparing it to a cubic power increase curve, since power increases as the cube of the decreasing distance between the sound source and the microphone.

[0121] Some mode switching embodiments may involve identifying the individual that the user is concerned about. In some instances, the individual of interest can be indirectly identified, for example, based on the direction of the sound source corresponding to the current position of the individual of interest. In some examples, the direction of the sound source may correspond to a position adjacent to the user, where the individual of interest is located. For example, for the use case of in-cabin movie playback, the user's selected sound source direction may correspond to a seat on the right or left side of the user, where the user's friend is sitting. The control system may be able to determine when an example of sound is received from the selected sound source direction and recognize this example as a mode switching indication. According to some such examples, the control system may be able to control an audio device (e.g., headphones) to pass sound from the selected sound source direction, while sound from other directions will not pass.

[0122] Thus, some embodiments may involve receiving an indication of a selected sound source direction from a user. Such embodiments may involve determining the direction of the sound source based at least in part on microphone data from two or more microphones. Some such embodiments may involve determining that the location of the sound source is a mode switch indication if the location of the sound source corresponds to the selected sound source direction.

[0123] Some mode switching implementations may involve speech recognition of microphone input audio data and / or identifying keywords based on the microphone input audio data recognized as speech. For example, the predetermined keywords may be mode switching instructions. Such keywords may correspond to emergency situations, potential dangers to the user, etc., such as "Help!" or "Caution!"

[0124] Some mode switching embodiments may involve a process of classifying microphone input audio data and basing a mode switching indication at least in part on the classification. Some such mode switching embodiments may involve recognizing the voice of an individual of interest to the user, which voice may also be referred to herein as the voice of a preselected individual. Alternatively or in addition, the classification may involve determining whether the microphone input audio data indicates another sound of potential consequence to the user, such as a car horn, the sound of an approaching vehicle, screaming, yelling, the voice of a preselected individual, a preselected keyword, and / or a public broadcast announcement.

[0125] Although Figure 3 The microphone arrangement shown in may provide satisfactory results, but other embodiments may include other microphone arrangements. Figure 4 An example of a microphone system comprising a pair of overlapping vertically stacked directional microphones is shown. In this example, microphone system 400a comprises an XY stereo microphone system having vertically stacked microphones 405a and 405b, each of which comprises a microphone capsule. Having a known vertical offset between microphones 405a and 405b is potentially advantageous because it allows detection of the time difference between the arrival of corresponding audio signals. Such a time difference can be used to determine the elevation angle of a sound source, for example, as described below.

[0126] In this embodiment, microphone 405a includes microphone capsule 410a and microphone 405b includes microphone capsule 410b. Due to the directional orientation of microphone 405b, microphone capsule 410b is positioned at Figure 4 In this example the longitudinal axis 415a of the microphone capsule 410a extends into and out of the page.

[0127] exist Figure 4 In the example shown in , an xyz coordinate system is shown relative to microphone system 400a. In this example, the z-axis of the coordinate system is the longitudinal axis. Accordingly, in this example, the vertical offset 420a between the longitudinal axis 415a of microphone capsule 410a and the longitudinal axis 415b of microphone capsule 410b extends along the z-axis. However, Figure 4 The orientation of the xyz coordinate system shown in and other coordinate systems disclosed herein is shown by way of example only. In other embodiments, the x or y axis may be the longitudinal axis. In yet other embodiments, a cylindrical or spherical coordinate system may be referenced instead of the xyz coordinate system.

[0128] In this embodiment, the microphone system 400a can be attached to a second device, such as a headset, a smartphone, etc. In some examples, the coordinate system of the microphone system 400a can coincide with the coordinate system of the inertial sensor system, such as Figure 3 . Here, mounting bracket 425 is configured for coupling with a second device. In this example, after microphone system 400a is physically connected to the second device via mounting bracket 425, an electrical connection can be established between microphone system 400a and the second device. Accordingly, audio data corresponding to the sound captured by microphone system 400a can be transmitted to the second device for storage, further processing, reproduction, etc.

[0129] Figure 5 Another example of a microphone system comprising a pair of coincident vertically stacked directional microphones is shown. Microphone system 400b comprises vertically stacked microphones 405e and 405f, each of which comprises a pair of coincident vertically stacked directional microphones. Figure 5 Microphone capsules not visible in FIG: Microphone 405e includes microphone capsule 410e and microphone 405f includes microphone capsule 410f. In this example, the longitudinal axis 415e of microphone capsule 410e and the longitudinal axis 415f of microphone capsule 410f extend in the x, y plane.

[0130] Here, the z-axis extends into and out of the page. In this example, the z-axis passes through the intersection 410 of the longitudinal axis 415e and the longitudinal axis 415f. This geometric relationship is an example of the microphones of the microphone system 400b "coinciding". The longitudinal axis 415e and the longitudinal axis 415f are vertically offset along the z-axis, however, this offset is Figure 5 Longitudinal axis 415e and longitudinal axis 415f are separated by an angle α, which may be 90 degrees, 120 degrees, or another angle depending on the particular embodiment.

[0131] In this example, microphone 405e and microphone 405f are directional microphones. The degree of directionality of a microphone can be represented by a "polar pattern," which indicates how sensitive the microphone is to sounds arriving at different angles relative to the longitudinal axis of the microphone. Figure 5 The polar patterns 405a and 405b illustrated in FIG. 4 represent the locus of points in the microphone that produce the same signal level output, given a given sound pressure level (SPL) generated from that point. In this example, the polar patterns 405a and 405b are cardioid polar patterns. In alternative embodiments, the microphone system may include coincident vertically stacked microphones having supercardioid or hypercardioid polar patterns or other polar patterns.

[0132] As used herein, the directionality of a microphone may sometimes refer to a "front" area and a "back" area. Figure 5 4. Sound source 515a is shown in FIG. 4, located in a region that will be referred to herein as the front region because it is in a region where the microphone is relatively more sensitive, as indicated by the greater extension of the polar pattern along longitudinal axes 415e and 415f. Sound source 515b is located in a region that will be referred to herein as the rear region because it is a region where the microphone is relatively less sensitive.

[0133] Determining the azimuth angle θ corresponding to the sound source direction may be based at least in part on the difference in sound pressure level (which may also be referred to herein as a difference in intensity or amplitude) between the sound captured by microphone capsule 410e and the sound captured by microphone capsule 410f. Some examples are described below.

[0134] Figure 6 An example of azimuth and elevation angles relative to a microphone system comprising a pair of overlapping, vertically stacked directional microphones is shown. For simplicity, only microphone capsules 410g and 410h of microphone system 400d are shown in this example, without supporting structures, electrical connections, etc. Here, the vertical offset 420c between the longitudinal axis 415g of microphone capsule 410g and the longitudinal axis 415h of microphone capsule 410h extends along the z-axis. In this example, the azimuth angle corresponding to the position of a sound source (e.g., sound source 515c) is measured in a plane parallel to the x- and y-planes. This plane may be referred to herein as the "azimuth plane." Accordingly, in this example, the elevation angle is measured in a plane perpendicular to the x- and y-planes.

[0135] Figure 7 705 is a graph showing an example of a curve indicating the relationship between azimuth and the ratio of intensity or level (L / R energy ratio) between right and left microphone audio signals produced by a pair of overlapping vertically stacked directional microphones. The right and left microphone audio signals are examples of first and second microphone audio signals referenced elsewhere herein. In this example, curve 705 corresponds to the relationship between azimuth and L / R ratio for signals produced by a pair of overlapping vertically stacked directional microphones, with longitudinal axes separated by 90 degrees in the azimuth plane.

[0136] refer to Figure 5 , for example, longitudinal axes 415e and 415f are separated by an angle α in the azimuthal plane. Figure 5 , sound source 515a is shown at an azimuth angle θ, which in this example is measured from axis 402 at a position midway between longitudinal axis 415e and longitudinal axis 415f. Curve 705 corresponds to the relationship between azimuth angle and L / R energy ratio for a signal produced by a similar pair of coincident vertically stacked directional microphones, where α is 90 degrees. Curve 710 corresponds to the relationship between azimuth angle and L / R ratio for a signal produced by another pair of coincident vertically stacked directional microphones, where α is 120 degrees.

[0137] It can be observed that Figure 7In the example shown in , both curves 705 and 710 have an inflection point at an azimuth angle of zero degrees, which in this example corresponds to an azimuth angle of a sound source positioned along an axis midway between the longitudinal axis of the left microphone and the longitudinal axis of the right microphone. Figure 7 As shown in , the local maximum occurs at an azimuth of -130 degrees or -120 degrees. Figure 7 In the example shown in , curves 705 and 710 also have local minima corresponding to azimuth angles of 130 and 120 degrees, respectively. The location of these minima depends in part on whether α is 90 or 120 degrees, and also on the directivity pattern of the microphone. Figure 7 The positions of the maxima and minima shown in correspond roughly to the microphone directivity pattern, e.g. Figure 5 The positions of the maxima and minima will be slightly different for microphones with different directivity patterns.

[0138] Reference again Figure 6 , we can see that the sound source 515c is at an elevation angle Located above microphone system 400d. Due to vertical offset 420c between microphone capsule 410g and microphone capsule 410h, sound emitted by sound source 515c will reach microphone capsule 410g before reaching microphone capsule 410h. Consequently, there will be a time difference between the microphone audio signal from microphone capsule 410g in response to the sound from sound source 515c and the corresponding microphone audio signal from microphone capsule 410g in response to the sound from sound source 515c.

[0139] Accordingly, some embodiments may involve determining an elevation angle corresponding to a sound source location based at least in part on a time difference between a first microphone audio signal and a second microphone audio signal. The elevation angle may be determined based on a vertical distance (also referred to herein as a vertical offset) between a first microphone and a second microphone of the pair of coincident vertically stacked directional microphones. According to some embodiments, Figure 1A The control system 110 may be capable of determining an elevation angle corresponding to a location of a sound source based at least in part on a time difference between the first microphone audio signal and the second microphone audio signal.

[0140] Various modifications to the embodiments described herein may be readily apparent to those skilled in the art. The general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the invention. Therefore, the claims are not intended to be limited to the embodiments shown herein, but are to be accorded the widest scope consistent with this invention, the principles, and the novel features disclosed herein.

Claims

1. A method for audio processing, comprising: receiving media input audio data corresponding to a media stream; receiving microphone input audio data from at least one microphone; calculating a media input excitation function for the media input audio data as a time-smoothed power of the media input audio data with a frequency-varying perceptual weight; calculating a microphone input excitation function for the microphone input audio data as a time-smoothed power of the microphone input audio data with a frequency-varying perceptual weight; determining a microphone input specific loudness for the microphone input audio data based at least in part on the microphone input excitation function and in accordance with a specific loudness function that models nonlinearities in human perception of loudness, the microphone input specific loudness corresponding to a perceived loudness of the microphone input audio data; determining a microphone portion specific loudness based at least in part on the media input excitation function and the microphone input excitation function, the microphone portion specific loudness corresponding to the perceived loudness of the microphone input audio data in the presence of the media input audio data; determining frequency-varying and time-varying media processing gains to be applied to the media input audio data to generate media output audio data; determining a frequency-varying and time-varying microphone processing gain to be applied to the microphone input audio data to produce microphone output audio data, wherein the media processing gain and the microphone processing gain are determined such that a first difference between the microphone input specific loudness and a perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the microphone input specific loudness and the microphone portion specific loudness; applying the determined media processing gain to the media input audio data to generate the media output audio data; applying the determined microphone processing gain to the microphone input audio data to generate the microphone output audio data; as well as The media output audio data and the microphone output audio data are mixed to generate mixed audio data. 2 . The method of claim 1 , wherein the perceived loudness of the microphone output audio data in the presence of the media output audio data is substantially equal to the perceived loudness of the microphone input audio data.

3. The method of claim 1, further comprising providing the mixed audio data to speakers of a headset.

4. The method according to claim 1, further comprising: Receive mode switching indication; as well as At least one of the receiving, determining, or mixing process is modified based at least in part on the mode switch indication. 5 . The method of claim 4 , wherein the modifying involves increasing the relative loudness of the microphone output audio data relative to the loudness of the media output audio data. 6 . The method of claim 5 , wherein increasing the relative loudness of the microphone output audio data involves suppressing the media input audio data or pausing the media stream.

7. The method of claim 4, wherein the mode switch indication is based at least in part on at least one of an indication of head movement or an indication of eye movement. The method of claim 4 , wherein the mode switch indication is based at least in part on inertial sensor data.

9. The method of claim 8, wherein the inertial sensor data corresponds to movement of a headset.

10. One or more non-transitory media having stored thereon software comprising instructions that control one or more devices for: receiving media input audio data corresponding to a media stream; receiving microphone input audio data from at least one microphone; calculating a media input excitation function for the media input audio data as a time-smoothed power of the media input audio data with a frequency-varying perceptual weight; calculating a microphone input excitation function for the microphone input audio data as a time-smoothed power of the microphone input audio data with a frequency-varying perceptual weight; determining a microphone input specific loudness for the microphone input audio data based at least in part on the microphone input excitation function and in accordance with a specific loudness function that models nonlinearities in human perception of loudness, the microphone input specific loudness corresponding to a perceived loudness of the microphone input audio data; determining a microphone portion specific loudness based at least in part on the media input excitation function and the microphone input excitation function, the microphone portion specific loudness corresponding to the perceived loudness of the microphone input audio data in the presence of the media input audio data; determining frequency-varying and time-varying media processing gains to be applied to the media input audio data to generate media output audio data; determining a frequency-varying and time-varying microphone processing gain to be applied to the microphone input audio data to produce microphone output audio data, wherein the media processing gain and the microphone processing gain are determined such that a first difference between the microphone input specific loudness and a perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the microphone input specific loudness and the microphone portion specific loudness; applying the determined media processing gain to the media input audio data to generate the media output audio data; applying the determined microphone processing gain to the microphone input audio data to generate the microphone output audio data; as well as The media output audio data and the microphone output audio data are mixed to generate mixed audio data.

11. The one or more non-transitory media of claim 10, wherein the perceived loudness of the microphone output audio data in the presence of the media output audio data is substantially equal to the perceived loudness of the microphone input audio data.

12. An audio processing device, comprising: Interface system; as well as A control system configured to: receiving, via the interface system, media input audio data corresponding to a media stream; receiving microphone input audio data from a microphone system comprising at least one microphone via the interface system; calculating a media input excitation function for the media input audio data as a time-smoothed power of the media input audio data with a frequency-varying perceptual weight; calculating a microphone input excitation function for the microphone input audio data as a time-smoothed power of the microphone input audio data with a frequency-varying perceptual weight; determining a microphone input specific loudness for the microphone input audio data based at least in part on the microphone input excitation function and in accordance with a specific loudness function that models nonlinearities in human perception of loudness, the microphone input specific loudness corresponding to a perceived loudness of the microphone input audio data; determining a microphone portion specific loudness based at least in part on the media input excitation function and the microphone input excitation function, the microphone portion specific loudness corresponding to the perceived loudness of the microphone input audio data in the presence of the media input audio data; determining frequency-varying and time-varying media processing gains to be applied to the media input audio data to generate media output audio data; determining a frequency-varying and time-varying microphone processing gain to be applied to the microphone input audio data to produce microphone output audio data, wherein the media processing gain and the microphone processing gain are determined such that a first difference between the microphone input specific loudness and a perceived loudness of the microphone output audio data in the presence of the media output audio data is less than a second difference between the microphone input specific loudness and the microphone portion specific loudness; applying the determined media processing gain to the media input audio data to generate the media output audio data; applying the determined microphone processing gain to the microphone input audio data to generate the microphone output audio data; as well as The media output audio data and the microphone output audio data are mixed to generate mixed audio data.

13. The audio processing apparatus of claim 12, wherein the perceived loudness of the microphone output audio data in the presence of the media output audio data is substantially equal to the perceived loudness of the microphone input audio data.

14. The audio processing apparatus according to claim 12, wherein the control system is further configured to: receive a mode switch indication; and At least one of the receiving, determining, or mixing process is modified based at least in part on the mode switch indication.

Citation Information

Patent Citations

  • Method and device for personalized hearing

    US20080137873A1

  • Directional sound modification

    US20160165336A1