Audio capture device selection
The method enhances voice processing in wearable audio systems by selecting audio capture devices based on correlation analysis to minimize noise interference, resulting in clearer voice reproduction.
Patent Information
- Application Number
- PCT/EP2025/069010
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-03
- Publication Date
- 2026-02-05
AI Technical Summary
Existing wearable audio listening systems face issues with external noise interference, such as wind noise and ambient noise, which deteriorate voice-processing algorithms and render speech unintelligible.
A method for selecting audio capture devices based on correlation analysis of signals from multiple microphones to identify and minimize unwanted noise, using techniques like beamforming and noise reduction algorithms to enhance voice processing.
Improves voice input processing by selecting audio signals with the least amount of unwanted noise, ensuring clearer reproduction of user voice.
Smart Images

Figure EP2025069010_05022026_PF_FP_ABST
Abstract
Description
[0001] AUDIO CAPTURE DEVICE SELECTION
[0002] Technical Field
[0003] The present application relates to audio capture device selection. In particular, the application relates to selection of an audio capture device for use in determining an output signal from a wearable audio listening system.
[0004] Background
[0005] A wearable audio listening system such as a set of headphones may use a number of audio capture devices such as microphones to capture audio, in particular the voice of a user. This enables the wearable audio listening system to be used for voice calls on a connected device, for example a smart phone, to receive voice commands for operation of the connected device, and the like. The use of multiple microphones to capture such audio may typically involve a voice-processing algorithm such as beamforming to generate a single signal from the multiple microphone signals, or denoising filtering steps using neural networks and other Artificial Intelligence approaches.
[0006] However, external noise (i.e. anything else but the voice) picked up by one or more of the microphones will deteriorate the output of the voice-processing algorithm and in some cases render the speech unintelligible. Examples of external noise might be: wind noise (typically generated by turbulence at the microphone inlets), vibration of part of the wearable audio listening system, distractor noise (somebody else talking), traffic noise, and other ambient noise e.g. from a restaurant, coffee shop, office, kitchen, etc.
[0007] It is therefore desired to provide a method for selecting audio capture devices in a wearable audio listening system that attempts to mitigate at least some of these issues.
[0008] Summary
[0009] This disclosure provides systems, methods and other approaches for selecting which audio capture devices of a wearable audio listening system should be used in generating an output signal. In particular, audio signals captured by different audio capture devices are analysed to determine which signals include the most unwanted audio, such as background noise, wind, etc. Based on the analysis, audio signals with the least amount of unwanted audio can be selected for use in generating an output signal including the user’s voice. Other signals can be ignored, or weighted lower, so that the unwanted audio does not adversely affect voice processing of the generated output signal. As a result, processing of voice inputs from a user is improved.
[0010] According to a first aspect of the disclosure, there is provided a computer-implemented method for selecting at least one audio capture device for use in determining an output signal from a wearable audio listening system, the computer-implemented method comprising, by processing circuitry of a computer system, acquiring a first sum of one or more first signals representing audio captured by one or more respective left-side audio capture devices of the wearable audio listening system, acquiring a second sum of one or more second signals representing audio captured by one or more respective right-side audio capture devices of the wearable audio listening system, determining a third signal based on a difference between the first sum and the second sum, determining a correlation between the third signal and each of the first and second sums and / or between the third signal and each of the first and second signals, and selecting at least one audio capture device for use in determining an output signal based on the determined correlations.
[0011] Optionally, acquiring the first sum comprises acquiring a single first signal, acquiring the second sum comprises acquiring a single second signal, the computer-implemented method comprising: determining the third signal based on a difference between the first and second signals, determining a first correlation between the third signal and the first signal, and determining a second correlation between the third signal and the second signal.
[0012] Optionally, the computer-implemented method further comprises filtering the first signal and the second signal, determining the third signal based on a difference between filtered first and second signals, determining a first correlation between the third signal and the filtered first signal, and determining a second correlation between the third signal and the filtered second signal.
[0013] Optionally, acquiring the first sum comprises acquiring a sum of a plurality of first signals, and acquiring the second sum comprises acquiring a sum of a plurality of second signals. Optionally, the computer-implemented method further comprises time synchronising the plurality of first signals and acquiring a sum of the plurality of time- synchronised first signals, and / or time synchronising the plurality of second signals and acquiring a sum of the plurality of time-synchronised second signals.
[0014] Optionally, the computer-implemented method comprises determining a first correlation between the third signal and the first sum, and determining a second correlation between the third signal and the second sum. Optionally, the computer-implemented method further comprises filtering the first sum and the second sum, determining the third signal based on a difference between filtered first and second sums, determining a first correlation between the third signal and the filtered first sum, and determining a second correlation between the third signal and the filtered second sum.
[0015] Optionally, the computer-implemented method comprises determining a respective correlation between the third signal and each of the plurality of first signals, and determining a respective correlation between the third signal and each of the plurality of second signals. Optionally, the computer-implemented method further comprises filtering the first sum and the second sum, filtering each of the plurality of first and second signals, determining the third signal based on a difference between the filtered first and second sums, determining a respective correlation between the third signal and each of the plurality of filtered first signals, and determining a respective correlation between the third signal and each of the plurality of filtered second signals. Optionally, the computer-implemented method further comprises time synchronising the plurality of first signals and determining a respective correlation between the third signal and each of the plurality of time-synchronised first signals, and / or time synchronising the plurality of second signals and determining a respective correlation between the third signal and each of the plurality of time-synchronised second signals.
[0016] Optionally, determining a correlation comprises determining a cross-correlation. Optionally, the computer-implemented method further comprises determining a time average of a result of the correlation.
[0017] Optionally, selecting at least one audio capture device for use in determining an output signal based on the determined correlations comprises: acquiring a first correlation result associated with the one or more left-side audio capture devices, acquiring a second correlation result associated with the one or more right-side audio capture devices, comparing the first and second correlation results, and selecting the one or more left-side audio capture devices or the one or more right-side audio capture devices for use in determining the output signal based on determining that one of the first correlation result and the second correlation result is higher than the other.
[0018] Optionally, selecting at least one audio capture device for use in determining an output signal based on the determined correlations comprises: acquiring a first correlation result associated with the one or more left-side audio capture devices, acquiring a second correlation result associated with the one or more right-side audio capture devices, determining a difference between the first and second correlation results, determining a sum of the first and second correlation results, determining a ratio between the difference and the sum of the first and second correlation results, and determining, based on the determined ratio, a gain to be applied to each of the one or more left-side audio capture devices and / or each of the one or more right-side audio capture devices for use in determining the output signal. Optionally, a sum of the respective gains of the one or more left-side audio capture devices and the one or more right-side audio capture devices is equal to one. Optionally, selecting at least one audio capture device for use in determining an output signal based on the determined correlations comprises selecting an audio capture device having a threshold signal to noise ratio.
[0019] Optionally, the wearable audio listening system comprises integrated left-side and rightside audio listening devices. Optionally, the wearable audio listening system comprises a left-side wearable audio listening device and a right-side wearable audio listening device separate to the left-side wearable audio listening device. Optionally, the wearable audio listening system comprises one or more sensors configured to determine if the one or more left-side and / or right-side audio capture devices are in a worn position.
[0020] According to a second aspect of the disclosure, there is provided a computer- implemented method for selecting an audio capture device for use in determining an output signal from a wearable audio listening system, the computer-implemented method comprising, by processing circuitry of a computer system, acquiring a first signal representing audio captured by a left-side audio capture device of the wearable audio listening system, acquiring a second signal representing audio captured by a right-side audio capture device of the wearable audio listening system, determining a third signal based on a difference between the first signal and the second signal, determining a first correlation between the third signal and the first signal, determining a second correlation between the third signal and the second signal, and selecting at least one audio capture device for use in determining an output signal based on the determined correlations.
[0021] According to a third aspect of the disclosure, there is provided a computer-implemented method for selecting an audio capture device for use in determining an output signal from a wearable audio listening system, the computer-implemented method comprising, by processing circuitry of a computer system, acquiring a first sum of a plurality of first signals representing audio captured by a plurality of respective left-side audio capture devices of the wearable audio listening system, acquiring a second sum of a plurality of second signals representing audio captured by a plurality of respective right-side audio capture devices of the wearable audio listening system, determining a third signal based on a difference between the first sum and the second sum, determining a first correlation between the third signal and the first sum, determining a second correlation between the third signal and the second sum, and selecting at least one audio capture device for use in determining an output signal based on the determined correlations.
[0022] According to a fourth aspect of the disclosure, there is provided a computer- implemented method for selecting an audio capture device for use in determining an output signal from a wearable audio listening system, the computer-implemented method comprising, by processing circuitry of a computer system, acquiring a first sum of a plurality of first signals representing audio captured by a plurality of respective leftside audio capture devices of the wearable audio listening system, acquiring a second sum of a plurality of second signals representing audio captured by a plurality of respective right-side audio capture devices of the wearable audio listening system, determining a third signal based on a difference between the first sum and the second sum, determining a respective first correlation between the third signal and each of the plurality of first signals, determining a respective second correlation between the third signal and each of the plurality of second signals, and selecting at least one audio capture device for use in determining an output signal based on the determined correlations.
[0023] According to a fifth aspect of the disclosure, there is provided a computer system for selecting at least one audio capture device for use in determining an output signal from a wearable audio listening system, the computer system comprising processing circuitry configured to acquire a first sum of one or more first signals representing audio captured by one or more respective left-side audio capture devices of the wearable audio listening system, acquire a second sum of one or more second signals representing audio captured by one or more respective right-side audio capture devices of the wearable audio listening system, determine a third signal based on a difference between the first sum and the second sum, determine a correlation between the third signal and each of the first and second sums and / or between the third signal and each of the first and second signals, and select at least one audio capture device for use in determining an output signal based on the determined correlations.
[0024] According to a sixth aspect of the disclosure, there is provided a wearable audio listening system comprising the computer system of the fifth aspect of the disclosure.
[0025] According to a seventh aspect of the disclosure, there is provided a computer program product comprising program code for performing, when executed by processing circuitry, the computer-implemented method of any of the first to fourth aspects of the disclosure.
[0026] According to a seventh aspect of the disclosure, there is provided a non-transitory computer-readable storage medium comprising instructions, which when executed by processing circuitry, cause the processing circuitry to perform the computer- implemented method of any of the first to fourth aspects of the disclosure.
[0027] The disclosed aspects, examples (including any preferred examples), and / or accompanying claims may be suitably combined with each other as would be apparent to anyone of ordinary skill in the art. Additional features and advantages are disclosed in the following description, claims, and drawings, and in part will be readily apparent therefrom to those skilled in the art or recognized by practicing the disclosure as described herein.
[0028] Brief Description of the Drawings.
[0029] Examples are described in more detail below with reference to the appended drawings.
[0030] FIG. 1A illustrates a top-down view of a user in a listening environment according to an example.
[0031] FIG. 1 B illustrates a side view of a wearable audio listening system according to an example.
[0032] FIG. 2 is a flow chart of a computer-implemented method according to an example. FIG. 3 is a signalling diagram for a wearable audio listening system comprising single left- and right-side audio capture devices according to an example.
[0033] FIG. 4 is a signalling diagram for a wearable audio listening system comprising multiple left- and right-side audio capture devices according to an example.
[0034] FIG. 5 is a signalling diagram for a wearable audio listening system comprising multiple left- and right-side audio capture devices according to an example.
[0035] FIGs. 6A to 60 illustrate top-down view of a user in different listening environments according to a number of examples.
[0036] FIG. 7 is a schematic diagram of a computer system for implementing examples disclosed herein.
[0037] Like reference numerals refer to like elements throughout the description.
[0038] DETAILED DESCRIPTION
[0039] The detailed description set forth below provides information and examples of the disclosed technology with sufficient detail to enable those skilled in the art to practice the disclosure.
[0040] FIG. 1A illustrates a top-down view of a user U in a listening environment 100. The user U is wearing a wearable audio listening system 102 comprising left and right audio listening devices 104, 106. The user U may use the wearable audio listening system 102 to listen to audio, such as music, podcasts and the like. For example, the wearable audio listening system 102 may comprise electronics for receiving signals from an associated electronic media playback device (not shown) to enable audio to be produced by the audio listening devices 104, 106. In particular, the wearable audio listening system 102 may comprise a control unit 108 comprising processing circuitry configured to implement functions of the wearable audio listening system 102 such as those described in this disclosure.
[0041] The media playback device may be a smartphone, tablet, laptop computer, desktop computer, or any other device capable of controlling audio playback via the wearable audio listening system 102. The wearable audio listening system 102 may communicate with the media playback device through a wired or wireless connection. The audio listening devices 104, 106 may each comprise a sound producing device, for example a speaker, for producing the audio. The wearable audio listening system 102 further comprises a number of audio capture devices 110, 112 such as microphones. In particular, the wearable audio listening system 102 comprises at least one left-side audio capture device 110 and at least one right-side audio capture device 112. The audio capture devices 110, 112 are disposed on respective audio listening devices 104, 106. The user II may use the audio capture devices 110, 112 to provide voice inputs, for example a voice call, voice recording, or commands relating to operation of the media playback device.
[0042] FIG. 1 B illustrates a side view of a wearable audio listening system 102, in particular a left audio listening device 104. As discussed above, at least one left-side audio capture device 110 is disposed on the left audio listening device 104. In the example of FIG. 1 B, two left-side audio capture devices 110a, 110b, are disposed on the left audio listening device 104. The first left-side audio capture device 110a may be positioned substantially towards the front of the left audio listening device 104 for proximity to the mouth of the user II. This positioning provides for better capture of voice inputs issued by the user II. The second left-side audio capture device 110b may be positioned elsewhere on the left audio listening device 104.
[0043] In some examples, the left-side audio capture devices 110a, 110b may be aligned such that an axis between them, indicated by the dashed arrow, is directed towards the mouth of the user II. This is an effective configuration for two-microphone beamforming (known as an end-fire configuration). In this configuration, the signal from the mouth of the user II arrives first at the first left-side audio capture device 110a (closer to the mouth) and then with a delay at the second left-side audio capture device 110b (further from the mouth), where the delay is equal to the acoustic propagation time between the two microphone locations. By summing a delayed version of the signal from the first left-side audio capture device 110a with the signal from the second left-side audio capture device 110b, an amplification of all acoustic sources aligned with the two audio capture devices in the forward direction is provided. A practical beamforming algorithm, while usually more complex, may operate on the same principle.
[0044] The second left-side audio capture device 110b may be considered as a secondary audio capture device for better capture of voice inputs issued by the user II. The first and / or second audio capture device 110a, 110b may be a forward facing microphone, in order to better capture voice inputs issued by the user II. The first and / or second audio capture device 110a, 110b may be an active noise cancellation (ANC) microphone, configured to reduce unwanted ambient sounds by generates an anti-noise signal to counteract any ambient audio signals. Whilst FIG. 1 B illustrates a side view of a left audio listening device 104, it will be appreciated that a similar construction and implementation may be employed for the right audio listening device 106.
[0045] In the example of FIGs. 1A and 1 B, the wearable audio listening system 102 is a set of over-ear headphones comprising a left cup 104 and a right cup 106 joined by a headband. In other examples, the wearable audio listening system 102 may comprise other types of audio listening devices, such as in-ear headphones. In some examples, the audio listening devices may be joined in other ways, such as by a neck-band, or may be separate devices, such as in the case of wireless, in-ear headphones.
[0046] In some examples, the wearable audio listening system 102 may further comprise one or more motion and / or orientation sensors (not shown) such as accelerometers, gyroscopes, and the like. These sensors may be configured to capture data indicative of a position and / or orientation of the wearable audio listening system 102 and / or the individual audio listening devices 104, 106. In this way, it is possible to determine if the audio listening devices 104, 106 (and, by association, the audio capture devices 110, 112) are in a worn position.
[0047] Returning to FIG. 1A, a number of different audio signals are present in the listening environment 100. A first audio signal 114 is the voice of the user II. Further audio signals include external audio signals 116a-d. Examples of external sources of external audio signals 116a-d include wind noise, vibration of part of the wearable audio listening system 102, distractor noise (somebody else talking), traffic noise, and other ambient noise e.g. from a restaurant, coffee shop, office, kitchen, etc. The audio capture devices 110, 112 may capture any of the different audio signals 114, 116a-d. As such, when the user II provides a voice input as part of the first audio signal 114, the external audio signals 116a-d may interfere with the captured audio. This may will deteriorate the output of the voice-processing algorithm and in some cases render the voice input unintelligible.
[0048] The acoustic radiation pattern present in the listening environment 100, made up of the audio signals 114, 116a-d, can be considered as mid (M) and side (S) signals as an alternative to the classical representation of left (L) and right (R) signals. This is based on the concepts described by A. D. Blumlein, and can be described as a spatial, or “true directional”, representation of an acoustic radiation. This is essentially a way to emulate human hearing, where humans detect phase and intensity differences in a sound field as the sound waves arrive at each of our left and right ears. The brain can then use this to determine the direction from which the sound is coming.
[0049] In these relations, the M signal represents the centre of a stereo image and the S signal represents the edges of that image. The M signal is a (weighted) sum of the left and right signals, whereas the S signal is a (weighted) difference between the left and right signals. These signals can be formulated as follows:
[0050] M = L + R
[0051] S = L — R
[0052] These representations are mathematically equivalent (the transformation from the L / R to the M / S base is bijective), as the left and right signals can be recovered from the M and S signals as
[0053] The above equations can be applied to an electronic signal registered from an acoustic domain, and have an inherent psycho-acoustic property (how they will be perceived by a listener when they are played back) that is dependent on how the registering microphones are placed geometrically in relation to each other and to the sound field to be captured. Assuming a symmetrical arrangement of audio capture devices on the left and right audio listening devices M and S signals representative of the sound field around the head can be determined.
[0054] As the mouth of the user II is substantially centred in the plane of the audio capture device arrangement, the voice of the user II (first audio signal 114) will be mainly represented in the M signal, while other audio (external audio signals 116a-d) will be mainly represented in the S signal. Put otherwise, the M signal will have higher signal-to-noise ratio (SNR) than any of the individual audio capture devices 110, 112, where the SNR may defined as the ratio of the respective signal to the S signal. As such, if the S signal can be removed from the captured audio in some way, then a cleaner representation of the voice of the user II can be obtained.
[0055] FIG. 2 is a flow chart of a computer-implemented method 200 according to an example. The method 200 is for selecting at least one audio capture device 110, 112 for use in determining an output signal from a wearable audio listening system 102. The method 200 enables audio capture devices 110, 112 to be selected based on the amount of unwanted audio that they capture, which allows an output signal to be generated where the voice of a user II is clearly reproduced. As a result, processing of voice inputs from the user II is improved. The method 200 may be implemented by processing circuitry of a computer system (e.g. the control unit 108 described in relation to FIG. 1).
[0056] At step 202, a first sum of one or more audio signals is acquired. In particular, the first sum is a sum of audio signals captured by one or more respective left-side audio capture devices 110 of the wearable audio listening system 102. In the case that the wearable audio listening system 102 comprises a single left-side audio capture devices 110, the audio signal captured by that device is used as the first sum (i.e. there are no other signals to be taken into account). In the case that the wearable audio listening system 102 comprises multiple left-side audio capture devices 110a-b, a sum is taken of the audio signals captured by each left-side audio capture device 110a-b. The first sum can be considered as a left signal of a classical stereo representation.
[0057] At step 204, a second sum of one or more audio signals is acquired, in this case audio signals captured by one or more respective right-side audio capture devices 112 of the wearable audio listening system 102. Again, in the case that the wearable audio listening system 102 comprises a single right-side audio capture device 112, the audio signal captured by that device is used as the second sum, while in the case that the wearable audio listening system 102 comprises multiple right-side audio capture devices, a sum is taken of the audio signals captured by each right-side audio capture device. The second sum can be considered as a right signal of a classical stereo representation.
[0058] At step 206, a difference signal is determined between the first sum acquired at step 202 and the second sum acquired at step 204. The difference signal can be considered as an S signal in the Blumlein representation discussed above, as it is a difference between the left and right signals from the classical stereo representation.
[0059] At step 208, a correlation is determined between the difference signal and each of the first and second sums. As discussed above, the S signal has most of the unwanted audio (external audio signals 116a-d). Therefore, a high correlation to the difference signal determined at 206 indicates that a signal includes a high amount of the unwanted audio. Based on this, it can be determined which of the first and second sums includes more of the unwanted audio. The correlation is determined to provide a measure of similarity between the respective signals. Any suitable method of determining a correlation or similarity may be used. In this disclosure, the term correlation is intended to cover determining a similarity in the time domain and / or the frequency domain. In some examples, a cross-correlation may be determined directly. In some examples, a correlation may be determined in the frequency domain, for example using a fast Fourier transform (this may also be referred to as determining a coherence between the signals). Determining correlation in the time domain is less computationally expensive, whereas determining a coherence may involve determining an average similarity over time, which is less sensitive to phase synchronisation (or time synchronisation).
[0060] At step 210, which can be performed additionally or alternatively to step 208, a correlation is determined between the difference signal and each individual signal from the audio capture devices 110, 112. The correlation may be performed in substantially the same way as in step 208, although in this case the individual signals are used instead of the sums. This can be useful in cases where there are multiple audio capture devices 110, 112 on each side. As a result of this correlation, it can be determined how much of the unwanted audio is being captured by each individual audio capture device 110, 112.
[0061] At step 212, at least one audio capture device 110, 112 is selected for use in determining an output signal based on the determined correlations. For example, if it is determined at step 208 that the first sum has a higher correlation to the difference signal than the second sum, it can be determined that the first sum includes more of the unwanted audio. This may be indicative of more external audio on the left side of the wearable audio listening system 102, for example that the left side is particularly exposed to wind, someone talking, or in general more exposed to an external noise source. In this case, one or more of the right-side audio capture devices 112 can be selected for use in determining an output signal. In some examples, a correlation threshold may be used, whereby any signals having a correlation above the threshold (i.e. indicating a sufficiently high correspondence to the “noisy” difference signal) are discounted.
[0062] In the case that the wearable audio listening system 102 comprises single left- and right-side audio capture devices 110, 112, one of the audio capture devices 110, 112 can be selected for use in determining an output signal. In the case that the wearable audio listening system 102 comprises multiple left- and right-side audio capture devices 110, 112, different combinations of devices can be selected as appropriate. For example, only left-side audio capture devices 110 may be selected, for example all left-side audio capture devices 110. Alternatively, only right-side audio capture devices 112 may be selected, for example all rightside audio capture devices 112. In some examples, a combination of left- and right-side audio capture devices 110, 112 may be selected.
[0063] Once an audio capture device 110, 112 is selected, the audio signal captured by that device can then be used to generate an output signal including the voice of the user II, as will be discussed below. The output signal may then be used in processing a voice input from a user II, for example as an input to a voice processing algorithm. In some examples, this may involve a beamforming step to focus the voice capture on the direction of the mouth of the user II, and a subsequently applied noise reduction algorithm, for example noise gating and / or Al noise reduction. In some examples, a noise reduction algorithm may receive input signals from several selected audio capture devices 110, 112 and output a single signal.
[0064] In some examples, an audio capture device 110, 112 may be selected based on an SNR of its captured audio signal (e.g. the ratio of the signal to the S signal). In particular, only audio capture devices 110, 112 having a captured audio signal with an SNR above a threshold value may be selected. In this way, if the level of a signal coming from an audio capture device 110, 112 is below a minimum threshold (e.g. close to noise floor of the audio capture device 110, 112, for example if the audio capture device 110, 112 is defective or blocked by water, dust, etc.), the audio capture device 110, 112 can be excluded from selection. The threshold may be an absolute value considered to provide a sufficiently high SNR to interpret the voice of the user II. A similar result could also be achieved by applying a threshold to the captured audio signal directly.
[0065] In some examples, an audio capture device 110, 112 may only be selected if it is determined that the associated audio listening device 104, 106 is in a worn position. As discussed above, one or more motion and / or orientation sensors may be configured to capture data indicative of a position and / or orientation of the wearable audio listening system 102 and / or the individual audio listening devices 104, 106. If it is determined that an audio listening device 104, 106 (and, by association, an audio capture device 110, 112) is not in a worn position, for example if the wearable audio listening system 102 is worn sideways or sloppily), it can be disregarded from selection.
[0066] In examples where at least one audio capture device from each side of the wearable audio listening system 102 is selected, an M signal can be determined. In the case of symmetrically placed audio capture devices 110, 112, this may be achieved by determining a sum of the signals as discussed above. In the case of a plurality of audio capture devices 110, 112 on each side, this may be achieved by first combining all signals on each side into one signal using a delay-and-sum algorithm and then summing the resulting left and right signals. When the background noise is distributed substantially evenly around the user II, the M signal will have higher SNR than any of the individual audio capture devices 110, 112. The M signal could then be used as the output signal.
[0067] However, when the background noise is not distributed evenly around the user II, one or more of the individual audio capture devices 110, 112 may have a higher SNR than the M signal. In that instance, the signals from the individual audio capture devices 110, 112 can be used. For example, in the case that a single audio capture device 110, 112 is selected, the audio signal from that device can be used as the output signal. In examples where a plurality of audio capture devices 110, 112 are selected, a weight sum of the respective signals can be determined. This can be achieved in a number of ways, as discussed below.
[0068] The generated output signal can then be used for voice processing, for example as an input to a voice processing algorithm. The method 200 enables audio signals captured by different audio capture devices to be analysed to determine which signals include the most unwanted audio, such as background noise, wind, etc. Based on the analysis, audio signals with the least amount of unwanted audio can be selected for use in generating an output signal including the user’s voice. Other signals can be ignored, or weighted lower, so that the unwanted audio does not adversely affect voice processing of the generated output signal. As a result, processing of voice inputs from a user II is improved.
[0069] The method 200 may be performed continuously or at intervals in order to take into account any changes in the listening environment 100 or the background noise. If a change in the selected audio capture devices is determined, a fade may be applied during the signal transition between devices.
[0070] FIG. 3 is a signalling diagram illustrating an example where the wearable audio listening system 102 comprises single left- and right-side audio capture devices 110, 112. As shown in FIG. 3, the left-side audio capture device 110 provides (captures) a first audio signal 302, while the right-side audio capture device 112 provides (captures) a second audio signal 304.
[0071] The audio signals 302, 304 are provided to a difference determiner 306. The difference determiner 306 is configured to determine a difference between the first audio signal 302 and the second audio signal 304. This corresponds to step 206 of the method 200. The difference determiner 306 produces a difference signal 308. As discussed above, the difference signal 308 can be considered as an S signal in the Blumlein representation discussed above, as it is a difference between the left and right signals from the classical stereo representation.
[0072] In some examples, and as illustrated in FIG. 3, the audio signals 302, 304 may optionally be filtered at a filter 310 before being provided to the difference determiner 306. A filter 310 may be configured to remove any noise that may degrade the output of a voice processing algorithm, and may ensure only the relevant frequency range for speech is taken into account. For example, a filter 310 may be configured to remove any DC and / or low frequency noise having similar properties to the voice of the user II. A filter 310 may additionally or alternatively be configured to remove any high frequency noise (e.g. wind). A filter 310 may be a band-pass filter, for example a biquad filter of 1stor 2ndorder. The band-pass filter may have a band-pass frequency range substantially corresponding to the human vocal range (e.g. 200 to 4000Hz). The filters 310 provide respective filtered signals 302a, 304a which are then provided to the difference determiner 306. The filtering may ensure only the relevant frequency range for speech is taken into account.
[0073] The difference signal 308 is then provided to a correlation determiner 312. The correlation determiner 312 also receives the first audio signal 302 and the second audio signal 304. While two instances of the correlation determiner 312 are shown in FIG. 3, it will be appreciated that the correlation determiner 312 may be embodied by a single module of a computer system. In examples where the difference signal 308 is determined based on the filtered signals 302a, 304a, the correlation determiner 312 receives the first filtered audio signal 302a and the second filtered signal 304a. The correlation determiner 312 then determines a correlation between the received signals. In this example, this corresponds to step 210 of the method 200, as the individual (filtered) signals from the audio capture devices 110, 112 are used to determine a correlation. A first correlation result 314 is determined for the left-side audio capture device 110, and a second correlation result 316 is determined for the right-side audio capture device 112.
[0074] In some examples, and as illustrated in FIG. 3, the correlation results 314, 316 are time- averaged. In particular, the correlation results 314, 316 may be provided to a time-averager 318 that is configured to provide respective time-averaged correlation results 314a, 316a. Any suitable time-averaging approach may be employed. For example, a root mean square (RMS) time-average of the correlation results 314, 316 may be determined. In other examples, a timeaverage of the maximum absolute value the correlation results 314, 316 may be determined. The time average may be determined using one or more of, for example, a low-pass filter, a moving-window, and the like. While two instances of the time-averager 318 are shown in FIG. 3, it will be appreciated that the time-averager 318 may be embodied by a single module of a computer system.
[0075] The (time-averaged) correlation results 314, 316 are then provided to a comparator 320. The comparator 320 may determine which of the correlation results 314, 316 is higher, i.e. which of the (filtered) audio signals 302, 304 has the highest correlation to the difference signal 308. As discussed above, the audio signal 302, 304 having the highest correlation can be considered to include more of the unwanted audio. Similarly, the audio signal 302, 304 having the lowest correlation can be considered to include less of the unwanted audio. The comparator 320 may then provide a comparison result 322.
[0076] The comparison result 322 is provided to a selector 324, which is configured to select at least one of the audio capture devices 110, 112 for use in determining an output signal 326. This corresponds to step 212 of the method 200. In some examples, only one of the audio capture devices 110, 112 is selected for use in determining an output signal 326. For example, the audio capture devices 110, 112 having the audio signal 302, 304 with the lowest correlation to the difference signal 308 may be selected. In some examples, a correlation threshold may be used to discount any signals having a particularly high correlation to the difference signal 308. This may be applied at the comparator 320 or the selector 324.
[0077] In some examples, both audio capture devices 110, 112 are selected. In this case, an M signal may be generated based on the left and right audio signals 302, 304 and used as the output 326. In other examples, the output 326 may be determined using a weighted sum of the audio signals 302, 304. For example, the audio signal 302, 304 having the highest correlation to the difference signal 308 may be weighted less than the other audio signal 302, 304. This will be discussed in more detail below. In some examples, the output 326 may be determined using a multiplexor with built-in smoothing. This progressively transitions between the signals when receiving a command to change between the selected signals. That reduces clipping and sharp transitions, for example when switching from using the left-side audio capture device 110 to using the right-side audio capture device 112.
[0078] FIG. 4 is a signalling diagram illustrating an example where the wearable audio listening system 102 comprises multiple left- and right-side audio capture devices 110a-b, 112a-b. As shown in FIG. 4, the first left-side audio capture device 110a provides (captures) a first left audio signal 401 , the second left-side audio capture device 110b provides (captures) a second left audio signal 402, the first right-side audio capture devices 112a provides (captures) a first right audio signal 403, and the second right-side audio capture device 112b provides (captures) a second right audio signal 404.
[0079] The left audio signals 401 , 402 are provided to a summer 430. The summer 430 is configured to determine a sum of first left audio signal 401 and the second left audio signal 402. This corresponds to step 202 of the method 200. The summer 430 produces a first sum 405. In some examples, one or both of the left audio signals 401 , 402 are provided to a delayer / synchroniser 428. The delayer / synchroniser 428 is configured to align the left audio signals 401 , 402 in time (time synchronisation) before they are provided to the summer 430. This may be useful when the left-side audio capture devices 110a-b are at different distances from the mouth of the user II, causing a delay between the signals 401 , 402. In the example shown in FIG. 4, the first left audio signal 401 is provided to a delayer / synchroniser 428, which provides a delayed first left audio signal 401a to the summer 430.
[0080] Similarly, the right audio signals 403, 404 are provided to a summer 430. The summer 430 is configured to determine a sum of first right audio signal 403 and the second right audio signal 404. This corresponds to step 204 of the method 200. The summer 430 produces a second sum 406. In some examples, one or both of the right audio signals 403, 404 are provided to a delayer / synchroniser 428 configured to align the right audio signals 403, 404 in time (time synchronisation). In the example shown in FIG. 4, the first right audio signal 403 is provided to a delayer / synchroniser 428, which provides a delayed first right audio signal 403a to the summer 430. While two instances of the summer 430 are shown in FIG. 4, it will be appreciated that the summer 430 may be embodied by a single module of a computer system.
[0081] The sums 405, 406 are provided to a difference determiner 407. The difference determiner 407 is configured to determine a difference between the first sum 405 and the second first right audio signal 406. This corresponds to step 206 of the method 200. The difference determiner 407 may function in substantially the same way as the difference determiner 306 described in relation to FIG. 3. The difference determiner 407 produces a difference signal 408. As discussed above, the difference signal 408 can be considered as an S signal in the Blumlein representation discussed above, as it is a difference between the left and right signals from the classical stereo representation.
[0082] In some examples, and as illustrated in FIG. 4, the sums 405, 406 may optionally be filtered at a filter 410 before being provided to the difference determiner 407. A filter 410 may be configured to remove any noise that may degrade the output of a voice processing algorithm, and may function in substantially the same way as the filters 310 described in relation to FIG. 3. The filters 410 provide respective filtered sums 405a, 406a which are then provided to the difference determiner 306.
[0083] The difference signal 408 is then provided to a correlation determiner 412. The correlation determiner 412 also receives the first sum 405 and the second sum 406. The correlation determiner 412 may function in substantially the same way as the correlation determiner 312 described in relation to FIG. 3. In examples where the difference signal 308 is determined based on the filtered sums 405a, 406a, the correlation determiner 312 receives the first filtered sum 405a and the second filtered sum 406a. The correlation determiner 412 then determines a correlation between the received signals. In this example, this corresponds to step 208 of the method 200, as the first and second sums 405, 406 are used for the correlation. A first correlation result 414 is determined for the left-side audio capture devices 110a-b, and a second correlation result 416 is determined for the right-side audio capture devices 112a-b.
[0084] In some examples, and as illustrated in FIG. 4, the correlation results 414, 416 are time- averaged. In particular, the correlation results 414, 416 may be provided to a time-averager 418 that is configured to provide respective time-averaged correlation results 414a, 416a. The time-averager 418 may function in substantially the same way as the time-averager 318 described in relation to FIG. 3.
[0085] The (time-averaged) correlation results 414, 416 are then provided to a comparator 420. The comparator 420 may determine which of the correlation results 414, 416 is higher, i.e. which of the (filtered) first and second sums 405, 406 has the highest correlation to the difference signal 408. As discussed above, if a sum 405, 406 has the highest correlation, it can be considered that the respective side of the wearable audio listening system 102 captures more of the unwanted audio. Similarly, if a sum 405, 406 has the lowest correlation, it can be considered that the respective side of the wearable audio listening system 102 captures less of the unwanted audio. The comparator 420 may then provide a comparison result 422.
[0086] The comparison result 422 is provided to a selector 424, which is configured to select at least one of the audio capture devices 110a-b, 112a-b for use in determining an output signal 426. This corresponds to step 212 of the method 200. In some examples, only one side of the wearable audio listening system 102 (i.e. either the left-side audio capture devices 110a-b or the right-side audio capture devices 112a-b) are selected for use in determining an output signal 426. In some examples, only one of the audio capture devices 110a-b, 112a-b is selected. In some examples, a number of the audio capture devices 110a-b, 112a-b are selected (for example one from each side, both from one side and one from the other side, or all audio capture devices). In examples where audio capture devices 110a-b, 112a-b from both sides are selected, an M signal may be generated based on the left and right audio signals 302, 304 and used as the output 426. Alternatively, the output 426 may be determined using a weighted sum of the audio signals 401 , 402, 403, 404, as will be discussed in more detail below. In one example, where all four audio capture devices 110a-b, 112a-b are selected, the first audio capture devices 110a, 112a may be used to determine a primary output (for example via a sum operation), while the second audio capture devices 110b, 112b may be used to determine a secondary output (for example via a sum operation). In some examples, a correlation threshold may be used to discount any signals having a particularly high correlation to the difference signal 408. This may be applied at the comparator 420 or the selector 424.
[0087] FIG. 5 is a signalling diagram illustrating another example where the wearable audio listening system 102 comprises multiple left- and right-side audio capture devices 110a-b, 112a-b. In this example, the signal processing performed before the correlation is determined is substantially the same as described in relation to FIG. 4. That is to say, the summer 430, delayer / synchroniser 428, difference determiner 407, and filters 410 may function in substantially the same way as described above, although it is noted that the delayer / synchroniser 428 and filters 410 have been omitted from FIG. 5.
[0088] In the example of FIG. 5, the correlation determiner 412 receives the difference signal 408 and the individual audio signals 401 , 402, 403, 404 from the respective left- and right-side audio capture devices 110a-b, 112a-b. The individual audio signals 401 , 402, 403, 404 may optionally be filtered by a filter 410 before they are provided to the correlation determiner 412. The correlation determiner 412 then determines a correlation between each of the respective individual audio signals 401 , 402, 403, 404 and the difference signal 408. In this example, this corresponds to step 210 of the method 200, as the individual (filtered) signals from the audio capture devices 110a-b, 112a-b are used to determine a correlation. A first correlation result 417 is determined for the first left-side audio capture device 110a, a second correlation result 419 is determined for the second left-side audio capture device 110b, a third correlation result 421 is determined for the first right-side audio capture devices 112a, and a fourth correlation result 423 is determined for the second right-side audio capture devices 112b.
[0089] In some examples, and as illustrated in FIG. 5, the correlation results 417, 419, 421 , 423 are time-averaged. In particular, the correlation results 414, 416 may be provided to a timeaverager 418 that is configured to provide respective time-averaged correlation results 417a, 419a, 421a, 423a. The (time-averaged) correlation results 417, 419, 421 , 423 are then provided to a comparator 420. The comparator 420 may determine an order for the correlation results 417, 419, 421 , 423, i.e. which of the (filtered) individual audio signals 401 , 402, 403, 404 have the highest correlations to the difference signal 408. The comparator 420 may then provide a comparison result 422.
[0090] The comparison result 422 is provided to a selector 424, which is configured to select at least one of the audio capture devices 110a-b, 112a-b for use in determining an output signal 426. This may be performed in the same manner as discussed above in relation to FIG. 4. It is noted that the selector 424, and the output signal 426 have been omitted from FIG. 5.
[0091] As will be appreciated from the description of FIGs. 3 to 5, a number of correlation results are provided, based on which one or more audio capture devices 110a-b, 112a-b are selected for use in determining an output signal. There are a number of possible approaches to how such an output signal may be generated.
[0092] In examples where at least one audio capture device 110, 112 from each side of the wearable audio listening system 102 is selected, an M signal can be determined using the approaches discussed above, i.e. based on a sum operation. This may be achieved using a delay-and-sum beamforming algorithm, either from a single audio capture device 110, 112 from each side (known as binaural beamforming) or a plurality of audio capture devices 110, 112 on each side. When the background noise is distributed substantially evenly around the user II, the M signal will have higher SNR than any of the individual audio capture devices 110, 112. The M signal could then be used as the output signal.
[0093] In the case of the embodiment of FIG. 3, one approach is to compare the correlation results 314, 316, and simply select the audio capture device 110, 112 having the audio signal with the lowest correlation. However, as discussed above, in some examples, both audio capture devices 110, 112 are selected, and the output 326 is determined using a weighted sum of the audio signals 302, 304. In these examples, the correlation results 314, 316 may be used to determine a weight (or gain) for each of the respective audio signals 302, 304.
[0094] One approach to determine a weight is based on a diff / sum operation. The diff / sum operation calculates a ratio between a difference between the correlation results 314, 316 and a sum of the correlation results 314, 316, which returns a value between -1 and 1. For example, if the first audio signal 302 from left-side audio capture device 110 has a correlation result 314 of 0.3 (on a scale of 0 to 1), while the second audio signal 304 from the right-side audio capture device 112 has a correlation result 316 of 0.5, the diff / sum operation returns a value of -0.25. This value can then be mapped to [0,1] (can be non-linear) and time-averaged, and used to define the gains for the microphones. The gains can be determined to provide a total value of 1 based on the mapped value. For example, the diff / sum value of -0.25 can be mapped to a [0,1] value of 0.375. A gain of 0.625 may therefore be applied to the first audio signal 302, while a gain of 0.375 may be applied to the second audio signal 304. This provides an output signal 326 that uses both available audio signals 302, 304, but is weighted based on their correlation to the difference signal 308.
[0095] In the case of the embodiment of FIG. 4, one approach is to compare the correlation results 414, 416, and simply select the side of the wearable audio listening system 102 having the sum 405, 406 with the lowest correlation. In that case, either the left-side audio capture devices 110a-b or the right-side audio capture devices 112a-b) are selected.
[0096] In some examples, only one of the audio capture devices 110a-b, 112a-b is selected. For example, the left side of the wearable audio listening system 102 is determined to have the lowest correlation to the difference signal 308. It may be subsequently identified that the second left-side audio capture device 110b has an SNR that is below a threshold value. In this case, only the first left-side audio capture device 110a may be selected for use in determining the output signal 426.
[0097] In some examples, a number of the audio capture devices 110a-b, 112a-b are selected (for example one from each side, both from one side and one from the other side, or all audio capture devices), and the output 426 is determined using a weighted sum of the audio signals 401 , 402, 403, 404. Take an example where the first sum 405 is determined to have a correlation to the difference signal 408 of 0.2, and the second sum 406 has a correlation of 0.4. A diff / sum operation returns a value of -0.33. This value can be mapped to a [0,1] value of 0.33. In this case, a gain of 0.67 may be applied to the first audio signal 401 and the second audio signal 402, while a gain of 0.33 may be applied to the third audio signal 403 and the fourth audio signal 404. This provides an output signal 426 that uses all available audio signals 401 , 402, 403, 404, but is weighted based on their correlation to the difference signal 408.
[0098] In the case of the embodiment of FIG. 5, one approach is to compare the correlation results 417, 419, 421 , 423, and simply select the audio capture device 110a-b, 112a-b having the audio signal 401 , 402, 403, 404 with the lowest correlation. Alternatively, a number of audio capture devices 110a-b, 112a-b having audio signals 401 , 402, 403, 404 with a correlation below a threshold value may be selected. In some examples, the output 426 is determined using a weighted sum of the audio signals 401 , 402, 403, 404. Take an example where the first audio signal 401 is determined to have a correlation to the difference signal 408 of 0.2, and the second audio signal 402 has a correlation of 0.4. A combined correlation value can then be determined for the left side, for example an average value of 0.3. Say the third audio signal 403 is determined to have a correlation to the difference signal 408 of 0.6, and the fourth audio signal 404 has a correlation of 0.8. A combined correlation value can then be determined for the right side, for example an average value of 0.7. A diff / sum operation on the combined values returns a value of -0.4. This value can be mapped to a to [0,1] value of 0.3. In this case, a gain of 0.7 may be applied to the first audio signal 401 and the second audio signal 402, while a gain of 0.3 may be applied to the third audio signal 403 and the fourth audio signal 404. This provides an output signal 426 that uses all available audio signals 401 , 402, 403, 404, but is weighted based on their correlation to the difference signal 408.
[0099] FIGs. 6A to 60 illustrate top-down view of a user in different listening environments 100 according to a number of examples. In each case, the user II is wearing a wearable audio listening system 102 comprising two left-side audio capture devices 1 10a-b and two right-side audio capture devices 112a-b.
[0100] In FIG. 6A, the external audio signals 116a-d are distributed substantially evenly around the user II. In this case, none of the audio capture devices 110a-b, 112a-b is particularly affected by external noise more than the others. In this case, all of the audio capture devices 110a-b, 1 12a-b may be selected for use in determining an output signal.
[0101] In FIG. 6B, the external audio signals 116c-d are distributed substantially on the right side of the user II. In this case, the right-side audio capture devices 1 12a-b will capture more of the external noise than the left-side audio capture devices 110a-b. In this case, only the left-side audio capture devices 110a-b may be selected for use in determining an output signal.
[0102] In FIG. 6C, the external noise comprises wind noise 1 16e coming from a front-left direction relative to the user II. In this case, the first left-side audio capture device 110a will capture more of the external noise than the other devices. In this case, the first leftside audio capture device 110a may be discounted and the other devices may be selected for use in determining an output signal. In some examples, it may also be determined that the second left-side audio capture device 110b captures a sufficient amount of the wind noise 116e to be discounted, in which case only the right-side audio capture devices 112a-b may be selected.
[0103] FIG. 7 is a block diagram illustrating an exemplary computer system 700 in which examples of the present disclosure may be implemented. In particular, the computer system 700 may, according to some examples, be configured to cause performance of the method 100 of FIG. 1. This example illustrates a computer system 700 such as may be used, in whole, in part, or with various modifications, to provide the functions of the disclosed system. For example, various functions may be controlled by the computer system 700, including, merely by way of example, acquiring, determining, selecting, identifying, receiving, etc. The computer system 700 is shown comprising hardware elements that may be electrically coupled via a bus 790. The hardware elements may include processing circuitry 710, one or more input devices 720 (e.g., a mouse, a keyboard, etc.), and one or more output devices 730 (e.g., a display device, a printer, etc.). The computer system 700 may also include one or more memories 740.
[0104] The processing circuitry 710 may include any number of hardware components for conducting data or signal processing or for executing computer code stored in memories 740. The processing circuitry 710 may include any number of hardware components for conducting data or signal processing or for executing computer code stored in memories 740. The processing circuitry 710 may, for example, include a general-purpose processor, an application specific processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a circuit containing processing components, a group of distributed processing components, a group of distributed computers configured for processing, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processing circuitry 710 may further include computer executable code that controls operation of the programmable device
[0105] The bus 790 provides an interface for system components including, but not limited to, the memories 740 and the processing circuitry 710. The bus 790 may be any of several types of bus structures that may further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and / or a local bus using any of a variety of bus architectures. T
[0106] The memories 740 may be one or more devices for storing data and / or computer code for completing or facilitating methods described herein. The memories 740 may include database components, object code components, script components, or other types of information structure for supporting the various activities herein. The memories 740 may include non- volatile memory (e.g., read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.), and volatile memory (e.g., random-access memory (RAM)), or any other medium which can be used to carry or store desired program code in the form of machine-executable instructions or data structures and which can be accessed by a computer or other machine with processing circuitry 710. The one or more input devices 720 are configured to receive input and selections to be communicated to the computer system 700 when executing instructions. The one or more output devices 730 are configured to forward output, such as to a display, a video display unit.
[0107] The computer system 700 may additionally include a computer-readable storage media reader 750, a communications system 760 (e.g., a modem, a network card (wireless or wired), an infrared communication device, Bluetooth™ device, cellular communication device, etc.), and a working memory 780, which may include RAM and ROM devices as described above. In some embodiments, the computer system 700 may also include a processing acceleration unit 770, which can include a digital signal processor, a special-purpose processor and / or the like.
[0108] The computer-readable storage media reader 750 can further be connected to a non-transitory computer-readable storage medium, together (and, optionally, in combination with the storage devices 740) comprehensively representing remote, local, fixed, and / or removable storage devices plus storage media for temporarily and / or more permanently containing computer- readable information. The communications system 760 may permit data to be exchanged with a network, system, computer and / or other component described above.
[0109] The computer system 700 may also comprise software elements, shown as being currently located within the working memory 780, including an operating system 788 and / or other code 784. Software of the computer system 700 may include code 784 for implementing any or all of the functions of the various elements of the architecture as described herein. For example, software, stored on and / or executed by a computer system such as the system 700, can provide the functions of the disclosed system. All or a portion of the examples disclosed herein may be implemented as a computer program stored on a transitory or non-transitory computer- usable or computer-readable storage medium (e.g., single medium or multiple media) which includes complex programming instructions (e.g., complex computer-readable program code) to cause the processing circuitry to carry out actions described herein. Thus, the computer- readable program code of the computer program can comprise software instructions for implementing the functionality of the examples described herein when executed by the processing circuitry 710. It should be appreciated that alternative embodiments of a computer system 700 may have numerous variations from that described above. For example, customised hardware might also be used and / or particular elements might be implemented in hardware, software (including portable software, such as applets), or both. The computer system 700 may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet. While only a single device is illustrated, the computer system 700 may include any collection of devices that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. Furthermore, connection to other computing devices such as network input / output and data acquisition devices may also occur.
[0110] The operational actions described in any of the exemplary aspects herein are described to provide examples and discussion. The actions may be performed by hardware components, may be embodied in machine-executable instructions to cause a processor to perform the actions, or may be performed by a combination of hardware and software. Although a specific order of method actions may be shown or described, the order of the actions may differ. In addition, two or more actions may be performed concurrently or with partial concurrence.
[0111] According to certain examples, there is also disclosed:
[0112] Example 1. A computer-implemented method (200) for selecting at least one audio capture device (110, 112) for use in determining an output signal (326, 426) from a wearable audio listening system (102), the computer-implemented method (200) comprising, by processing circuitry (710) of a computer system (108, 700): acquiring (202) a first signal (302) representing audio captured by a leftside audio capture device (110) of the wearable audio listening system (102); acquiring (204) a second signal (304) representing audio captured by a right-side audio capture device (112) of the wearable audio listening system (102); determining (206) a third signal (308) based on a difference between the first signal (302) and the second signal (304); determining (210) a first correlation (314) between the third signal (308) and the first signal (302); determining (210) a second correlation (316) between the third signal (308) and the second signal (304); and selecting (212) at least one audio capture device (110, 112) for use in determining an output signal (326) based on the determined correlations (314, 316).
[0113] Example 2. A computer-implemented method (200) for selecting at least one audio capture device (110, 112) for use in determining an output signal (326, 426) from a wearable audio listening system (102), the computer-implemented method (200) comprising, by processing circuitry (710) of a computer system (108, 700): acquiring (202) a first sum (405) of a plurality of first signals (401 , 402) representing audio captured by a plurality of respective left-side audio capture devices (110) of the wearable audio listening system (102); acquiring (204) a second sum (406) of a plurality of second signals (403, 404) representing audio captured by one or more respective right-side audio capture devices (112) of the wearable audio listening system (102); determining (206) a third signal (408) based on a difference between the first sum (405) and the second sum (406); determining (208) a first correlation (414) between the third signal (408) and the first sum (405); determining (208) a second correlation (416) between the third signal (408) and the second sum (406); and selecting (212) at least one audio capture device (110, 112) for use in determining an output signal (426) based on the determined correlations (414, 416).
[0114] Example 3. A computer-implemented method (200) for selecting at least one audio capture device (110, 112) for use in determining an output signal (326, 426) from a wearable audio listening system (102), the computer-implemented method (200) comprising, by processing circuitry (710) of a computer system (108, 700): acquiring (202) a first sum (405) of a plurality of first signals (401 , 402) representing audio captured by a plurality of respective left-side audio capture devices (110) of the wearable audio listening system (102); acquiring (204) a second sum (406) of a plurality of second signals (403, 404) representing audio captured by one or more respective right-side audio capture devices (112) of the wearable audio listening system (102); determining (206) a third signal (408) based on a difference between the first sum (405) and the second sum (406); determining (210) a respective first correlation (417, 419) between the third signal (408) and each of the plurality of first signals (401 , 402); determining (210) a respective second correlation (421 , 423) between the third signal (408) and each of the plurality of second signals (403, 404); and selecting (212) at least one audio capture device (1 10, 112) for use in determining an output signal (426) based on the determined correlations (417, 419, 421 , 423).
[0115] Terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including" when used herein specify the presence of stated features, integers, actions, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, actions, steps, operations, elements, components, and / or groups thereof.
[0116] It will be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the present disclosure.
[0117] Relative terms such as "below" or "above" or "upper" or "lower" or "horizontal" or "vertical" may be used herein to describe a relationship of one element to another element as illustrated in the Figures. It will be understood that these terms and those discussed above are intended to encompass different orientations of the device in addition to the orientation depicted in the Figures. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0118] It is to be understood that the present disclosure is not limited to the aspects described above and illustrated in the drawings; rather, the skilled person will recognize that many changes and modifications may be made within the scope of the present disclosure and appended claims.
[0119] In the drawings and specification, there have been disclosed aspects for purposes of illustration only and not for purposes of limitation, the scope of the disclosure being set forth in the following claims.
Claims
Claims1. A computer-implemented method (200) for selecting at least one audio capture device (110, 112) for use in determining an output signal (326, 426) from a wearable audio listening system (102), the computer-implemented method (200) comprising, by processing circuitry (710) of a computer system (108, 700): acquiring (202) a first sum (302, 405) of one or more first signals (302, 401 , 402) representing audio captured by one or more respective left-side audio capture devices (110) of the wearable audio listening system (102); acquiring (204) a second sum (304, 406) of one or more second signals (302, 403, 404) representing audio captured by one or more respective rightside audio capture devices (112) of the wearable audio listening system (102); determining (206) a third signal (308, 408) based on a difference between the first sum (302, 405) and the second sum (304, 406); determining (208, 210) a correlation (314, 316, 414, 416, 417, 419, 421 , 423) between the third signal (308, 408) and each of the first and second sums (302, 405, 304, 406) and / or between the third signal (308, 408) and each of the first and second signals (302, 401 , 402, 302, 403, 404); and selecting (212) at least one audio capture device (110, 112) for use in determining an output signal (326, 426) based on the determined correlations (314, 316, 414, 416, 417, 419, 421 , 423).
2. The computer-implemented method (200) of claim 1 , wherein: acquiring (202) the first sum (302) comprises acquiring a single first signal (302); acquiring (204) the second sum (304) comprises acquiring a single second signal (304); the computer-implemented method (200) comprising: determining (206) the third signal (308) based on a difference between the first and second signals (302, 304); determining (210) a first correlation (314) between the third signal (308) and the first signal (302); and determining (210) a second correlation (316) between the third signal (308) and the second signal (204).
3. The computer-implemented method (200) of claim 2, further comprising: filtering the first signal (302) and the second signal (304);the computer-implemented method (200) comprising: determining (206) the third signal (308) based on a difference between filtered first and second signals (302a, 304a); determining (210) a first correlation (314) between the third signal (308) and the filtered first signal (302a); and determining (210) a second correlation (316) between the third signal (308) and the filtered second signal (304a).
4. The computer-implemented method (200) of claim 1 , wherein: acquiring (202) the first sum (405) comprises acquiring a sum of a plurality of first signals (401 , 402); and acquiring (204) the second sum (406) comprises acquiring a sum of a plurality of second signals (403, 404).
5. The computer-implemented method (200) of claim 4, further comprising: time synchronising the plurality of first signals (401 , 402) and acquiring (202) a sum (405) of the plurality of time-synchronised first signals (401a, 402); and / or time synchronising the plurality of second signals (403, 404) and acquiring (204) a sum (406) of the plurality of time-synchronised second signals (403a, 404).
6. The computer-implemented method (200) of claim 4 or 5, comprising: determining (210) a first correlation (414) between the third signal (408) and the first sum (405); and determining (210) a second correlation (416) between the third signal (408) and the second sum (406).
7. The computer-implemented method (200) of claim 6, further comprising: filtering the first sum (405) and the second sum (406); the computer-implemented method (200) comprising: determining (206) the third signal (408) based on a difference between filtered first and second sums (405a, 406a); determining (210) a first correlation (414) between the third signal (408) and the filtered first sum (405a); and determining (210) a second correlation (416) between the third signal (408) and the filtered second sum (406a).
8. The computer-implemented method (200) of claim 4 or 5, comprising:determining (210) a respective correlation (417, 419) between the third signal (408) and each of the plurality of first signals (401 , 402); and determining (210) a respective correlation (421 , 423) between the third signal (408) and each of the plurality of second signals (403, 404).
9. The computer-implemented method (200) of claim 8, further comprising: filtering the first sum (405) and the second sum (406); and filtering each of the plurality of first and second signals (401 , 402, 403, 404); the computer-implemented method (200) comprising: determining (206) the third signal (408) based on a difference between filtered first and second sums (405a, 406a); determining (210) a respective correlation (417, 419) between the third signal (408) and each of the plurality of filtered first signals; and determining (210) a respective correlation (421 , 423) between the third signal (408) and each of the plurality of filtered second signals.
10. The computer-implemented method (200) of claim 8 or 9, further comprising: time synchronising the plurality of first signals (401 , 402) and determining (210) a respective correlation (417, 419) between the third signal (408) and each of the plurality of time-synchronised first signals; and / or time synchronising the plurality of second signals (403, 404) and determining (210) a respective correlation (421 , 423) between the third signal (408) and each of the plurality of time-synchronised second signals.
11. The computer-implemented method (200) of any preceding claim, wherein determining (208, 210) a correlation comprises determining a cross-correlation.
12. The computer-implemented method (200) of any preceding claim, further comprising determining a time average (314a, 316a, 414a, 416a, 417a, 419a, 421a, 423a) of a result (314, 316, 414, 416, 417, 419, 421 , 423) of the correlation.
13. The computer-implemented method (200) of any preceding claim, wherein selecting (212) at least one audio capture device (110, 112) for use in determining an output signal (326, 426) based on the determined correlations (314, 316, 414, 416, 417, 419, 421 , 423) comprises:acquiring a first correlation result (314, 414, 417, 419) associated with the one or more left-side audio capture devices (110); acquiring a second correlation result (316, 416, 421 , 423) associated with the one or more right-side audio capture devices (112); comparing the first and second correlation results (314, 316, 414, 416, 417, 419, 421 , 423); and selecting (212) the one or more left-side audio capture devices (110) or the one or more right-side audio capture devices (112) for use in determining the output signal (326, 426) based on determining that one of the first correlation result (314, 414, 417, 419) and the second correlation result (316, 416, 421 , 423) is higher than the other.
14. The computer-implemented method (200) of any of claims 1 to 12, wherein selecting (212) at least one audio capture device (110, 112) for use in determining an output signal (326, 426) based on the determined correlations (314, 316, 414, 416, 417, 419, 421 , 423): acquiring a first correlation result (314, 414, 417, 419) associated with the one or more left-side audio capture devices (110); acquiring a second correlation result (316, 416, 421 , 423) associated with the one or more right-side audio capture devices (112); determining a difference between the first and second correlation results (314, 316, 414, 416, 417, 419, 421 , 423); determining a sum of the first and second correlation results (314, 316, 414, 416, 417, 419, 421 , 423); determining a ratio between the difference and the sum of the first and second correlation results (314, 316, 414, 416, 417, 419, 421 , 423); and determining, based on the determined ratio, a gain to be applied to each of the one or more left-side audio capture devices (110) and / or each of the one or more right-side audio capture devices (112) for use in determining the output signal (326, 426).
15. The computer-implemented method (200) of claim 14, wherein a sum of the respective gains of the one or more left-side audio capture devices (110) and the one or more right-side audio capture devices (112) is equal to one.
16. The computer-implemented method (200) of any preceding claim, wherein selecting (212) at least one audio capture device (110, 112) for use indetermining an output signal (326, 426) based on the determined correlations (314, 316, 414, 416, 417, 419, 421 , 423) comprises selecting an audio capture device (110, 112) having a threshold signal to noise ratio.
17. The computer-implemented method (200) of any preceding claim, wherein the wearable audio listening system (102) comprises integrated left-side and rightside audio listening devices (104, 106).
18. The computer-implemented method (200) of any of claims 1 to 16, wherein the wearable audio listening system (102) comprises a left-side wearable audio listening device (104) and a right-side wearable audio listening device (106) separate to the left-side wearable audio listening device (104).
19. The computer-implemented method (200) of any preceding claim, wherein the wearable audio listening system (102) comprises one or more sensors configured to determine if the one or more left-side and / or right-side audio capture devices (110, 112) are in a worn position.
20. A computer system (108, 700) for selecting at least one audio capture device (110, 112) for use in determining an output signal (326, 426) from a wearable audio listening system (102), the computer system (108, 700) comprising processing circuitry (710) configured to: acquire a first sum (302, 405) of one or more first signals (302, 401 , 402) representing audio captured by one or more respective left-side audio capture devices (110) of the wearable audio listening system (102); acquire a second sum (304, 406) of one or more second signals (302, 403, 404) representing audio captured by one or more respective right-side audio capture devices (112) of the wearable audio listening system (102); determine a third signal (308, 408) based on a difference between the first sum (302, 405) and the second sum (304, 406); determine a correlation (314, 316, 414, 416, 417, 419, 421 , 423) between the third signal (308, 408) and each of the first and second sums (302, 405, 304, 406) and / or between the third signal (308, 408) and each of the first and second signals (302, 401 , 402, 302, 403, 404); and select at least one audio capture device (110, 112) for use in determining an output signal (326, 426) based on the determined correlations (314, 316, 414, 416, 417, 419, 421 , 423).
21. A wearable audio listening system (102) comprising the computer system (108, 700) of claim 20.
22. A computer program product comprising program code for performing, when executed by processing circuitry (710), the computer-implemented method (200) of any of claims 1 to 19.
23. A non-transitory computer-readable storage medium comprising instructions, which when executed by processing circuitry (710), cause the processing circuitry (710) to perform the computer-implemented method (200) of any of claims 1 to 19.
Citation Information
Patent Citations
Motor-vehicle voice-control system and microphone-selecting method therefor
US20120221341A1
Voice Sensing using Multiple Microphones
US20180047381A1
Automatic noise cancellation using multiple microphones
US20180114518A1