Apparatus, method and computer program for controlling noise reduction
By dividing the audio signal into multiple intervals and controlling the noise reduction method according to the noise characteristic parameters, the noise reduction problem in the capture audio signal in multiple microphones in the prior art is solved, and high-quality audio signal processing is achieved.
Patent Information
- Application Number
- CN202510187334.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-20
- Filing Date
- 2019-12-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively control noise reduction in audio signals captured by multiple microphones, especially with different noise characteristics in different frequency intervals.
By dividing the audio signal into multiple intervals, parameters related to the noise characteristics of different intervals are determined, and noise reduction methods for different intervals are controlled based on these parameters.
Personalized noise reduction processing for different frequency intervals is realized, improving the quality of the audio signal and the perceptibility of user perception.
Smart Images

Figure CN119993183A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese invention patent application entitled "Device, method and computer program for controlling noise reduction" (application number 201980092414.2, application date December 13, 2019). Technical Field
[0002] Examples of the present disclosure relate to apparatus, methods, and computer programs for controlling noise reduction.Some relate to apparatus, methods, and computer programs for controlling noise reduction in an audio signal including audio captured by multiple microphones. Background Art
[0003] Audio signals comprising audio captured by multiple microphones may be used to provide spatial audio signals to a user. The quality of these signals may be adversely affected by unwanted noise captured by the multiple microphones. Summary of the invention
[0004] According to various but not necessarily all examples of the present disclosure, a device is provided, which includes a module for: obtaining one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; dividing the obtained one or more audio signals into multiple intervals; determining one or more parameters related to one or more noise characteristics of different intervals; and controlling noise reduction applied to different intervals based on the one or more parameters determined in different intervals.
[0005] The intervals may comprise time-frequency intervals.
[0006] The noise characteristic may include a noise level.
[0007] Parameters related to one or more noise characteristics may be determined independently for different intervals.
[0008] Determining one or more parameters associated with the one or more noise characteristics may include determining whether the one or more parameters are within a threshold range.
[0009] Different thresholds for one or more parameters related to noise characteristics may be used for different frequency ranges within the plurality of intervals.
[0010] One or more parameters related to one or more noise characteristics may include one or more of the following: the noise level within the interval, the noise level in the interval before the analyzed interval, the noise reduction method used for the previous frequency interval, the duration that the current noise reduction method has been used within the frequency band, and the orientation of a microphone capturing one or more audio signals.
[0011] The noise reduction applied to the first interval may be independent of the noise reduction applied to the second interval, where the first and second intervals have different frequencies but overlap in time.
[0012] Different noise reduction can be applied to different intervals, where the different intervals have different frequencies but overlap in time.
[0013] Controlling the noise reduction applied to the interval may include selecting a method for noise reduction within the interval.
[0014] Controlling the noise reduction applied to the intervals may include determining when to switch between different methods of noise reduction for one or more intervals.
[0015] Controlling the noise reduction applied to the interval may include one or more of: providing a spatial output with noise reduction, providing a spatial output without noise reduction, providing a mono audio output with noise reduction, providing a beamformed output, providing a beamformed output with noise reduction.
[0016] The noise that is reduced may include noise within the one or more audio signals that has been detected by one or more microphones of the plurality of microphones capturing the audio.
[0017] The noise may include one or more of wind noise, touch noise.
[0018] According to various but not necessarily all examples of the present disclosure, there is provided a device comprising: a processing circuit; and a memory circuit comprising a computer program code, wherein the memory circuit and the computer program code are configured to, together with the processing circuit, enable the device to: obtain one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; divide the obtained one or more audio signals into multiple intervals; determine one or more parameters associated with one or more noise characteristics of different intervals; and control noise reduction applied to different intervals based on the one or more parameters determined within the different intervals.
[0019] According to various but not necessarily all examples of the present disclosure, there is provided an electronic device comprising the apparatus as described above and a plurality of microphones.
[0020] The electronic device may include a communication device.
[0021] According to various but not necessarily all examples of the present disclosure, a method is provided, comprising: obtaining one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; dividing the obtained one or more audio signals into multiple intervals; determining one or more parameters related to one or more noise characteristics of different intervals; and controlling noise reduction applied to different intervals based on the one or more parameters determined within the different intervals.
[0022] Parameters related to one or more noise characteristics may be determined independently for different intervals.
[0023] According to various but not necessarily all examples of the present disclosure, there is provided a computer program comprising computer program instructions, which, when executed by a processing circuit, causes: obtaining one or more audio signals, wherein the one or more audio signals comprise audio captured by a plurality of microphones; dividing the obtained one or more audio signals into a plurality of intervals; determining one or more parameters associated with one or more noise characteristics of different intervals; and controlling noise reduction applied to different intervals based on the one or more parameters determined within the different intervals.
[0024] Parameters related to one or more noise characteristics may be determined independently for different intervals.
[0025] According to various but not necessarily all examples of the present disclosure, a device is provided, including a module for: obtaining one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; dividing the obtained one or more audio signals into multiple intervals; determining one or more parameters related to one or more noise characteristics of different intervals; and determining whether to provide mono audio output or spatial audio output based on the determined one or more parameters.
[0026] The intervals may comprise time-frequency intervals.
[0027] The noise characteristic may include a noise level.
[0028] Providing a mono audio output may include determining a microphone signal having minimal noise and providing the mono audio output using the determined microphone signal.
[0029] Providing a mono audio output may include combining microphone signals from two or more microphones of the plurality of microphones, wherein the two or more microphones of the plurality of microphones are located proximate to each other.
[0030] The spatial audio output may include one or more of the following: a stereo signal, a binaural signal, an Atmos signal.
[0031] Determining one or more parameters associated with the one or more noise characteristics for different intervals may include determining whether energy differences between microphone signals from different microphones within the plurality of microphones are within a threshold range.
[0032] Determining one or more parameters associated with one or more noise characteristics for different intervals may include determining whether switching between mono audio output and spatial audio output has occurred within a threshold time.
[0033] Different threshold ranges may be used for different frequency bands.
[0034] A mono audio output may be provided for a first frequency band within the interval, and a spatial audio output may be provided for a second frequency band within the interval, wherein the first interval and the second interval have different frequencies but overlap in time.
[0035] According to various but not necessarily all examples of the present disclosure, there is provided a device comprising: a processing circuit; and a memory circuit comprising a computer program code, wherein the memory circuit and the computer program code are configured to, together with the processing circuit, enable the device to: obtain one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; divide the obtained one or more audio signals into multiple intervals; determine one or more parameters related to one or more noise characteristics of different intervals; and determine whether to provide a mono audio output or a spatial audio output based on the determined one or more parameters.
[0036] According to various but not necessarily all examples of the present disclosure, there is provided an electronic device comprising the above apparatus and a plurality of microphones.
[0037] The electronic device may include a communication device.
[0038] According to various but not necessarily all examples of the present disclosure, a method is provided, comprising: obtaining one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; dividing the obtained one or more audio signals into multiple intervals; determining one or more parameters related to one or more noise characteristics of different intervals; and controlling noise reduction applied to different intervals based on the one or more parameters determined within the different intervals.
[0039] Providing a mono audio output may include determining a microphone signal having minimal noise, and providing the mono audio output using the determined microphone signal.
[0040] According to various but not necessarily all examples of the present disclosure, there is provided a computer program comprising computer program instructions which, when executed by a processing circuit, result in: obtaining one or more audio signals, wherein the one or more audio signals comprise audio captured by a plurality of microphones; dividing the obtained one or more audio signals into a plurality of intervals; determining one or more parameters associated with one or more noise characteristics of different intervals; and controlling noise reduction applied to different intervals based on the one or more parameters determined within the different intervals.
[0041] Providing the mono audio output may include determining a microphone signal having minimal noise, and providing the mono audio output using the determined microphone signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Some example embodiments will now be described with reference to the accompanying drawings, in which:
[0043] Figure 1 An example device is described;
[0044] Figure 2 An example electronic device is illustrated
[0045] Figure 3 An example method is illustrated;
[0046] Figure 4 Another example method is illustrated;
[0047] Figure 5 Another example method is illustrated;
[0048] Figure 6 Another example electronic device is described; and
[0049] Figure 7 Another example method is illustrated. DETAILED DESCRIPTION
[0050] Examples of the present disclosure relate to an apparatus 101, method, and computer program for controlling noise reduction in an audio signal comprising audio captured by multiple microphones. The apparatus 101 comprises an apparatus for obtaining 301 one or more audio signals, wherein the one or more audio signals comprise audio captured by multiple microphones 203 and dividing 303 the obtained one or more audio signals into a plurality of intervals. The apparatus may also be configured to determine 305 one or more parameters associated with one or more noise characteristics of different intervals, and to control 307 noise reduction applied to different intervals based on the one or more parameters determined in the different intervals.
[0051] Thus, the apparatus 101 may enable different noise reduction methods to be applied to different intervals within the obtained audio signal. This may take into account the perceptibility differences of noise in different frequency bands, the perceptibility of switching between different noise reduction methods in different frequency bands, and any other suitable factors to improve the perceptible quality of the output signal.
[0052] Figure 1 The device 101 according to an example of the present disclosure is schematically shown. Figure 1 In the example of , the device 101 includes a controller 103. Figure 1 In some examples, the controller 103 may be implemented as a controller circuit. In some examples, the controller 103 may be implemented solely in hardware, with certain aspects in software (including separate firmware), or may be a combination of hardware and software (including firmware).
[0053] like Figure 1 As shown, the controller 103 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 109 in a general or special purpose processor 105, which executable instructions can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such a processor 105.
[0054] The processor 105 is configured to read and write to the memory 107. The processor 105 may also include an output interface and an input interface. The processor 105 outputs data and / or commands through the output interface, and the data and / or commands are input to the processor 105 through the input interface.
[0055] The memory 107 is configured to store a computer program 109 comprising computer program instructions (computer program code 111) that control the operation of the apparatus 101 when loaded into the processor 105. The computer program instructions of the computer program 109 provide for enabling the apparatus 101 to perform Figure 3 , 4 , 5 and 7. The processor 105 can load and execute the computer program 109 by reading the memory 107.
[0056] Therefore, the device 101 includes: at least one processor 105; and at least one memory 107 including computer program code 111, and the at least one memory 107 and the computer program code 111 are configured to use the at least one processor 105 to cause the device 101 to at least perform: obtain 301 one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones; divide 303 the obtained one or more audio signals into multiple intervals; determine 305 one or more parameters related to one or more noise characteristics of different intervals; and control 307 noise reduction applied to different intervals based on the one or more parameters determined in different intervals.
[0057] In some examples, device 101 may include at least one processor 105; and at least one memory 107 including computer program code 111, wherein the at least one memory 107 and the computer program code 111 are configured to utilize at least one processor 105 to cause device 101 to at least perform: obtaining 501 one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones 203; dividing 503 the obtained one or more audio signals into multiple intervals; determining 505 one or more parameters associated with one or more noise characteristics of different intervals; and determining 507 whether to provide mono audio output or spatial audio output based on the determined one or more parameters.
[0058] like Figure 1 As shown, the computer program 109 can reach the device 101 through any suitable delivery mechanism 113. The delivery mechanism 113 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a storage device, a recording medium such as a compact disk read-only memory (CD-ROM) or a digital versatile disk (DVD) or a solid-state memory, an article of manufacture containing or tangibly embodying the computer program 109. The delivery mechanism can be a signal configured to reliably transmit the computer program 109. The device 101 can propagate or send the computer program 109 as a computer data signal. In some examples, the computer program 109 can be sent to the device 101 using a wireless protocol, such as Bluetooth, Bluetooth low energy, Bluetooth smart, 6LoWPan (IPv6 over low power personal area network), ZigBee, ANT+, near field communication (NFC), radio frequency identification, wireless local area network (WLAN), or any other suitable protocol.
[0059] The computer program 109 includes computer program instructions for causing the device 101 to perform at least the following operations: obtain 301 one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones 203; divide 303 the obtained one or more audio signals into multiple intervals; determine 305 one or more parameters related to one or more noise characteristics of different intervals; and control 307 noise reduction applied to different intervals based on the one or more parameters determined in different intervals.
[0060] In some examples, the computer program 109 includes computer program instructions for causing the device 101 to perform at least the following operations: obtain 501 one or more audio signals, wherein the one or more audio signals include audio captured by multiple microphones 203; divide 503 the obtained one or more audio signals into multiple intervals; determine 505 one or more parameters associated with one or more noise characteristics of different intervals; and determine 507 whether to provide mono audio output or spatial audio output based on the determined one or more parameters.
[0061] The computer program instructions may be included in a computer program 109 , a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program 109 .
[0062] Although memory 107 is shown as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage.
[0063] Although processor 105 is shown as a single component / circuit, it may be implemented as one or more separate components / circuit, some or all of which may be integrated / removable.Processor 105 may be a single-core or multi-core processor.
[0064] References to "computer-readable storage medium," "computer program product," "tangibly embodied computer program," etc., or "controller," "computer," "processor," etc., should be understood to include not only computers having different structures, such as single / multi-processor structures and serial (von Neumann) / parallel architectures, but also special-purpose circuits, such as field-programmable gate arrays (FPGAs), application-specific circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc. should be understood to include software or firmware for programmable processors, such as the programmable content of hardware devices, whether instructions for a processor, or configuration settings for fixed-function devices, gate arrays, or programmable logic devices, etc.
[0065] In this application, the term "circuitry" may refer to one or more or all of the following:
[0066] (a) pure hardware circuit implementation (e.g., implementation only in analog and / or digital circuits) and
[0067] (b) a combination of hardware circuitry and software, such as (where applicable):
[0068] (i) a combination of analog and / or digital hardware circuitry and software / firmware, and
[0069] (ii) any portion of a hardware processor (including a digital signal processor) with software, software and memory that work together to enable a device (such as a mobile phone or server) to perform various functions, and
[0070] (c) Hardware circuits and / or processors, such as a microprocessor or portion of a microprocessor, that require software (eg, firmware) to operate, but where software is not required to operate, the software may not be present.
[0071] The definition of circuitry applies to all uses of the term in this application, including in any claims. As another example, as used in this application, the term circuitry also covers an implementation of hardware-only circuits or processors and their (or their) accompanying software and / or firmware. If applicable to the particular claim element telephone, the term circuitry also covers, for example, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0072] Figure 2 An example electronic device 201 is shown. The example electronic device 201 includes Figure 1 The apparatus 101 shown. The apparatus 101 may include a processor 105 and a memory 107 as described above. The example electronic device also includes a plurality of microphones 203.
[0073] The electronic device 201 may be a communication device such as a mobile phone. It should be understood that the communication device may include Figure 2 Components not shown in the drawings, for example, a communication device may include one or more transceivers that enable wireless communication.
[0074] In some examples, the electronic device 201 may be an image capture device. In such examples, the electronic device 201 may include one or more cameras that may enable the capture of images. The images may be video images, still images, or any other suitable type of images. The images captured by the camera module may be accompanied by audio captured by the plurality of microphones 203.
[0075] The plurality of microphones 203 may include any device configured to capture sound and enable providing an audio signal. The audio signal may include an electrical signal representing at least some of the sound fields captured by the plurality of microphones 203. The output signal provided by the microphone 203 may be modified to provide an audio signal. For example, the output signal from the microphone 203 may be filtered or equalized, or any other suitable processing may be performed thereon.
[0076] The electronic device 201 is configured to provide an audio signal including audio from a plurality of microphones 203 to the apparatus 101. This enables the apparatus 101 to process the audio signal. In some examples, it may enable the apparatus 101 to process the audio signal in order to reduce the effects of noise captured by the microphones 203.
[0077] Multiple microphones 203 may be located within the electronic device 201 to enable the capture of spatial audio. For example, the locations of the multiple microphones 203 may be throughout the electronic device 201 to enable the capture of spatial audio. Spatial audio includes one or more audio signals that may be presented so that a user of the electronic device 201 may perceive spatial properties of the one or more audio signals. For example, spatial audio may be presented so that the user may perceive the source direction and distance from the audio source.
[0078] exist Figure 2 In the example shown, the electronic device 201 includes three microphones 203. The first microphone 203A is arranged at the first end on the first surface of the electronic device 201. The second microphone 203B is arranged at the first end on the second surface of the electronic device 201. The second surface is located on the side of the electronic device 201 opposite to the first surface. The third microphone 203C is arranged at the second end of the electronic device 201. The second end is the end of the electronic device 201 opposite to the first end. The third microphone 203C is arranged on the same surface as the first microphone 203A. It should be understood that other configurations of multiple microphones 203 may be provided in other examples of the present disclosure. Similarly, in other examples, the electronic device 201 may include different numbers of microphones 203. For example, the electronic device 201 may include two microphones 203 or may include more than three microphones 203.
[0079] Multiple microphones 203 are coupled to the device 101. This may enable signals captured by the multiple microphones 203 to be provided to the device 101. This may enable an audio signal including the audio captured by the microphones 203 to be stored in the memory 107. This may also enable the processor 105 to perform noise reduction on the obtained audio signal. Example methods of noise reduction are as follows: Figure 3 and Figure 4 shown.
[0080] exist Figure 2 In the example shown, the microphone 203 that captures audio and the processor 105 that performs noise reduction are provided in the same electronic device 201. In other examples, the microphone 203 and the processor 105 that performs noise reduction may be provided in different electronic devices 201. For example, the audio signals may be transmitted from the plurality of microphones 203 to the processing device via a wireless connection or some other suitable communication link.
[0081] Figure 3 An example method of controlling noise reduction is shown. The method may use Figure 1 The device 101 shown and / or Figure 2 The electronic device 201 shown is implemented.
[0082] The method includes, at block 301, obtaining one or more audio signals, wherein the one or more audio signals represent sound signals captured by a plurality of microphones 203. In some examples, the one or more audio signals include audio obtained from microphones 203, which are provided within the same electronic device 201 as the apparatus 101. In other examples, the one or more audio signals may include audio obtained from microphones 203 provided in one or more separate devices. In such examples, the audio signals may be sent to the apparatus 101.
[0083] The one or more audio signals obtained may include electrical signals representing at least some of the sound fields captured by the plurality of microphones 203. The output signals provided by the microphones 203 may be modified to provide audio signals. For example, the output signals from the microphones 203 may be filtered or equalized or any other suitable processing may be performed thereon.
[0084] The one or more audio signals obtained may include audio captured by the spatially distributed microphones 203, so that a spatial audio signal may be provided to the user. The spatial audio signal may be a stereo signal, a binaural signal, an panoramic signal, or any other suitable type of spatial audio signal.
[0085] The method further includes, at block 303, dividing the obtained one or more audio signals into a plurality of intervals. The obtained one or more audio signals may be divided into intervals using any suitable process. The intervals may be time-frequency intervals, time intervals, or any other suitable type of intervals.
[0086] In some examples, the intervals may be of different sizes. For example, where the intervals include time-frequency intervals, the frequency bands used to define the time-frequency intervals may have different sizes for different frequencies. For example, a lower frequency interval may cover a smaller frequency band than a higher frequency interval.
[0087] At block 305, the method includes determining one or more parameters associated with one or more noise characteristics of different intervals. In some examples, the parameters may be determined for each interval. In other examples, the parameters may be determined only for a subset of the intervals.
[0088] In some examples, one or more parameters associated with one or more noise characteristics may provide an indication of whether noise is present in different intervals. In other examples, the method may include determining whether noise is present and then, if noise is present, determining one or more parameters associated with one or more noise characteristics of the determined noise for different intervals.
[0089] In some examples, one or more parameters associated with one or more noise characteristics may be determined while determining the presence of noise. In other examples, the presence of noise may be determined separately from one or more parameters associated with one or more noise characteristics.
[0090] In some examples, one or more parameters associated with one or more noise characteristics may be a noise presence parameter, which may be a binary variable having a value equivalent to noise or no noise. The no noise value may be a level of the only noise that the user does not perceive to be present. In other examples, noise presence may have a range of values. In some examples, the noise presence variable value may be associated with signal energy. One or more parameters associated with noise characteristics may provide a ratio or energy value indicating the amount of external sound in the audio signal captured at different intervals, in which case it may be assumed that the remaining energy is noise.
[0091] The noise characteristics being analyzed are related to the noise detected by one or more of the multiple microphones 203 that capture the audio of the audio signal. The noise may be an unwanted sound in the audio signal captured by the microphone 203. The noise may include noise that does not correspond to the sound field captured by the multiple microphones 203. For example, the noise may be wind noise, operating noise, or any other suitable type of noise. In some examples, the noise may include noise caused by other components of the electronic device 201. For example, the noise may include noise caused by a focusing camera within the electronic device 201. The noise characteristics being analyzed may exclude noise introduced by the microphone 203.
[0092] In some examples, the one or more parameters related to noise characteristics may include an energy ratio parameter that determines a proportion of external sounds at the captured audio signal, which may include external sounds and noise.
[0093] The one or more parameters related to the one or more noise characteristics may include any parameters that provide an indication of noise level and / or noise reduction methods that will improve the audio quality of the interval being analyzed.
[0094] In some examples, one or more parameters related to noise characteristics may include noise level in an interval. The noise level may be determined by monitoring signal level differences between frequency bands, monitoring correlations between audio captured by different microphones 203, or any other suitable method.
[0095] In some examples, the noise level in intervals prior to the interval being analyzed can be monitored. For example, to determine the noise level in a given frequency band, the noise in the previous time period can be determined. The probability of a significant change in the noise level in the next interval can then be predicted based on the noise level in the previous interval. This can therefore take into account the fact that a single interval may show a small amount of noise, but this may be an anomaly in other noisy time periods.
[0096] In some examples, the one or more parameters related to the noise characteristics may include parameters related to the noise reduction method currently being used or previously used. In such examples, the one or more parameters may include the noise reduction method used for the previous time interval in the frequency band, the duration that the current noise reduction method has been used, or any other suitable parameters.
[0097] Using the parameters associated with the noise reduction method, it is possible to implement how often switching between different types of noise reduction methods occurs. This can reduce artifacts caused by switching between different types of noise reduction and therefore can improve the audio quality perceived by the user.
[0098] Other types of parameters related to noise characteristics may also be used in other examples of the present disclosure. For example, in some examples, the orientation of the microphone 203 capturing the audio or any other suitable parameter may be used. The orientation of the microphone may give an indication of effects such as occlusion, which may affect the level of audio captured by the microphone from different directions and therefore affect the detection of noise captured by the microphone.
[0099] Parameters related to noise characteristics may be determined independently for different intervals. For example, the analysis performed for a first interval may be independent of the analysis performed for a second interval. This may mean that the analysis and determination performed for the first interval does not affect the analysis and determination performed for the second interval.
[0100] In some examples, determining one or more parameters related to the noise characteristic includes determining whether the one or more parameters are within a threshold range. Determining whether a parameter is within a threshold range may include determining whether a value of the parameter is above or below a threshold. In some examples, determining whether a parameter is within a threshold range may include determining whether the value of the parameter is between an upper limit value and a lower limit value.
[0101] For different intervals, the value of the threshold value may be different. For example, different threshold values of one or more parameters related to noise characteristics can be used for different frequency ranges within multiple time-frequency intervals. This can take into account the fact that different frequency bands may be more affected by noise than other frequency bands. For example, wind noise may be more easily perceived in lower frequency bands than in higher frequency bands. In addition, due to the presence of a higher phase difference, the switching between different noise reduction methods may be more easily perceived by the user at a higher frequency band. Since the acoustic shielding effect of the electronic device 201 is larger for a higher frequency band, the level difference may also be higher at a higher frequency band. This may make it undesirable to switch between different noise reduction methods too frequently for a higher frequency band. Therefore, in an example of the present disclosure, different threshold values for the time period between switching can be used for different frequency bands.
[0102] At block 307 , the method includes controlling noise reduction applied to the different intervals based on one or more parameters determined in the different time-frequency intervals.
[0103] Controlling the noise reduction applied to the interval may include using the determined parameter to select a noise reduction method to be applied to the interval. The selection of the noise reduction method may be based on whether the parameter related to the noise characteristic is determined to be within a threshold range.
[0104] The noise reduction method may include any process that reduces the amount of noise within the interval. In some examples, the noise reduction method may include one or more of the following: providing a spatial output with noise reduction, providing a spatial output without noise reduction, providing a mono audio output with noise reduction, providing a beamformed output, providing a beamformed output with noise reduction. The type of noise reduction available may depend on the type of spatial audio available, the type of microphone 203 used to capture the audio, the noise level, and any other suitable factors.
[0105] In an example of the present disclosure, parameters related to noise characteristics are determined differently for different intervals. This can enable different noise reduction methods to be used for different intervals. This enables different frequency bands to use different types of noise reduction at the same time. Thus, for example, a first type of noise reduction can be applied to a first frequency band, while at the same time, a second type of noise reduction can be applied to a second frequency band. This can make the noise reduction applied to the first interval independent of the noise reduction applied to the second interval, where the first and second intervals have different frequencies but overlapping times.
[0106] In some examples, controlling the noise reduction applied to the intervals may include determining when to switch between different methods of noise reduction for one or more intervals. In such examples, two or more different noise reduction methods may be used and device 101 may use Figure 3The method shown in the figure determines when to switch between different methods. The method can make different switching time intervals used for different frequency bands. For example, switching between different noise reduction methods may be more easily perceived by the user on higher frequency bands because there are larger phase differences in these frequency bands, so the time period between switching between different noise reduction methods for higher frequency bands is longer than that for lower frequency bands.
[0107] Figure 4 Another example method of controlling noise reduction is shown. The method may use Figure 1 The device 101 shown and / or Figure 2 The electronic device 201 shown is implemented.
[0108] At block 401, a plurality of audio signals are obtained. The audio signals may include audio obtained from a plurality of microphones 203. The plurality of microphones 203 may be spatially distributed so as to be able to provide a spatial audio signal.
[0109] In blocks 403 and 405, the obtained audio signal is divided into a plurality of intervals. Figure 4 In the example of , the audio signal is divided into multiple time-frequency intervals. These time-frequency intervals can also be referred to as time-frequency tiles. In box 403, the audio signal is divided into time intervals. Once the audio signal is divided into time intervals, the time intervals are converted into the frequency domain. The time domain to frequency domain conversion of the time interval can use more than one time interval. For example, short-time Fourier transform (STFT) can use the current and previous time intervals, and use analysis windows (on two time intervals) and fast Fourier transform (FFT) to perform conversion. Other conversions can use more than two time intervals. In box 405, the frequency domain signal is grouped into frequency subbands. Subbands in different time frames now provide multiple time-frequency intervals.
[0110] At block 407 , it is estimated whether noise is present in the different time-frequency bins. The noise may be wind noise, processing noise, or any other unwanted noise that may be captured by the plurality of microphones 203 .
[0111] Any suitable process may be used to estimate the presence of noise. In some examples, differences in signal levels between different microphones 203 for different frequency bands may be used to determine whether noise exists in different time-frequency intervals. If there are large signal differences between frequency bands, it may be estimated that noise exists in the louder signal.
[0112] In some examples, correlation between microphones 203 may be used to estimate whether noise is present in a time-frequency interval. This may be in addition to or instead of comparing different signal levels.
[0113] In such an example, the plurality of microphones 203 provide signals x m (n'), where m is the microphone index and n' is the sample index. In this example, the time interval is N samples long, n represents the time interval index of the frequency transformed signal. When estimating whether there is noise, the processor 105 is configured to apply a sine window on each input from a different microphone 203 for sample index n'=(n-1)N, ..., (n+1)N-1 for the time interval index n, and transform these windowed input signal sequences into the frequency domain by Fourier transform. This results in a frequency transformed signal X m (k,n), where k is the frequency bin index. This process is called short-time Fourier transform. The frequency domain representation is grouped into B subbands with indices b=0,…,B-1, where each subband has the lowest frequency bin k b,low and the highest frequency bin k b,high , and also include the frequency bins between them.
[0114] For lower frequency bands, the distance between the microphones 203 is short compared to the wavelength of the sound in the frequency band. For such frequency bands, the correlation estimate between the first microphone 203A and the second microphone 203B is In comparison, the signal from the first microphone 203A A high power estimate of indicates the presence of noise in the signal from the first microphone 203A.
[0115] The process of determining whether noise is present may also take into account other factors that may affect the difference in signal levels. For example, the body of the electronic device 201 will mask the audio, making the audio from the source to the electronic device 201 louder in the microphone 203 on the same side as the source, and the audio is attenuated by the masking of the electronic device 201 in the microphone 203 on the other side. This masking effect is greater at higher frequencies, and the signal level differences caused by the masking need to be considered when estimating whether noise is present. This may mean using different thresholds in the signal level for different frequency bands to estimate whether noise is present. For example, there may be a higher threshold for higher frequency bands, so that a larger difference between the signal levels must be detected before estimating the presence of noise compared to lower frequency bands.
[0116] At block 409, it is determined whether noise reduction was used in a previous time-frequency interval.The previous time-frequency interval may be a time-frequency interval immediately preceding the previous time-frequency interval in a given frequency band.
[0117] If noise reduction is used, it is determined at block 411 whether the current previous time-frequency interval being analyzed requires noise reduction. For example, it may be determined whether the noise level within the time-frequency interval is low enough that noise reduction is not required. This may be determined by determining whether the noise level is above or below a threshold.
[0118] In some examples, determining whether noise reduction is needed may include determining the number of microphones 203 that have provided signals with low noise levels. For example, if there are two or more microphones 203 with low noise levels, this may enable providing sufficiently high quality signals without applying noise reduction.
[0119] For example, if the difference between the least noisy microphone signal and the next noisiest microphone signal does not exceed expected masking effects, then it can be estimated that the two signals include sufficiently low noise levels that noise reduction is not needed. The two low noise microphone signals can be used to create a spatial audio signal.
[0120] Masking can depend on the arrangement of microphones 203 and the frequency of the captured sound. In some examples, masking can be determined experimentally, such as by playing audio from different directions toward electronic device 201 in an anechoic chamber. In some examples, the expected energy difference between the signal obtained by first microphone 203A and the signal obtained by second microphone 203B can be estimated using a lookup equation:
[0121] ShadowAB=ShdAB(direction)*ratio.
[0122] For highly directional sounds the ratio increases towards 1 and for weakly correlated inputs the ratio decreases towards 0. Table ShdAB values may be determined by laboratory measurements or any other suitable method.
[0123] In other examples, a different value may be used as a threshold for determining whether noise reduction is needed. This different value may be used instead of or in addition to the desired effect of masking. Other values that may be used include any one or more of: a frequency-dependent but signal-independent fixed threshold adjusted for the electronic device 201 based on testing, a correlation-based measure (which takes into account that microphone signals naturally become less correlated at high frequencies and in the presence of wind noise), a maximum phase shift between microphone signals, where the maximum phase shift depends on frequency and microphone distance, or any other suitable value.
[0124] In other examples, determining whether noise reduction is needed may include determining whether the cross-correlation between the microphone signals is above a threshold. This may be used for low frequencies where the wavelength of the captured sound is long relative to the spacing between the microphones 203. In such an example, the cross-correlation between the signals captured by a pair of microphones 203 may be normalized relative to the microphone energy to produce a normalized cross-correlation value between 0 and 1, where 0 represents an orthogonal signal and 1 represents a perfectly correlated signal. When the normalized cross-correlation is above a threshold such as 0.8, it may indicate that the noise level captured by the microphone pair 203 is low enough that noise reduction is not needed.
[0125] If it is determined in box 411 that noise reduction is needed, then in box 413, it is determined whether the noise reduction method currently needed is the same as the method used in the previous time-frequency interval. This can include determining whether the best method for noise reduction in the time-frequency interval is the same as the method used in the previous time-frequency interval. For example, it can be determined whether the same microphone signal is used for the noise reduction method in the previous time-frequency interval. This can be achieved by checking whether the microphone 203 providing the lowest noise signal is the same as the microphone 203 providing the lowest noise signal in the previous time-frequency interval.
[0126] If it is determined in block 413 that the noise reduction methods are different, then it is determined in block 415 whether the noise reduction time limit has been exceeded. That is, it is determined whether the same noise reduction method has been used for a time period exceeding a threshold. Different time periods may be used for thresholds in different frequency bands.
[0127] The threshold for the time period can be selected by estimating whether switching to a different noise reduction method would result in more perceptual artifacts than leaving noise without the switch. In the example where the noise reduction method includes switching between different microphones, this can be estimated by the following equation:
[0128]
[0129] in
[0130] prevenergy is the energy in the current time-frequency interval of the microphone 203 used in the previous time-frequency interval
[0131] CurrentEnergy is the energy of the current time-frequency interval of the microphone 203 with the minimum noise
[0132] maxphase is the maximum phase shift that can occur when switching from the microphone 203 used in the previous time-frequency interval to the microphone 203 that currently has the least noise. The phase takes into account the distance between the microphones 203 and the frequency band of the time-frequency interval. For frequencies where half the wavelength of the sound is greater than the distance between the microphones 203, this is a maximum phase shift of 180°,
[0133] ·w phase is the weighting factor,
[0134] time is the time (in seconds) when the most recent switch occurred and the threshold time TH The threshold time is selected so that switching between different microphones 203 does not occur every time the lowest noise microphone 203 changes. The threshold time may be between 10 and 100 milliseconds or in any other suitable range.
[0135] ·w time is the weighting factor,
[0136] ●shadow is the maximum acoustic shadow caused by the electronic device 201,
[0137] ●safety is a constant that is used to estimate the error in the estimate and slow down the switching speed according to the wrong estimate.
[0138] The values in the equation may be calculated for a single microphone 203 or for multiple microphones 203. In the case where values are calculated for multiple microphones 203, average values may be used for the terms in the equation.
[0139] If the time limit has not been exceeded, the noise reduction method used for the previous time-frequency interval is applied to the current time-frequency interval at block 417. That is, there will not be any switching in the noise reduction method used to avoid artifacts being perceived by the user.
[0140] If the time limit is exceeded, then the best noise reduction method for the current time-frequency interval is selected and applied to the current time-frequency interval at block 419. In such an example, it may have been determined that switching between different noise reduction methods will result in fewer artifacts than noise within the audio signal.
[0141] If it is determined in box 413 that the current best noise reduction method is the same as the method used in the previous time-frequency interval, the method proceeds to box 419, and the best noise reduction method for the current time-frequency interval is selected and applied to the current time-frequency interval. In this case, there will be no switching between different types of noise reduction.
[0142] If it is determined at block 411 that noise reduction is not required, then it is determined at block 421 whether a switching threshold is exceeded. It may be determined whether switching from applying noise reduction to not applying noise reduction would result in more perceptual artifacts than applying noise reduction. The threshold may be a comparison between an estimated noise level in a time-frequency bin and an estimated artifact caused by the switch.
[0143] If the threshold is not exceeded, the method will proceed to block 413 and follow the procedures described in blocks 413, 415, 417 and 419. If it is determined that noise reduction is best applied in this case, it will be that no noise reduction is applied in this case.
[0144] If it is determined at block 421 that the threshold is not exceeded, then at block 423 the noise reduction is controlled so that no noise reduction is applied to the time-frequency bin. This may be applied without the process of blocks 413, 415 and 419.
[0145] If it is determined at block 409 that noise reduction was not used in the previous time-frequency interval, the method moves to block 425. A determination is made at block 425 whether noise reduction is required. The process used at block 425 may be the same as the process used at block 411.
[0146] If it is determined at block 425 that noise reduction is needed, the process moves to block 427. At block 427, it is determined whether a switching threshold is exceeded. It may be determined that switching from not applying noise reduction to applying noise reduction will cause more perceptual artifacts than not applying noise reduction. The threshold may be a comparison between an estimated noise level in a time-frequency bin and an estimated artifact caused by the switch.
[0147] In some examples, the switching threshold can be a fixed time limit that must elapse since the last switch between different noise reduction methods. The time limit can be 0.1 seconds or any other suitable time limit. In other examples, the time limit can be estimated based on different signal levels and artifacts caused by the switch. In some examples, different time limits can be used for different frequency bands.
[0148] The switching threshold for switching from not applying noise reduction to applying some noise reduction can be a shorter time limit than the switching threshold for switching from applying some noise reduction to not applying noise reduction. This is because noise may appear suddenly, so it is beneficial to be able to turn noise reduction on more quickly than to turn it off quickly.
[0149] If the switching threshold is not exceeded, the process moves to block 423 and no noise reduction is applied to the current time-frequency interval. In this case, there is no switching between different noise reduction methods, as this is considered to provide a lower quality signal than the noise itself.
[0150] If the switching threshold is exceeded, sufficient time has passed since the last switch in the noise reduction method, and the process moves to block 429. At block 429, noise reduction is applied to the current time-frequency interval. The noise reduction applied may be the noise reduction that has been determined to be optimal for the noise level within the current noise frequency interval.
[0151] If it is determined at block 425 that noise reduction is not required, the process moves to block 431 and noise reduction is not applied to the current time-frequency bin.
[0152] Once based on Figure 4 The process shown determines whether to apply or not apply noise reduction, the method moves to block 433 and the time-frequency bins are converted back to the time domain. The time domain signal may then be stored in the memory 107 and / or provided to a rendering device for rendering to a user.
[0153] It will be appreciated that blocks 407 to 433 will be repeated as needed for each time-frequency interval. In some examples, the method may be repeated for each time-frequency interval. In some examples, the method may be repeated for only a subset of time-frequency intervals.
[0154] Figure 3 and Figure 4 An example of the method shown in . The advantage provided is that it is possible to use different noise reduction methods for different frequency bands. The method also allows the use of different criteria to determine when to switch between different noise reduction methods for different frequency bands. Therefore, this provides an improved quality audio signal with a reduced noise level.
[0155] Figure 5 Another example method of controlling noise reduction is shown. The method may use Figure 1 The device 101 shown and / or Figure 2 The electronic device 201 shown is implemented.
[0156] The method includes, at block 501, obtaining one or more audio signals, wherein the one or more audio signals represent sound signals captured by a plurality of microphones 203. In some examples, the one or more audio signals include audio obtained from a microphone 203, which is disposed in the same electronic device 201 as the apparatus 101. In other examples, the one or more audio signals include audio obtained from a microphone 203 disposed in one or more separate devices. In such examples, the one or more audio signals may be sent to the apparatus 101.
[0157] The obtained audio signal may include an electrical signal representing at least some of the sound field captured by the plurality of microphones 203. The output signal provided by the microphone 203 may be modified to provide the audio signal. For example, the output signal from the microphone 203 may be filtered or equalized, or any other suitable processing may be performed thereon.
[0158] The obtained audio signal can be captured by the spatially distributed microphones 203, so that a spatial audio signal can be provided to the user. The spatial audio signal can be a stereo signal, a binaural signal, a panoramic sound signal or any other suitable type of spatial audio signal.
[0159] The method further includes, at block 503, dividing the obtained one or more audio signals into a plurality of intervals. The obtained one or more audio signals may be divided into intervals using any suitable process. The intervals may be time-frequency intervals, time intervals, or any other suitable type of intervals.
[0160] In some examples, the intervals may be of different sizes. For example, the frequency bands used to define the intervals may have different sizes for different frequencies. For example, a lower frequency interval may cover a smaller frequency band than a higher frequency interval.
[0161] At block 505, the method includes determining one or more parameters associated with one or more noise characteristics of different intervals. In some examples, the parameters may be determined for each region. In other examples, the parameters may be determined only for a subset of the intervals.
[0162] The noise characteristics being analyzed are related to the noise detected by one or more of the multiple microphones 203 that capture the audio of the one or more audio signals. The noise may be an unwanted sound in the audio signal captured by the microphone 203. The noise may include noise that does not correspond to the sound field captured by the multiple microphones 203. For example, the noise may be wind noise, operating noise, or any other suitable type of noise. In some examples, the noise may include noise caused by other components of the electronic device 201. For example, the noise may include noise caused by the focusing of a camera within the electronic device 201. The noise characteristics being analyzed may exclude noise introduced by the microphone 203.
[0163] The one or more parameters related to the one or more noise characteristics may include any parameters that provide an indication of noise level and / or noise reduction methods that will improve the audio quality of the interval being analyzed.
[0164] In some examples, one or more parameters related to noise characteristics may include noise level in an interval. The noise level may be determined by monitoring signal level differences between frequency bands, monitoring correlations between audio signals captured by different microphones 203, or any other suitable method.
[0165] In some examples, the noise level in intervals prior to the interval being analyzed can be monitored. For example, to determine the noise level in a given frequency band, the noise in the previous time period can be determined. The probability of a significant change in the noise level in the next interval can then be predicted based on the noise level in the previous interval. This can therefore take into account the fact that a single time interval may show a small amount of noise, but this may be an anomaly in other noisy time periods.
[0166] In some examples, the one or more parameters related to the noise characteristics may include parameters related to the noise reduction method currently being used or previously used. In such examples, the one or more parameters may include the noise reduction method used for the previous time interval in the frequency band, the duration that the current noise reduction method has been used, or any other suitable parameters.
[0167] Using parameters associated with the noise reduction method, it is possible to implement how often switching between different types of noise reduction methods occurs. This can reduce artifacts caused by switching between different types of noise reduction, and therefore can improve the audio quality perceived by the user.
[0168] Other types of parameters related to noise levels may also be used in other examples of the present disclosure. For example, in some examples, the direction of the microphone 203 capturing the audio signal or any other suitable parameter may be used. The direction of the microphone may give an indication of effects such as occlusion, which may affect the level of audio captured by the microphone from different directions, thereby affecting the noise captured by the microphone.
[0169] Parameters related to noise characteristics may be determined independently for the intervals. For example, the analysis performed for a first interval may be independent of the analysis performed for a second interval. This may mean that the analysis and determination performed for the first interval does not affect the analysis and determination performed for the second interval.
[0170] In some examples, determining one or more parameters associated with the noise characteristic includes determining whether the one or more parameters are within a threshold range. Determining whether a parameter is within a threshold range may include determining whether a value of the parameter is above or below a threshold. In some examples, determining whether a parameter is within a threshold range may include determining whether the value of the parameter is between an upper limit value and a lower limit value.
[0171] The value of the threshold may be different for different intervals. For example, different thresholds for one or more parameters related to noise characteristics can be used for different frequency ranges within multiple intervals. This can take into account the fact that different frequency bands may be more affected by noise than other frequency bands. For example, wind noise may be more noticeable in lower frequency bands than in higher frequency bands. Switching between different noise reduction methods may also be more noticeable to the user in higher frequency bands. This may make it undesirable to switch between different noise reduction methods too frequently for higher frequency bands. Therefore, in the examples of the present disclosure, different thresholds for the time period between switches can be used for different frequency bands.
[0172] At block 507, the method includes determining whether to provide mono audio output or spatial audio output based on the determined one or more parameters. The mono audio output may include an audio signal including audio from two or more channels, wherein the audio signal for each channel is substantially identical.
[0173] A mono audio output may be more robust than a spatial audio output and may therefore provide a reduced noise level. Thus, providing a mono audio output rather than a spatial audio output may provide a reduced noise output for the audio signal.
[0174] In some examples, if it is determined to provide a mono audio output, the microphone signal with the least noise can be determined so that it can be used to provide the mono audio output. In some examples, a mono audio output can be provided by combining two or more microphone signals from multiple microphones 203. In such examples, the microphones 203 can be placed close to each other. For example, the microphones 203 can be located at the same end of the electronic device.
[0175] In examples of the present disclosure, different parameters may be determined differently for different frequency bands within multiple intervals. In such examples, this may enable a mono audio output to be provided for a first frequency band while a spatial audio output may be provided for a second frequency band. This enables a mono audio output to be provided for a first frequency band within an interval while a spatial audio output is provided for a second frequency band within an interval, wherein the first and second intervals have different frequencies but overlap in time.
[0176] Figure 6 Another example electronic device 601 is shown. The example electronic device 601 may be used to implement Figure 5 and Figure 7 In some examples, the electronic device 601 may also implement Figure 3 and Figure 4 It should also be understood that, for example, Figure 2 Other electronic devices shown in the electronic device 201 can be used to implement Figure 5 and7 The method shown in .
[0177] Figure 6 The example electronic device 601 includes an apparatus 101, which can be as follows Figure 1 As shown. The apparatus 101 may include a processor 105 and a memory 107 as described above. The example electronic device also includes a plurality of microphones 203. Figure 6 In the example of 601, the electronic device 601 includes two microphones.
[0178] The electronic device 601 may be a communication device such as a mobile phone. It should be understood that the communication device may include Figure 6 Components not shown in the figure, for example, the communication device may include one or more transceivers capable of wireless communication.
[0179] In some examples, the electronic device 601 may be an image capture device. In such an example, the electronic device 601 may include one or more cameras capable of capturing images. The images may be video images, still images, or any other suitable type of images. The images captured by the camera module may be accompanied by sound signals captured by the plurality of microphones 203.
[0180] The plurality of microphones 203 may include any device configured to capture sound and capable of providing one or more audio signals. The one or more audio signals may include electrical signals representing at least some of the sound fields captured by the plurality of microphones 203. The output signals provided by the microphones 203 may be modified to provide audio signals. For example, the output signals from the microphones 203 may be filtered or equalized, or any other suitable processing may be performed thereon.
[0181] The electronic device 601 is configured so that an audio signal including audio from the plurality of microphones 203 is provided to the apparatus 101. This enables the apparatus 101 to process the audio signal. In some examples, it may enable the apparatus 101 to process the audio signal to reduce the effects of noise captured by the microphones 203.
[0182] Multiple microphones 203 may be located within the electronic device 601 to enable the capture of spatial audio. For example, the locations of the multiple microphones 203 may be distributed throughout the electronic device 601 to enable the capture of spatial audio. Spatial audio includes an audio signal that may be rendered so that a user of the electronic device 601 may perceive the spatial characteristics of the audio signal. For example, spatial audio may be rendered so that a user may perceive the source direction and distance from the audio source.
[0183] exist Figure 6In the example shown, the electronic device 601 includes two microphones 203. A first microphone 203A is disposed at a first end on a first surface of the electronic device 601. A second microphone 203B is provided at a second end of the electronic device 601. The second end is an end of the electronic device 601 opposite to the first end. The second microphone 203B is disposed on the same surface as the first microphone 203A. It should be understood that in other examples of the present disclosure, other configurations of multiple microphones 203 may be provided.
[0184] Multiple microphones 203 are coupled to the device 101. This can enable audio signals captured by the multiple microphones 203 to be provided to the device 101. This can enable the audio signals to be stored in the memory 107. This can also enable the processor 105 to perform noise reduction on the obtained audio signals. Example methods of noise reduction are as follows Figure 5 and Figure 7 shown.
[0185] exist Figure 6 In the example shown, the microphone 203 that captures audio and the processor 105 that performs noise reduction are provided in the same electronic device 601. In other examples, the microphone 203 and the processor 105 that performs noise reduction may be provided in different electronic devices 601. For example, the audio signal may be transmitted from the plurality of microphones 203 to the processing device via a wireless connection or some other suitable communication link.
[0186] Figure 7 Another example method of controlling noise reduction is shown. The method may use Figure 1 The device 101 shown and / or Figure 6 The electronic device 601 shown is implemented.
[0187] At block 701, a plurality of audio signals are obtained. The audio signals may include audio obtained from a plurality of microphones 203. The plurality of microphones 203 may be spatially distributed so as to be able to provide a spatial audio signal. Figure 7 In the example, two audio signals are obtained.
[0188] In blocks 703 and 705, the obtained audio signal is divided into a plurality of intervals. Figure 7In the example of , the audio signal is divided into multiple time-frequency intervals. These time-frequency intervals can also be referred to as time-frequency tiles. In box 703, the audio signal is divided into time intervals. Once the audio signal has been divided into time intervals, the time intervals are converted into the frequency domain. The time domain to frequency domain conversion of the time interval can use more than one time interval. For example, a short-time Fourier transform (STFT) can use the current and previous time intervals, and use an analysis window (on two time intervals) and a fast Fourier transform (FFT) to perform the conversion. Other conversions can use time intervals other than two time intervals. In box 705, the frequency domain signal is grouped into frequency subbands. The subbands in different time frames now provide multiple time-frequency intervals.
[0189] At block 707, microphone signal energies are calculated for different time-frequency bins. Once the microphone signal energies have been calculated, the energies of different time-frequency bins may be compared.
[0190] At block 709 , it is estimated whether noise is present in the time-frequency bin. The noise may be wind noise, processing noise, or any other unwanted noise that may be captured by the plurality of microphones 203 .
[0191] Any suitable process can be used to estimate whether there is noise. In some examples, a comparison of the energy of different time-frequency intervals can be used to determine whether there is noise. If there is a large energy difference between frequency bands, it can be estimated that there is noise in the larger signal.
[0192] The process of determining whether there is noise can take into account factors that may affect the difference in signal levels, such as shielding. For example, the body of the electronic device 601 will shield the audio so that the audio from the source to the electronic device 601 is louder in the microphone 203 on the same side as the source, and the audio on the other side of the microphone 203 is attenuated by the shielding of the electronic device 601. This shielding effect is greater at higher frequencies, and the signal level differences caused by shielding need to be considered when estimating whether there is noise. This may mean using different thresholds for signal level differences for different frequency bands to estimate whether there is noise. For example, there may be a higher threshold for higher frequency bands, so that a larger difference between signal levels must be detected before estimating the presence of noise compared to lower frequency bands.
[0193] In other examples, other methods for determining whether noise is present may be used instead. For example, cross-correlation of energy in different time-frequency bins may be used.
[0194] For different frequency bands in multiple time-frequency intervals, the threshold value for determining whether there is noise in the time-frequency interval can be different. The threshold value is selected so that compared with the high frequency band, the device 101 is more likely to use the mono audio output for the low frequency band. For example, compared with the lower frequency, a higher frequency can use a higher signal difference threshold. In some examples, the threshold value can be 10dB for the low frequency band, and can be 15dB for the high frequency band. In other examples, the threshold value can be 5dB for the low frequency band, and can be 10dB for the high frequency band. It should be understood that other values of the threshold value can be used in other examples of the present disclosure. This takes into account the fact that the lower frequency band is more susceptible to noise than the higher frequency band. This may also take into account that it may be more difficult to accurately detect the presence of noise in the higher frequency band.
[0195] If it is estimated at block 709 that there is noise, the method moves to block 711. At block 711, the microphone signal with the least noise is used to provide a mono audio output. In some examples, there may be two or more microphones 203 that provide signals with low noise. However, if these microphones 203 are close together, for example if they are located at the same end of the electronic device 601, the two microphone signals may be combined to provide a mono audio output. The microphone signals may be combined by summing or using any other suitable method.
[0196] If it is estimated that no noise is present at block 709, or if the estimated noise presence is below a threshold, the method moves to block 713. At block 713, the two or more microphone signals are used to provide a spatial audio output. The spatial audio output may be a stereo signal, a binaural signal, an panoramic signal, or any other suitable spatial audio output. It should be appreciated that any suitable process may be used to generate a spatial audio output from the obtained audio signal.
[0197] Once Figure 7 If the process shown provides a mono audio output or a spatial audio output, the method moves to block 715 and converts the time-frequency interval back to the time domain. The time domain signal may then be stored in the memory 107 and / or provided to a rendering device for rendering to a user.
[0198] It will be appreciated that blocks 707 to 714 will be repeated as needed for each time-frequency interval. In some examples, the method may be repeated for each time-frequency interval. In some examples, the method may be repeated for only a subset of time-frequency intervals.
[0199] Thus, examples of the present disclosure provide an audio output signal with an improved noise level by controlling switching between spatial audio output and mono audio output for different frequency bands. This takes into account that lower frequency bands are more susceptible to noise than higher frequency bands.
[0200] Since humans are less sensitive to the direction of higher frequency sounds, limiting the lower frequencies to a mono audio output may also result in fewer artifacts perceived by the user.
[0201] It should be understood that the above-described example methods and apparatus 101 may be modified. For example, when capturing audio signals, the effects of noise may depend on the orientation of the electronic device 201, 601. This may mean that some microphones 203 are more likely to be affected by noise when the electronic device 201, 601 is used in a first orientation than when the electronic device 201, 601 is used in a second orientation. This information may then be used when selecting a noise reduction method to be used or when selecting between a mono audio output and a spatial audio output. For example, it may enable the application of different thresholds and / or weighting factors in order to bias the use of microphone signals that are less likely to be affected by noise for a given orientation of the electronic device 201, 601.
[0202] The above examples find application in the following components: automotive systems; telecommunications systems; electronic systems, including consumer electronics; distributed computing systems; media systems for generating or presenting media content, including audio, video and audiovisual content and mixed, mediated, virtual and / or augmented reality; personal systems, including personal health systems or personal fitness systems; navigation systems; user interfaces also known as human-machine interfaces; networks, including cellular, non-cellular and optical networks; ad hoc networks; the Internet; the Internet of Things; virtualized networks; and related software and services.
[0203] The term "comprising" is used in this document in an inclusive rather than exclusive sense. That is, any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If the exclusive sense of "comprising" is intended, this will be made clear in the context by reference to "comprising only one..." or by the use of "consisting of..."
[0204] In this specification, reference is made to various examples. The description of features or functions associated with an example indicates that these features or functions are present in the example. The use of the term "example" or "for example" or "may" or "might" in the text indicates that such features or functions are present in at least the example described, whether or not explicitly stated, whether or not described as an example, and they may but not necessarily be present in some or all other examples. Therefore, "example" or "for example" or "may" or "might" refers to a specific instance in a class of examples. The attributes of an instance may be attributes of only the instance or attributes of a class or attributes of a subclass of a class, the subclass including some but not all instances in the class. Therefore, features described with reference to one example rather than with reference to another example are implicitly disclosed, and may be used as part of a working combination in the other examples where possible, but do not necessarily have to be used in the other examples.
[0205] Although the embodiments have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
[0206] Features described in the preceding description may be used in combinations other than the combinations explicitly described above.
[0207] Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
[0208] Although features have been described with reference to certain embodiments, those features may also be present in other embodiments whether described or not.
[0209] The terms "a", "an" or "the" used in this document have an inclusive rather than exclusive meaning. That is, any reference to X that includes one / the Y indicates that X may include only one Y or may include more than one Y, unless the context clearly indicates otherwise. If the exclusive meaning of "a" or "the" is intended, it will be clearly stated in the context. In some cases, "at least one" or "one or more" may be used to emphasize the inclusive meaning, but the absence of these terms should not be regarded as inferring and exclusive meaning.
[0210] The presence of a feature (or combination of features) in a claim is a reference to the feature or (combination of features) itself, and also a reference to features that achieve substantially the same technical effect (equivalent features). Equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function in substantially the same way to achieve substantially the same result.
[0211] In this specification, reference has been made to various examples using adjectives or adjective phrases to describe features of examples. Such descriptions of characteristics associated with examples indicate that the characteristics exist in some examples exactly as described, and exist in other examples substantially as described.
[0212] While an effort has been made in the foregoing description to call attention to those features regarded as essential, it should be understood that the applicant may seek protection by way of claims for any patentable feature or combination of features mentioned above and / or shown in the drawings whether or not this has been emphasized.
Claims
1. A device comprising: at least one processor; as well as at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: obtaining one or more audio signals, wherein the one or more audio signals include audio captured using a plurality of microphones; Dividing the obtained one or more audio signals into a plurality of intervals; determining one or more parameters associated with one or more noise characteristics of each of the plurality of intervals; and Noise reduction applied to each of the plurality of intervals is controlled based on one or more parameters determined within each of the plurality of intervals.
2. The device according to claim 1, wherein: Controlling the applied noise reduction comprises at least the instructions, when executed by the at least one processor, causing the apparatus to: Based on the determined one or more parameters, it is determined whether to provide a basic mono signal or a spatial signal.
3. The device according to claim 2, wherein: The basic mono signal is from a smaller subset of the plurality of microphones than the spatial signal.
4. The device according to claim 2, wherein: The basic mono signal includes two or more channels, wherein audio signals of each of the two or more channels are substantially the same.
5. The device according to claim 1, wherein: The plurality of bins include time-frequency bins.
6. The device according to claim 1, wherein: The one or more noise characteristics include a noise level.
7. The device according to claim 1, wherein: The one or more parameters related to the one or more noise characteristics are independently determined for each of the plurality of intervals.
8. The device according to claim 1, wherein: The instructions, when executed by the at least one processor, cause the apparatus to: A determination is made as to whether the one or more parameters are within a threshold range.
9. The device according to claim 8, wherein: Different thresholds for the one or more parameters associated with the one or more noise characteristics are used for different frequency ranges within the plurality of intervals.
10. The device according to claim 1, wherein: The one or more parameters related to the one or more noise characteristics include one or more of the following: Noise level within the interval; the noise level in the interval preceding the interval being analyzed; noise reduction method for the previous frequency interval; the duration that the current noise reduction method has been used in the frequency band; or Orientations of the plurality of microphones capturing the one or more audio signals.
11. The device according to claim 1, wherein: The noise reduction applied to the first interval is independent of the noise reduction applied to the second interval, wherein the first interval and the second interval have different frequencies but overlap in time.
12. The device according to claim 1, wherein: Different noise reduction is applied to different intervals, where the different intervals have different frequencies but overlap in time.
13. The device according to claim 1, wherein: Controlling the noise reduction applied to the intervals includes the instructions, when executed by the at least one processor, causing the apparatus to perform one or more of: selecting a method for noise reduction within the interval; determining when to switch between different methods for noise reduction within one or more intervals; Provides noise-reduced spatial output; Provides spatial output without noise reduction Provides noise-reduced mono audio output; Providing beamforming output; or Provides noise-reduced beamformed output.
14. The device according to claim 1, wherein: The one or more noise characteristics are associated with one or more of: noise within the one or more audio signals that has been detected using one or more of the plurality of microphones capturing audio; Wind noise; or Dealing with noise.
15. A method comprising: obtaining one or more audio signals, wherein the one or more audio signals include audio captured using a plurality of microphones; Dividing the obtained one or more audio signals into a plurality of intervals; determining one or more parameters associated with one or more noise characteristics of each of the plurality of intervals; and Noise reduction applied to each of the plurality of intervals is controlled based on one or more parameters determined within each of the plurality of intervals.
16. An apparatus comprising: at least one processor; as well as at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: obtaining one or more audio signals, wherein the one or more audio signals include audio captured using a plurality of microphones; Dividing the obtained one or more audio signals into a plurality of intervals; determining one or more parameters associated with one or more noise characteristics of each of the plurality of intervals; and Based on the determined one or more parameters, it is determined whether to provide a basic mono signal or a spatial signal.
17. The device according to claim 16, wherein: The basic mono signal is from a smaller subset of the plurality of microphones than the spatial signal.
18. The device according to claim 16, wherein: The basic mono signal includes two or more channels, wherein audio signals of each of the two or more channels are substantially the same.
19. The device according to claim 16, wherein: The plurality of bins include time-frequency bins.
20. The device according to claim 16, wherein: Determining one or more parameters related to the one or more noise characteristics includes the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to: determining whether energy differences between microphone signals from different microphones within the plurality of microphones are within a threshold range; and A determination is made whether a switch has been made between the base mono signal and the spatial signal within a threshold time.