Apparatus, method and computer program for noise suppression
By using multiple microphone signals and program codes to predict the output signal in the audio communication setting, the problems of high noise suppression complexity and strict delay requirements in the prior art are solved, and effective noise suppression effect with low complexity and loose delay requirements are achieved.
Patent Information
- Application Number
- CN202411869736.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems of high complexity and strict delay requirements in noise suppression of audio signals, making it difficult to effectively improve voice intelligibility and quality in communication settings.
By using at least two microphone signals, the audio signals of the current frame or the previous frame are obtained and the output signals are predicted in part based on these signals for noise suppression processing of future frames. The device uses an output signal to process the microphone signal of a future frame in the first audio signal processing, and continues using the output signal to process the output future frames outputted for the previous processing in the second audio signal processing.
Realizes noise suppression with low complexity and loose delay requirements, suitable for audio communication settings, and improves voice intelligibility and quality.
Smart Images

Figure CN120199266A_ABST
Abstract
Description
Technical Field
[0001] Examples of the present disclosure relate to apparatuses, methods, and computer programs for noise suppression. Some relate to apparatuses, methods, and computer programs for noise suppression of audio signals in communication setups. Background Art
[0002] Noise suppression for audio signals can be used in communication setups to improve the intelligibility of speech and / or other desired sounds. Program code such as machine learning programs can be used to implement noise suppression. Summary of the Invention
[0003] According to various but not necessarily all examples of the present disclosure, there is provided an apparatus for noise suppression, including components for the following operations:
[0004] Obtaining at least one audio signal for the current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames;
[0005] Using program code to predict an output signal for a future frame based at least in part on at least one audio signal for the current frame or one or more previous frames; and
[0006] Using the output signal to process future frames of the at least two microphone signals in a first audio signal processing and using the output signal to process future frames of the output of the first audio signal processing in a second audio signal processing to achieve noise suppression.
[0007] The at least one audio signal may include the output of the first audio signal processing.
[0008] The first audio signal processing and the second audio signal processing may be consecutive processes, where the output of the first audio signal processing is provided as the input of the second audio signal processing.
[0009] The first audio signal processing may include beamforming processing, and wherein the beamforming processing includes using the output signal to process future frames of the at least two microphone signals.
[0010] The output signal may include a gain to be applied to at least one of the at least two microphone signals to be used in the beamforming processing.
[0011] The output signal may include an amplitude of at least one of the at least two microphone signals to be used in the beamforming processing.
[0012] The second audio signal processing may include spectral noise suppression processing.
[0013] The output signal may include a gain for an input to be applied to spectral noise suppression processing.
[0014] The output signal may include an amplitude used for spectral noise suppression processing.
[0015] The output signal may be applied to future frames of each audio signal processing in the frequency domain.
[0016] The program code may receive a single input and provide a single output.
[0017] The program code may include a machine learning program.
[0018] The machine learning program may include a neural network circuit.
[0019] The same output signal may be applied to future frames of multiple audio signal processes.
[0020] The number of current frames or previous frames in the acquired audio signals used to predict the output signal for future frames may be selected at least partially based on a latency requirement.
[0021] The apparatus may be used in an audio communication setup.
[0022] The audio communication setup may be at least one of the following;
[0023] A one-way communication setup; or
[0024] A two-way communication setup.
[0025] According to various but not necessarily all examples of the present disclosure, an electronic device including the apparatus as described herein is provided, where the electronic device is at least one of the following: a telephone, a camera, a computing device, a teleconferencing device, a television, a virtual reality device, an augmented reality device.
[0026] According to various but not necessarily all examples of the present disclosure, a method is provided, including:
[0027] Obtaining at least one audio signal for a current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames;
[0028] Using program code to predict an output signal for a future frame at least partially based on at least one audio signal for the current frame or one or more previous frames; and
[0029] Using the output signal to process future frames of at least two microphone signals in a first audio signal process, and using the output signal to process future frames of the output of the first audio signal process in a second audio signal process to achieve noise suppression.
[0030] According to various but not necessarily all examples of the present disclosure, a computer program including instructions is provided, which when executed by a device cause the device to at least perform:
[0031] Obtain at least one audio signal for a current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames;
[0032] Use program code to predict an output signal for a future frame based at least in part on at least one audio signal for the current frame or one or more previous frames; and
[0033] Use the output signal in a first audio signal processing to process future frames of the at least two microphone signals, and use the output signal in a second audio signal processing to process future frames of the output of the first audio signal processing to achieve noise suppression.
[0034] Although the above examples and optional features of the present disclosure are described separately, it should be understood that they are provided in all possible combinations and permutations within the present disclosure. It should be understood that various examples of the present disclosure may include any or all features described in other examples of the present disclosure, and vice versa. Moreover, it should be understood that any one or more or all features in any combination may be implemented / included in / executable by a device, method, and / or computer program instructions as needed and appropriately. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Some examples will now be described with reference to the drawings, where:
[0036] Figure 1 An example system is shown;
[0037] Figure 2 An example of multi-channel noise suppression is shown;
[0038] Figure 3 An example method is shown;
[0039] Figure 4 An implementation of an example of the present disclosure is shown;
[0040] Figure 5 An example computing pipeline is shown;
[0041] Figure 6 An example system is shown;
[0042] Figure 7 An example method is shown;
[0043] Figure 8 An example machine learning program is shown;
[0044] Figure 9 shows an example architecture of a machine learning program;
[0045] Figure 10 shows an exemplary room and microphone array;
[0046] Figure 11 shows a graph of example results;
[0047] Figure 12 shows a graph of example results; and
[0048] Figure 13 shows an example device.
[0049] The drawings are not necessarily to scale. For clarity and conciseness, some features and views of the drawings may be shown schematically or enlarged in scale. For example, the dimensions of some elements in the figures may be enlarged relative to other elements to aid in illustration. Corresponding reference numerals are used in the drawings to denote corresponding features. For clarity, it is not necessary to show all reference numerals in all the drawings. DETAILED DESCRIPTION
[0050] Noise suppression of an audio signal can be used in a communication setting to improve the intelligibility and / or quality of speech and / or other desired sounds. Speech intelligibility reflects how well the speech content can be understood. Speech quality is used to describe the comfort of a person listening to speech. A communication setting may require simultaneous capture and playback of an audio signal, which can be challenging for performing noise suppression. Program code such as a machine learning program can be used to implement noise suppression.
[0051] Examples of the present disclosure provide improved noise suppression. Examples of the present disclosure can be used in communication settings and / or any other suitable settings that use simultaneous capture and playback of an audio signal. Examples of the present disclosure enable program code with low complexity and relaxed latency requirements to be used to implement noise suppression and / or any other suitable processing.
[0052] Figure 1 Shows an example system 101 that can be used to implement examples of the present disclosure. The example system 101 provides a communication setting.
[0053] Figure 1 The system 101 shown can be used for audio communication. Audio communication can include voice communication. Audio from a proximal user can be detected, processed, and transmitted for presentation and playback to a distal user. In some examples, audio from a proximal user can be stored in an audio file for later use. Examples of the present disclosure can also be used in other systems and / or variations of the system 101.
[0054] The system 101 includes a first user device 103A and a second user device 103B. InFigure 1 In the example shown, each of the first user device 103A and the second user device 103B includes a mobile phone. In other examples of the present disclosure, other types of user devices 103 may be used. For example, the user device 103 may be a telephone, a tablet computer, a soundbar, a microphone array, a camera, a computing device, a teleconferencing device, a television, a virtual reality (VR) / augmented reality (AR) device, or any other suitable type of communication device.
[0055] The user devices 103A, 103B include one or more microphones 105A, 105B and one or more speakers 107A, 107B. The one or more microphones 105A, 105B are configured to detect an acoustic signal and convert the acoustic signal into an output electrical audio signal. The output signals from the microphones 105A, 105B may provide microphone signals. The one or more speakers 107A, 107B are configured to convert an input electrical signal into an output acoustic signal that can be heard by the user.
[0056] The user devices 103A, 103B may also be coupled to one or more peripheral playback devices 109A, 109B. The playback devices 109A, 109B may be headphones, speaker devices, or any other suitable type of playback device 109A, 109B. The playback devices 109A, 109B may be configured such that spatial audio or any other suitable type of audio can be played back for the user to listen to. In an example where the user devices 103A, 103B are coupled to the playback devices 109A, 109B, the microphone signals may be processed and provided to the playback devices 109A, 109B instead of being provided to the speakers 107A, 107B of the user devices 103A, 103B.
[0057] The user devices 103A, 103B also include audio processing components 111A, 111B. The processing components 111A, 111B may include any components suitable for processing the microphone signals from the microphones 105A, 105B and / or processing components 111A, 111B configured to process the audio signals provided to the speakers 107A, 107B and / or the playback devices 109A, 109B. The processing components 111A, 111B may include one or more devices 1301 as Figure 13 shown and described below and / or any other suitable components.
[0058] The processing components 111A, 111B can be configured to perform any suitable processing on the microphone signals and / or any other suitable signals. For example, the processing components 111A, 111B can be configured to perform noise suppression, acoustic echo cancellation, residual echo suppression, speech enhancement, speech dereverberation, wind noise reduction, sound source separation, and / or any other suitable processing on the microphone signals and / or any other suitable signals. The processing components 111A, 111B can be configured to perform spatial rendering and dynamic range compression on the input electrical signals of the speakers 107A, 107B and / or the playback devices 109A, 109B. The processing components 111A, 111B can be configured to perform other processing, such as active gain control, source tracking, head tracking, audio focusing, or any other suitable processing.
[0059] The processing components 111A, 111B can be configured to process the microphone signals using a computer program such as a machine learning program. The machine learning program can be configured as described or in any other suitable manner.
[0060] The processed audio signals can be transmitted between the user devices 103A, 103B using any suitable communication network. In some examples, the communication network can include a 4G or 5G or other suitable type of network. The communication network can include one or more codecs 113A, 113B, which can be configured to encode and decode the audio signals appropriately. In some examples, the codecs 113A, 113B can be IVAS (Immersive Voice Audio System) codecs or any other suitable type of codec.
[0061] In a system such as Figure 1 the system 101, noise suppression can be applied by considering signals from multiple microphones. This can be referred to as multi-channel noise suppression.
[0062] Figure 2 An example digital signal processing (DSP) chain 201 that can be used for multi-channel noise suppression is shown. The DSP chain 201 includes multiple audio signal processes. The multiple audio signal processes can be applied to the signals from multiple microphones 105. In Figure 2 the example, three microphones 105_1, 105_2, 105_3 are used. Other numbers of microphones 105 can be used in other examples.
[0063] The microphones 105_1, 105_2, 105_3 are configured to capture multiple audio signals. Each of the individual audio signals can include a speech component (s x (t), x = 1, 2, 3) and a noise component (n x (t), x = 1, 2, 3). The speech component s x(t) can originate from the person 203 who is speaking. The person 203 can participate in a communication session such as a conference call. More than one person can also participate in the communication session. The noise component n x (t) can include any unwanted sound. In Figure 2 the example, the unwanted sounds include traffic noise 205, music 207, crosstalk from other people 209. Other types of wanted and unwanted sounds can be used in other examples.
[0064] The corresponding microphone signal s x (t) + n x (t) is provided as an input to the corresponding short-time Fourier transform (STFT) transforms 211_1, 211_2, 211_3. The STFT transforms 211_1, 211_2, 211_3 are configured to transform the microphone signal s x (t) + n x (t) to the STFT domain. The STFT transforms 211_1, 211_2, 211_3 provide the transformed microphone signal E x (f) as an output. Other suitable transforms or filter banks can be used to convert the signal into a frequency-domain representation.
[0065] The transformed microphone signal E x (f) is provided as an input to the first audio signal processing. In this example, the first audio signal processing includes beamforming processing 213. The beamforming processing 213 is configured to combine the transformed microphone signal E x (f) into a single signal. The beamforming processing 213 can include a minimum variance distortionless response (MVDR) beamformer or any other suitable type of beamformer.
[0066] The beamforming processing 213 provides a single beamformed signal E(f) as an output. The single output E(f) of the beamforming processing 213 or other first audio signal processing is based on multiple microphone inputs.
[0067] The output E(f) of the first audio signal processing is provided as an input to the second audio signal processing. The second audio signal processing can include spectral noise suppression processing 215 or any other appropriate type of audio signal processing.
[0068] The spectral noise suppression process 215 includes an element-wise multiplication of the beamformed signal E(f) with a frequency-dependent mask M(f). If the beamformed signal E(f) mainly includes speech (or other desired sounds), the mask M(f) in a particular frequency bin has a value of 1. If the beamformed signal E(f) mainly includes noise (or other unwanted sounds), the mask M(f) in a certain frequency bin has a value of 0. If the beamformed signal E(f) includes both noise and speech in a frequency bin, a mask value between 0 and 1 is used.
[0069] The spectral noise suppression process 215 provides a noise-reduced signal as an output. The noise-reduced signal is given by the following formula:
[0070]
[0071] The noise-reduced signal is provided as an input to the inverse STFT 217. The inverse STFT 217 is configured to convert the noise-reduced signal back to the time domain. The inverse STFT 217 provides a time-domain noise-reduced signal as an output.
[0072] Figure 2 The example DSP chain 201 enables signals from multiple microphones 105 to be combined by the beamforming process 213 in order to suppress noise with minimal speech distortion. Using the spectral noise suppression process 215 in addition to the beamforming process 213 can provide even better noise suppression.
[0073] The mask M(f) used in the spectral noise suppression process 215 can be generated using any suitable component such as a machine learning program. The mask can also be used in the beamforming process 213. This may require generating multiple masks for a single DSP chain 201. This may require a complex machine learning program.
[0074] The examples of the present disclosure address these issues and provide an efficient method for implementing a computer program such as a machine learning program in an audio DSP chain, which can be used for noise suppression processing or other types of processing.
[0075] Figure 3 An example method that can be implemented in the examples of the present disclosure is shown. The method can be implemented by an apparatus 1301 as shown in Figure 13 and the apparatus 1301 can be provided in, as shown in Figure 1in the client device 103 shown or in any other suitable type of device. The apparatus 1301 or the client device 103 can be used for audio communication setup. The audio communication setup can be a one-way communication setup or a two-way communication setup. The client device 103 can be an electronic device such as a telephone, a camera, a computing device, a teleconferencing device, a television, a virtual reality device, an augmented reality device, and / or any other suitable type of device.
[0076] The method includes, at block 301, obtaining at least one audio signal for the current frame or one or more previous frames. The at least one obtained audio signal is based on at least two microphone signals for the current frame or one or more previous frames. The audio signal can be based on two or more microphone signals and some processing that can be performed on the two or more microphone signals to provide at least one audio signal. The at least two microphone signals can be combined such that they are provided as a single input. Beamforming processing 213 or any other suitable type of processing can be used to combine the microphone signals.
[0077] In some examples, the at least one audio signal includes the output of a first audio signal processing. The first audio signal processing can include beamforming processing 213 or any other suitable type of processing.
[0078] At block 303, the method includes using program code to predict an output signal for future frames. The predicted output signal is at least partially based on the at least one audio signal for the current frame or one or more previous frames. Future frames are frames that occur later than the frames for which their audio signals have been obtained.
[0079] The program code can include a machine learning program such as a neural network circuit. The neural network circuit can be a deep neural network (DNN) or can have any other suitable type of architecture.
[0080] At block 305, the method includes using the output signal to process future frames of the at least two microphone signals in a first audio signal processing and also using the output signal to process future frames of the output of the first audio signal processing in a second audio signal processing. The processing of the future frames of the output of the first audio signal processing in the second audio signal processing enables noise suppression. The same output signal is applied to future frames of multiple audio signal processings.
[0081] The first audio signal processing and the second audio signal processing can be consecutive processings, where the output of the first audio signal processing is provided as the input of the second audio signal processing. The first audio signal processing and the second audio signal processing can be part of the DSP chain 201 as Figure 4 shown, or can be any other suitable configuration.
[0082] The first audio signal processing may include beamforming processing 213. The beamforming processing 213 may be configured to combine multiple microphone signals into a combined microphone signal. The beamforming processing 213 may include processing future frames of at least two microphone signals using an output signal from program code. In such an example, the output signal may include a gain to be applied to at least one of the at least two microphone signals for the beamforming processing 213. The gain may be a multiplier applied to two or more microphone signals. In some examples, the output signal may include an amplitude of at least one of the at least two microphone signals to be used for the beamforming processing 213. For example, the amplitude or the product with the gain may be used to calculate the complex power spectral density of each microphone signal. The amplitudes or the products with the gain of all microphones may be used to calculate the complex power spectral density matrix of the noise, speech, or interference signal. These matrices may be used to obtain a good configuration for the beamforming processing.
[0083] The second audio signal processing may include spectral noise suppression processing 215. The spectral noise suppression processing 215 may be configured to further reduce the noise in the beamformed microphone signal. In such an example, the output signal from the program code may include a gain to be applied to the input for the spectral noise suppression processing 215. The gain may include a multiplier that may be applied to the beamformed microphone signal to obtain a denoised version of the beamformed microphone signal. In some examples, the output signal of the program code may include an amplitude to be used for the spectral noise suppression processing. The amplitude may correspond to the amplitude of the denoised signal, and a prediction of the phase is added to the denoised signal to obtain the denoised signal. The prediction of the phase may be the phase of the input signal for the spectral noise suppression processing 215.
[0084] In some examples, the output signal may be applied to future frames of the respective audio signal processing in the frequency domain. Any suitable transform may be used to transform the microphone signal to the frequency domain and transform the processed signal back to the time domain. The output signal may be provided in a format such that the output signal from the computer program can be used for the respective audio signal processing in the frequency domain.
[0085] In some examples, the program code receives a single input and provides a single output. This may reduce the complexity of the program code. The single input may include any number of features. The single input is single because it is derived from a single audio signal (e.g., the beamformed microphone signal). The single output of the program code may be provided in a format such that it can be applied to the signals in the DSP chain 210. The single output may be applied to multiple audio signal processing.
[0086] The number of current frames or previous frames in the acquired audio signal for predicting the output signal for future frames may be selected at least partially based on the delay requirement.
[0087] Figure 4 Schematically shows a DSP chain 201 that can be used in the examples of the present disclosure. In this example, program code 401 is used to predict the output signals of the first audio signal processing and the second audio signal processing. The program code 401 uses the input signals based on one or more current frames or previous frames to predict the output signals. Then, the output signals can be applied to future frames of the respective audio signal processing. The example DSP 201 can be used in communication setups and / or any other suitable type of setup.
[0088] In Figure 4 the example of, the DSP chain 201 includes three microphones 105_1, 105_2, 105_3. Other numbers of microphones 105 can be used in other examples.
[0089] The microphones 105_1, 105_2, 105_3 are configured to capture multiple copies of the audio signal. Each of the respective copies of the audio signal includes a speech component (s x (t), x = 1, 2, 3) and a noise component (n x (t), x = 1, 2, 3). The speech component (s x (t), x = 1, 2, 3) can include the desired sounds to be retained in the processed audio signal, such as one or more people speaking. The noise component (n x (t), x = 1, 2, 3) can include the unwanted sounds to be suppressed in the processed audio signal, such as crosstalk noise or traffic noise or non-target people.
[0090] The corresponding microphone signals s x (t) + n x (t) are provided as inputs to respective short-time Fourier transform (STFT) transforms 211_1, 211_2, 211_3. The STFT transforms 211_1, 211_2, 211_3 are configured to transform the microphone signals s x (t) + n x (t) into the STFT domain. The STFT transforms 211_1, 211_2, 211_3 provide the transformed microphone signals E1(f, τ), E2(f, τ), E3(f, τ) as outputs, where τ indicates the time frame and f indicates the frequency bin. As an alternative to the STFR, other suitable transforms or filter banks can also be used to convert the signal into a frequency domain representation.
[0091] The transformed multiple microphone signals E1(f, τ), E2(f, τ), E3(f, τ) are provided as inputs to a first audio signal processing. In this example, the first audio signal processing includes a beamforming process 213. In other examples, other types of audio signal processing may be used. The beamforming process 213 is configured to combine the transformed microphone signals E1(f, τ), E2(f, τ), E3(f, τ) into a single beamformed signal E(f, τ). The beamforming process 213 may include a minimum variance distortionless response (MVDR) beamformer, a multi-channel Wiener filter, or any other suitable type of beamformer.
[0092] The beamformed signal E(f, τ) is provided as an input to a second audio signal processing. The beamformed signal E(f, τ) is also provided as an input to a program code 401.
[0093] The program code 401 receives the beamformed signal E(f, τ) for the current frame and / or one or more previous frames. The program code 401 receives a single audio signal as an input. The single audio signal includes information derived from the single audio signal. In this example, the single audio signal is the beamformed signal E(f, τ).
[0094] In this example, the program code 401 includes a machine learning program. The machine learning program may be a deep neural network (DNN) or any other suitable type of program code 401. Examples of DNNs are shown in Figure 8 and 9 are shown.
[0095] The program code 401 is configured to predict the output signals of both the first audio signal processing and the second audio signal processing for future frames. The function of the program code 401 is to calculate the output signals to be applied during a future frame τ + T based on a given input (e.g., E(f, τ)) for a frequency bin f and the current frame τ. In an example where the program code 401 includes a DNN, this can be expressed as: M τ (f, τ - T) = DNN E(f, τ - T) or M τ+T (f, τ) = DNN E(f, τ).
[0096] The output signal M τ (f, τ - T) indicates the output signal predicted by the machine learning program or DNN based on the current frame τ - T or previous frames for a future frame τ. Thus, in Figure 4 at frame τ, the output signal M τ (f, t - T) of the DNN is used, where the output signal M τ(f, t-T) is predicted at least in part based on the output of the first audio signal processing. In this case, the output signal M for a future frame (τ relative to τ-T) is predicted at least in part based on the beamformed signal E(f, τ-T). The output signal M τ (f, τ-T) is predicted based on the current frame τ-T and / or one or more previous frames of the output of the first audio signal processing for the future frame τ. τ (f, τ-T) for the future frame τ.
[0097] The points indicated in the output of the computer program 401 represent the fact that the computer program 401 receives an input signal (e.g., E(f, τ)) for a particular current (and previous frames). The computer program 401 can then start to calculate the corresponding output signal that only has to be ready at a certain future frame τ+T. However, during the current frame τ, the output signal M τ (f, τ-T) calculated for the current frame based on past frames is applied.
[0098] In Figure 4 the example, the output signal M τ (f, τ-T) can include a mask applied to multiple signals. The output signal M τ (f, τ-T) can be used as a microphone-related, frequency-related, or time-related mask for the first audio signal processing and the second audio signal processing. The same output signal M τ (f, τ-T) is provided to the first audio signal processing and the second audio signal processing.
[0099] The output signal M of the program code 401 τ (f, τ-T) can be fed back to the first audio processing. In Figure 4 the example, the output signal M τ (f, τ-T) can be fed back to the beamforming processing 213. The output signal M τ (f, τ-T) can include the gain for future frames of the input microphone signals E1(f, τ), E2(f, τ), E3(f, τ) to be applied to the beamforming processing 213. The output signal M τ (f, τ-T) can include the amplitude of future frames of at least one of the input microphone signals E1(f, τ), E2(f, τ), E3(f, τ) to be used for the beamforming processing 213.
[0100] The output signal M of the program code 401 τ (f, τ-T) is also provided as an input to the second audio signal processing. In this example, the second audio signal processing includes the spectral noise suppression processing 215. Other types of audio signal processing can be used in other examples. The output signal M τ(f, τ-T) can include the gain for future frames of the beamformed signal E(f, τ) to be applied in the spectral noise suppression process 215. The output signal M τ (f, τ-T) can include the amplitude for future frames of at least one beamformed signal E(f, τ) in the spectral noise suppression process 215. The product of the amplitude or with the gain can be used to calculate the complex power spectral density of each microphone signal. The multiplication of the amplitude or with the gain for all microphones can be used to calculate the complex power spectral density matrix of the noise, speech, or interference signal. These matrices can be used to obtain a good configuration for the beamforming process.
[0101] The spectral noise suppression process 215 can use the output signal M from the program code 401 τ (f, τ-T) to control the reduction of noise in the beamformed signal E(f, τ). The spectral noise suppression process 215 provides a noise-reduced signal as output. The noise reduction signal can be obtained by multiplying the beamformed signal E(f, τ) and the output signal M τ (f, τ-T) such that:
[0102] The noise-reduced signal is converted back to the time domain by the inverse STFT 217. The inverse STFT 217 provides a time-domain noise-reduced signal as output.
[0103] In the Figure 4 example, the notation M τ (f, τ-T) represents the fact that the program code 401 predicts the output signal for time frame τ based on the input available up to time frame τ-T. The advantage of this is that the program code 401 has significantly relaxed delay requirements for calculating the output signal. Figure 5 Illustrates how this relaxes the delay requirements.
[0104] In examples such as Figure 4 , the program code 401 can be a single-channel noise suppression DNN that predicts the STFT mask of a single input STFT signal. This allows the DNN or other types of program code 401 to be smaller and optimized for single-channel noise suppression. This can be trained more efficiently compared to multi-channel noise suppression DNNs. Single-channel noise suppression based on DNN can achieve good noise suppression performance. Single-channel noise suppression based on DNN can outperform traditional signal processing techniques. The computational complexity of the single-channel noise suppression DNN is much lower compared to that of the multi-channel noise suppression DNN. The weights of a single noise suppression DNN can be trained (using some loss function) to follow the following relationship:
[0105] ο and
[0106] where S(f) represents the target speech signal, represents the predicted speech signal, M(f) represents the spectral mask output of the DNN, E(f) represents the input signal of the DNN, Z(f) represents the target noise signal, and represents the predicted noise signal.
[0107] These relationships imply that the mask output of the DNN is such that when applied to the microphone STFT signal, it predicts speech, and when the difference between 1 and the mask is applied to the microphone STFT signal, it predicts noise. Some DNN training loss functions for achieving this are by considering the following loss functions L1, L2, L3, or L4, where s(t) represents the target speech in the time domain: or L2 = (|M(f)E(f)| - |S|) 2 or L3 = (|M(f)E(f)| 0.3 - |S| 0.3 ) 2 or L4 = (ISTFT(M(f)E(f)) - s(t)) 2
[0108] In an example of the present disclosure, the output of program code 401 provides a single output signal that can be used for multiple audio signal processing. In Figure 4 's example, the predicted output signal M τ (f, τ - T) is used as the mask of the spectral post-filter in the spectral noise reduction process 215 and also as the microphone, frequency, and time-dependent mask of the beamforming process 213. For multiple audio signal processing, reusing the single output signal M τ (f, τ - T) enables the use of a very low-complexity DNN or other types of program code 401.
[0109] In Figure 4 's example, during frame τ, the beamforming process 213 needs to calculate the output E(f, τ) from the inputs E1(f, τ), E2(f, τ), E3(f, τ) based on the output signal M τ (f, τ - T). The program code 401 can pre-compute the output signal M τ (f, τ - T) based on the inputs available at frame τ - T, where T refers to the number of predicted frames. This gives the program code 401 approximately T time frames to compute the output signal M τ (f, τ - T). The examples of the present disclosure are not limited to single-frame-ahead prediction. For example, in some examples, four-frame-ahead prediction can be used. In such an example, the output signal M τ(f, τ-4), and the program code 401 can use information derived from frame τ-4 and earlier frames to predict the output signal of frame τ. This relaxes the latency requirement calculated by the program code 401 to four frames, which may be approximately 40 ms.
[0110] Figure 5 Illustrates how using different numbers of frames in frame-advance prediction can affect the latency requirement and the complexity of the program code 401. Figure 5 Illustrates example calculation pipelines for different methods. The first pipeline 501 is shown for an implementation where there is no frame-advance prediction. In this pipeline, the output of the program code 401 is predicted for the current frame.
[0111] Other pipelines 503, 505, 507 are for methods using examples of the present disclosure. The second pipeline 503 uses 1-frame advance prediction, the third pipeline 505 uses 2-frame advance prediction, and the fourth pipeline 507 uses 3-frame advance prediction.
[0112] In Figure 5 the example, for each frame, as indicated by the arrows, the corresponding signals are transmitted to and from the program code 401. The signals transmitted to the program code 401 can include the STFT signal. The signals transmitted to the program code 401 can be based on the microphone signal. Before transmitting the microphone signal to the program code 401, some processing can be performed on the microphone signal, and the signals transmitted to the program code 401 can include the beamformed signal E(f, τ) or any other suitable type of signal.
[0113] The signals transmitted from the program code 401 can also include the STFT signal or its compressed version (e.g., Mel-frequency cepstral coefficients, equivalent rectangular bandwidth (ERB) coefficients, or Bark scale coefficients). The signals transmitted from the program code 401 can include the output signal M τ (f, τ-T). The output signal M τ (f, τ-T) can include the STFT mask or its compressed version (e.g., in ERB or Bark scale).
[0114] The program code 401 has a time interval to perform the inference to obtain the predicted output signal M τ (f, τ-T). The time interval depends on the number of future frame predictions and also on the time required to transmit the input signal to and the output signal from the program code, respectively. The time interval available for the program code 401 to perform the inference is represented by Figure 5 the horizontal width of the boxes in the respective pipelines in Figure 5is represented by the horizontal length of the arrow 515 in each pipeline.
[0115] The first pipeline 501 corresponds to an example where no future frame prediction is used. In this example, the program code 401 for predicting a mask or other output signal includes multiple single-channel DNNs, each of which processes one of the microphone signals in parallel. The multiple single-channel DNNs are represented by Figure 5 the dots in, and in this pipeline 501, the program code 401 has much less time than a full time frame to compute the output signal.
[0116] The second pipeline 503 shows an example of the present invention, where the program code 401 includes a single-channel DNN that uses 1-frame-ahead prediction. The microphone signals are combined into a single audio signal and provided as an audio input signal to the program code 401. For example, the microphone signals can be combined by beamforming processing 213. Only one single-channel DNN is used in the second pipeline 503 because only one audio signal input is used. This reduces the complexity of the program code 401 compared to the program code used in the first pipeline 501.
[0117] In the second pipeline 503, the input at frame t-3 is transmitted to the only single-channel DNN to infer the mask at frame t-2, and the program code 401 has approximately a full time frame to compute the output signal. The delay indicated by the horizontal length of the arrow 515 is allowed to be greater in the second pipeline 503 because there is a full time frame available for transmitting the input and output signals and for DNN inference, and thus the delay is relaxed.
[0118] The single-channel DNN used in the second pipeline 503 can have the same architecture as the single-channel DNN used in the first pipeline 501, however different weights will be used. The corresponding single-channel DNNs will have the same number of computations. The computations can be multiply-accumulate (MAC) computations, non-linear activation functions, or any other suitable type of computation.
[0119] Figure 5 It is shown that compared to the first pipeline 501 that does not use any future frame prediction, the program code 401 in the second pipeline 503 has considerably relaxed delay and computational complexity requirements. This is indicated by the significant increase in the width of the boxes in the second pipeline 503 compared to the first pipeline 501. The delay indicated by the horizontal length of the arrow 515 is allowed to be greater in the third pipeline 505 because there are two full time frames available for transmitting the input and output signals and for DNN inference. Transmitting the same amount of input and output data over a longer time, thus this provides significantly relaxed delay requirements.
[0120] The third pipeline 505 shows an example of the present disclosure, where the program code 401 includes a single-channel DNN using 2-frame-ahead prediction. In this example, the input at frame t-3 is transmitted to the program code 401 to infer the output signal for frame t-1, and the program code 401 has approximately two complete time frames to calculate the output signal. In the third pipeline 505, the program code 401 includes two single-channel DNNs deployed in parallel. The first single-channel DNN is indicated by the first row of boxes 509, and the second single-channel DNN is indicated by the second row of boxes 511. Compared with the program code 401 used in the second pipeline 503, using two single-channel DNNs in parallel increases the complexity of the program code 401. However, compared with the time available in the second pipeline 503, the time available for performing the calculations has approximately doubled. The delay indicated by the horizontal length of the arrow 515 is allowed to be even greater in the fourth pipeline 507 because there are three complete time frames available for transmitting the input and output signals and for DNN inference. The same amount of input and output data is transmitted over a longer period of time, and thus this provides even more relaxed delay requirements.
[0121] The fourth pipeline 507 shows another example of the present disclosure, where the program code 401 includes a single-channel DNN using 3-frame-ahead prediction. In this example, the input at frame t-3 is transmitted to the program code 401 to infer the mask for frame t. The program code 401 has approximately three complete time frames to calculate the output signal. For this version, the program code 401 includes three single-channel DNNs deployed in parallel. The first single-channel DNN is indicated by the first row of boxes 509, the second single-channel DNN is indicated by the second row of boxes 511, and the third single-channel DNN is indicated by the third row of boxes 513. Compared with the program code 401 used in the second pipeline 503 and the third pipeline 505, using three single-channel DNNs in parallel increases the complexity of the program code 401. However, compared with the time available in the second pipeline 503, the time available for performing the calculations has approximately tripled. The delay indicated by the horizontal length of the arrow 515 is allowed to be greater than the delays in the first, second, and third pipelines 501, 503, 505.
[0122] In an example of the present disclosure, the program code 401 can be configured to predict any suitable number of ahead frames. When determining the number of previous frames for prediction, the performance of audio signal processing is considered. The earlier the predicted output, the more difficult it is to predict the output signal. The size of the frame is also considered. If the frame has a smaller size, a larger frame-ahead value can be used. Increasing the frame-ahead value provides more relaxed delay requirements because this gives more time before the output signal needs to be ready. Thus, determining the frame-ahead value is a trade-off between delay and prediction accuracy.
[0123] Figure 6System 601 is schematically shown that can be used to implement some examples of the present disclosure. Figure 6 The illustrated system 601 can be implemented in a user equipment 103 as shown in Figure 1 and / or can be implemented in any other suitable type of device or combination of devices.
[0124] System 601 includes a plurality of microphones 105. Three microphones are shown in Figure 6 . Other numbers of microphones 105 can be used in other examples.
[0125] The microphones 105 provide a plurality of microphone input signals 603. Each of the plurality of microphones 105 provides a corresponding microphone input signal 603. In the example of Figure 6 , there are three microphones 105, and thus three microphone input signals 603 are provided. Other numbers of microphone input signals 603 can be used in other examples.
[0126] The microphone input signals 603 are provided as inputs to a central processing unit (CPU) 605 or a DSP system. The CPU 605 is configured to implement a first audio signal processing and a second audio signal processing. The first audio signal processing and the second audio signal processing can be performed on two or more microphone input signals 603. The first audio signal processing and the second audio signal processing can be consecutive processes so as to provide the output of the first audio signal processing as the input of the second audio signal processing.
[0127] The first audio signal processing can include beamforming processing 213, and the second audio signal processing can include spectral noise reduction processing 215 or any other suitable type of processing. The beamforming processing 213 and the spectral noise reduction processing 215 can be as shown in Figure 4 . Other types of audio signal processing can be used in other examples.
[0128] Each audio signal processing performed by the CPU 605 can be based on a mask or other output signals output from the program code 401.
[0129] The mask or output signals for audio signal processing to be performed by the CPU 605 are obtained from the DNN engine 609. The DNN engine includes program code 401, which is configured to predict output signals for future frames based on audio signals for the current frame or one or more previous frames.
[0130] The CPU 605 provides an audio signal 607 to the DNN engine 609. The audio signal 607 is based on at least two microphone signals 603. The audio signal 607 may include a combined microphone signal, or features extracted from the combined microphone signal, and / or any other suitable input. The combined microphone signal can be obtained by performing beamforming or any other suitable processing on two or more microphone signals 603.
[0131] The audio signal 607 provided to the DNN engine 609 includes the current frame and / or one or more previous frames. A single audio signal 607 can be provided from the CPU to the DNN engine 609.
[0132] The DNN engine 609 may include software routines in a DSP or a graphics processing unit (GPU), or may be a hardware accelerator or any other suitable device.
[0133] The DNN engine 609 includes program code 401 configured to process the audio signal 607 to generate an output signal 611. The program code 401 may include one or more single-channel DNNs or any other suitable type of program code 401. The number of single-channel DNNs included in the program code 401 can be determined by the number of frames predicted before it and / or any other suitable factors. For example, if one-frame-ahead prediction is used, the program code 401 may include only one single-channel DNN. If two-frame-ahead prediction is used, the program code 401 may include two single-channel DNNs.
[0134] The output signal 611 provided by the DNN engine 609 may include masks that can be used for respective audio signal processing by the CPU 605. The masks are predicted based on previous frames, but can be used for the current frame. The same output signal can be used for future frames of both the first audio signal processing and the second audio signal processing. The output signal 611 may include masks, gains, amplitudes, and / or any other information suitable for the first audio signal processing and the second audio signal processing.
[0135] The CPU 605 uses the output signal 611 in the first audio signal processing and the second audio signal processing. The CPU 605 provides a processed signal 613 as an output. In this example, the processed signal 613 is a noise suppression signal. Other types of processing may be performed in other examples.
[0136] Figure 7 An example method that can be used in some examples of the present disclosure is shown.
[0137] In block 701, the method includes obtaining the current frame of an audio signal or one or more previous frames. The audio signal can be based on microphone signals. The audio signal can include a combined microphone signal.
[0138] The acquired audio signal including the current frame or one or more previous frames may include the output of a first audio signal processing. The first audio signal processing may be performed on two or more input microphone signals. The first audio signal processing may include beamforming processing 213 or any other suitable processing.
[0139] The first audio signal processing may be part of the DSP chain 201. The DSP chain 201 may include at least the first audio signal processing and the second audio signal processing. The first audio signal processing and the second audio signal processing may be consecutive processes, where the output of the first audio signal processing is provided as the input of the second audio signal processing. The first audio signal processing may include beamforming processing 213, and the second audio signal processing may include spectral noise reduction processing 215. The DSP chain 201 may be as Figure 4 shown, or may be any other suitable DSP chain 201.
[0140] In block 703, features are extracted from the audio signal and provided as input to the program code 401. Features may be extracted from the audio signal to provide an input in a format suitable for the program code 401. This may include compressing the dimensionality of the audio signal by extracting Mel-frequency cepstral coefficients, equivalent rectangular bandwidth (ERB) grid coefficients, or Bark scale grid coefficients. This may also include normalizing or standardizing the audio signal by subtracting the online calculated mean or by dividing by the online calculated standard deviation. This may also include a log-power transformation of the audio signal.
[0141] In block 705, the program code 401 uses the features extracted from the audio signal to predict an output signal. The output signal may include a mask or any other suitable information, which may be used for future frames of at least two audio signal processings in the DSP chain 201. The audio signal processing may include beamforming processing 213 and spectral noise reduction processing 215 and / or any other suitable audio signal processing.
[0142] The program code 401 may include a DNN as Figure 8 shown in and / or 9 or any other suitable type of program code.
[0143] At block 707, the output signal is transmitted from the program code 401 to the DSP chain 201. At block 709, the output signal is applied to future frames of both the first audio signal processing and the second audio signal processing for future frames. In an example of the present disclosure, the program code 401 provides a single output signal. The same output signal is used for future frames of both the first audio signal processing and the second audio signal processing. The output signal can be used as a mask in each processing. The mask can be microphone-related, frequency-related, time-related, or have any other suitable configuration.
[0144] At block 711, the DSP chain 201 outputs a processed signal. In this example, the processed signal can include a noise-reduced signal. Other types of processed signals can be provided in other examples.
[0145] In some examples, the program code 401 can include a machine learning program, Figure 8 An example machine learning program that can be used in some examples of the present disclosure is shown. In this example, the machine learning program includes a deep neural network (DNN) 801. The DNN 801 includes an input layer 803, an output layer 807, and a plurality of hidden layers 805. The hidden layers 805 are provided between the input layer 803 and the output layer 807. Figure 8 The illustrated example DNN 801 includes two hidden layers 805, but in other examples, the DNN 801 can include any number of hidden layers 805.
[0146] Each layer within the DNN 801 includes a plurality of nodes 809. The nodes 809 within each layer are connected together by a plurality of connections 811 or edges, as Figure 8 shown. Each connection 811 represents a multiplication with a weight configuration. Within the nodes 809 of the hidden layers 805 and the output layer 807, a non-linear activation function is applied to obtain a multi-dimensional non-linear mapping between the input and the output.
[0147] In an example of the present disclosure, the DNN 801 is trained or configured to map a single input signal to a corresponding output signal. The input signal can include any suitable input, such as the output of the beamforming process 213 or any other input based on the microphone signal. The output signal can include a mask or any other suitable information for the beamforming process 213 and the spectral noise suppression process 215 and / or any other suitable process.
[0148] Figure 9 An example architecture of a machine learning program that can be used as the program code 401 in some examples of the present disclosure is shown.
[0149] In this example, the program code 401 includes a DNN. Figure 9The architecture in includes multiple interconnected residual gated recurrent (GRU) networks. Other architectures of program code 401 may be used in other implementations of the present disclosure.
[0150] Program code 401 receives a single input 901. The input 901 includes an audio signal. The input 901 is a single input because it includes information derived from a single signal. The single signal may be based on two or more microphone signals. For example, the single signal may be the output of beamforming processing 213 or other types of processing that combine multiple microphone signals.
[0151] The input 901 may be provided in any suitable format. For example, the input 901 may be provided in STFT format. The input 901 may be provided as the log power of STFT frames. In some examples, the input 901 may also be pre-conditioned by STFT to ERB (equivalent rectangular bandwidth) grid conversion or STFT to Bark scale grid conversion or STFT to Mel-frequency cepstral coefficient conversion to reduce the complexity of program code 401.
[0152] The input 901 is provided to a series of four consecutive gated recurrent unit (GRU) layers 903, 905, 907, 909. Each GRU layer 903, 905, 907, 909 has a residual connection 917, 919, 921, 923. For each corresponding GRU layer 903, 905, 907, 909, the input is also added to the output.
[0153] The output of each of the GRU layers 903, 905, 907, 909 is provided as the input to a linear terminal layer 911. The linear terminal layer 911 combines the corresponding inputs.
[0154] The output of the linear terminal layer 911 is provided as the input to a sigmoid (S-shaped) activation function 913 to generate the output 915 of program code 401. Program code 401 provides a single output 915. The single output 915 may be provided in any suitable format. The single output may have dimensions that enable it to be applied to signals with similar dimensions. For example, the output 915 may be in STFT format so that the output 915 can be applied to other STFT signals. The output 915 is a single signal, but it may be applied to multiple audio processes, such as beamforming processing 213 and spectral noise suppression processing 915.
[0155] In this example, the output signal 915 includes a mask that can be used for first audio signal processing and second audio signal processing. The mask may be used for future frames of respective audio signal processing.
[0156] Any suitable process can be used to train the program code 401. Example training objectives or loss functions that can be used to train such program code 401 can be as follows:
[0157]
[0158] where T refers to the number of frames that the program code 401 needs to predict in advance.
[0159] Another example training objective or loss function that can be used to train such program code 401 can be as follows:
[0160]
[0161] Figure 10 An exemplary room 1001 and a microphone array 1003 are shown, which are used together with examples of the present disclosure to obtain Figure 11 and 12 the exemplary results shown in. The room 1001 has dimensions 5m (x-axis) × 5m (y-axis) × 3m (height).
[0162] The microphone array 1003 includes a linear array of four microphones. Each microphone is spaced 10 cm apart within the linear array.
[0163] Figure 10 Four signals 1005A-D are shown. The first signal 1005A is a speech signal. The other signals 1005B-D are noise signals. In this example, the second signal 1005B includes multiplexed noise, and the third signal 1005C and the fourth signal 1005D are white noise signals.
[0164] To obtain Figure 11 and 12 the exemplary results shown, STFT domain processing with frames and hop sizes of 1024 and 512 samples with a Hann window was used respectively. When considering a sampling frequency of 48 kHz, this corresponds to a frame of approximately 20 ms. The examples of the present disclosure are applied with 1-frame advance prediction.
[0165] Using 1-frame advance prediction relaxes the latency requirement to less than 15 - 20 ms instead of less than 5 ms.
[0166] Figure 11 A graph of the exemplary results is shown on a linear scale, Figure 12 A graph of the exemplary results is shown on a dB scale. Figure 11 The first curve 1101 in Figure 12 and the first curve 1201 in Figure 11 show the captured signal on the first microphone 105. This indicates that there is a large amount of noise in the signal.Figure 12 The second curve 1203 in shows the signal after beamforming processing 213 and spectral noise suppression processing 215 using the examples of the present disclosure. These show that the noise is significantly reduced without much speech distortion.
[0167] The prediction of the output signal for future frames for each audio signal processing provides benefits for the program code 401 for making the prediction. As described herein, the prediction of future frames can relax the latency requirements and can also reduce complexity if the program code 401 is required.
[0168] Figure 13 FIG. schematically shows an apparatus 1301 that can be used to implement the examples of the present disclosure. In this example, the apparatus 1301 includes a controller 1303. The controller 1303 can be a chip or a chipset. In some examples, the controller can be provided within a user equipment 103 such as Figure 1 shown in .
[0169] In Figure 13 the example of , the implementation of the controller 1303 can be as a controller circuit. In some examples, the controller 1303 can be implemented solely in hardware, with certain aspects of software including separate firmware, or can be a combination of hardware and software (including firmware).
[0170] As Figure 13 shown, the controller 1303 can be implemented using instructions that implement hardware functions, for example, by using executable instructions of a computer program 1309 in a general-purpose or a special-purpose processor 1305, and the executable instructions can be stored on a computer-readable storage medium (disk, memory, etc.) to be executed by such a processor 1305.
[0171] The processor 1305 is configured to read from and write to the memory 1307. The processor 1305 may also include an output interface and an input interface, through which the processor 1305 outputs data and / or commands and inputs data and / or commands to the processor 1305.
[0172] The memory 1307 is configured to store a computer program 1309 including computer program instructions (computer program code 401), which, when loaded into the processor 1305, control the operation of the controller 1303. The computer program instructions of the computer program 1309 provide the logic and routines that enable the controller 1303 to execute the methods shown in the figures, and the processor 1305 can load and execute the computer program 1309 by reading the memory 1307.
[0173] Accordingly, apparatus 1301 includes: at least one processor 1305; and at least one memory 1307 including computer program code 401, the at least one memory 1307 and the computer program code 401 being configured to, with the at least one processor 1305, cause the apparatus 1301 to at least perform:
[0174] obtain 301 at least one audio signal for the current frame or one or more previous frames, based on at least two microphone signals for the current frame or one or more previous frames;
[0175] use the program code 401 to predict 303 an output signal for a future frame, at least in part based on the at least one audio signal for the current frame or one or more previous frames; and
[0176] use the output signal to process a future frame of the at least two microphone signals in a first audio signal processing, and use the output signal to process a future frame of the output of the first audio signal processing in a second audio signal processing, to achieve noise suppression.
[0177] As Figure 13 shown, the computer program 1309 may reach the controller 1303 via any suitable delivery mechanism 1311. The delivery mechanism 1311 may be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a recording medium such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD) or a solid-state memory, an article of manufacture including or tangibly embodying the computer program 1309. The delivery mechanism may be a signal configured to reliably transmit the computer program 1309. The controller 1303 may propagate or transmit the computer program 1309 as a computer data signal. In some examples, the computer program 1309 may be transmitted to the controller 1303 using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (IPv6 over Low-Power Personal Area Network), ZigBee, ANT+, Near Field Communication (NFC), Radio Frequency Identification, Wireless Local Area Network (Wireless LAN) or any other suitable protocol.
[0178] The computer program 1309 includes computer program instructions that, when executed by the apparatus 1301, cause the apparatus 1301 to at least perform the following operations:
[0179] obtain 301 at least one audio signal for the current frame or one or more previous frames, based on at least two microphone signals for the current frame or one or more previous frames;
[0180] Use program code 401 to predict 303 an output signal for a future frame based at least in part on at least one audio signal for a current frame or one or more previous frames; and
[0181] Use the 305 output signal to process future frames of at least two microphone signals in a first audio signal processing, and use the output signal to process future frames of the output of the first audio signal processing in a second audio signal processing to achieve noise suppression.
[0182] Computer program instructions may be included in a computer program 1309, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program 1309.
[0183] Although the memory 1307 is shown as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / removed and / or may provide permanent / semi-permanent / dynamic / cache storage.
[0184] Although the processor 1305 is illustrated as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / removable. The processor 1305 may be a single-core or multi-core processor.
[0185] References to "computer-readable storage medium", "computer program product", "tangibly embodied computer program", etc. or "controller", "computer", "processor", etc. should be understood to cover not only computers having different architectures such as single / multi-processor architectures and sequential (von Neumann) / parallel architectures, but also dedicated circuits such as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc. should be understood to cover software or firmware for programmable processors, such as programmable content of hardware devices, whether instructions for a processor or configuration settings for fixed function devices, gate arrays, or programmable logic devices, etc.
[0186] As used in this application, the term "circuit" may refer to one or more or all of the following:
[0187] (a) only hardware circuit implementations (e.g., only in analog and / or digital circuits) and
[0188] (b) combinations of hardware circuits and software, such as (where applicable):
[0189] (i) combinations of analog and / or digital hardware circuits and software / firmware, and
[0190] (ii) Any part of a hardware processor (including a digital signal processor), software, and memory with software, which work together to enable a device such as a mobile phone or a server to perform various functions, and
[0191] (c) A hardware circuit and / or a processor (such as a microprocessor or a part of a microprocessor) that requires software (such as firmware) for operation, but when software is not required for operation, the software does not need to be present.
[0192] This definition of a circuit applies to all uses of the term in this application, including in any claim. As a further example, as used in this application, the term circuit also covers an implementation of only a hardware circuit or a processor and its (or their) accompanying software and / or firmware. The term circuit also covers, for example and if applicable to a particular claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network devices.
[0193] As Figure 13 The device 1301 shown as such can be provided within any suitable device. In some examples, the device 1301 can be provided within an electronic device, such as a mobile phone, a teleconferencing device, a camera, a computing device, or any other suitable device.
[0194] The boxes shown in the figure can represent steps in a method and / or code portions in a computer program 1309. The specification of a particular order of the boxes does not necessarily imply the existence of a required or preferred order of the boxes, and the order and arrangement of the boxes can be changed. Additionally, some boxes can be omitted.
[0195] The term "comprising" is used in this document in an inclusive rather than an exclusive sense. That is, any reference to X comprising Y means that X can include only one Y or can include more than one Y. If an exclusive meaning of "comprising" is intended, it is made explicit in the context by referring to "including only one..." or by using "consisting of".
[0196] In this specification, the phrases "connected", "coupled", and "communicating" and their derivatives mean operatively connected / coupled / communicating. It should be understood that any number of intermediate components or combinations of intermediate components (including no intermediate components) can exist, i.e., so as to provide a direct or indirect connection / coupling / communication. Any such intermediate component can include hardware and / or software components.
[0197] As used herein, the term "determine" (and grammatical variations thereof) can include, but is not limited to: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, etc. Moreover, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), obtaining, etc. Moreover, "determine" can include parsing, selecting, picking, establishing, etc.
[0198] In this specification, various examples have been referred to. The description of a feature or function related to an example indicates that those features or functions exist in that example. The use of the terms "example" or "for example" or "may" or "might" in the text indicates that, whether or not explicitly stated, such features or functions exist in at least the described example, whether or not described as an example, and they may but do not necessarily exist in some or all other examples. Thus, "example", "for example", "may" or "might" refer to a particular instance within a class of examples. The attributes of an instance can be attributes of only that instance or attributes of the class or of a subclass of the class that includes some but not all instances of the class. Thus, it is implicitly disclosed that features described with reference to one example rather than another can, where possible, be used as part of a working combination in that other example, but do not necessarily have to be used in that other example.
[0199] Although examples have been described in the foregoing paragraphs with reference to various examples, it should be understood that the given examples can be modified without departing from the scope of the claims.
[0200] The features described in the above description can be used in combinations different from those explicitly described above.
[0201] Although functions have been described with reference to certain features, those functions can be performed by other features whether or not described.
[0202] Although features have been described with reference to certain examples, those features can also exist in other examples whether or not described.
[0203] The terms "a", "an", or "the" are used in this document in an inclusive rather than an exclusive sense. That is, any reference to X including a / an / the Y means that X can include only one Y or can include more than one Y, unless the context clearly indicates the contrary. If it is intended to use "a", "an", or "the" with an exclusive meaning, then it will be clearly stated in the context. In some cases, the use of "at least one" or "one or more" can be used to emphasize the inclusive meaning, but the absence of these terms should not be taken as inferring any exclusive meaning.
[0204] The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and also to features (equivalent features) that achieve substantially the same technical effect. Equivalent features include, for example, features that are variants and that achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function in substantially the same way to achieve substantially the same result.
[0205] In this specification, adjectives or adjective phrases have been used to describe the features of the various examples for reference to those examples. Such a description of a feature related to an example indicates that the feature exists in some examples that are exactly the same as those described and in other examples that are substantially the same as those described.
[0206] Some examples of the present disclosure have been described above. However, those of ordinary skill in the art will recognize possible alternative structural and method features that provide functionality equivalent to that of the specific examples of such structures and features described above, and for the sake of brevity and clarity, have been omitted from the above description. Nevertheless, the above description should be understood to implicitly include a reference to such alternative structural and method features that provide equivalent functionality, unless such alternative structural or method features are expressly excluded in the above description of the examples of the present disclosure.
[0207] Although efforts have been made in the foregoing specification to focus on those features that are considered important, it should be understood that the applicant may seek protection by claims for any patentable feature or combination of features mentioned above and / or shown in the drawings, whether or not they have been emphasized.
Claims
1. A device for noise suppression, comprising: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least another processor, cause the apparatus to at least: acquiring at least one audio signal for a current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames; using program code to predict an output signal for a future frame based at least in part on the at least one audio signal for the current frame or one or more previous frames; as well as The output signal is used in a first audio signal processing to process the future frames of the at least two microphone signals, and the output signal is used in a second audio signal processing to process the future frames of the output of the first audio signal processing to achieve noise suppression.
2. The device according to claim 1, wherein: The at least one audio signal comprises an output of the first audio signal processing.
3. The device according to claim 1, wherein: The first audio signal processing and the second audio signal processing are consecutive processes, wherein an output of the first audio signal processing is provided as an input of the second audio signal processing.
4. The device according to claim 1, wherein: The first audio signal processing comprises a beamforming process, and wherein the beamforming process comprises processing the future frames of the at least two microphone signals using the output signal.
5. The device according to claim 4, wherein: The output signal comprises a gain of at least one of the at least two microphone signals to be applied to the beamforming process.
6. The device according to claim 4, wherein: The output signal comprises an amplitude of at least one of the at least two microphone signals to be used for the beamforming process.
7. The device according to claim 1, wherein: The second audio signal processing includes spectral noise suppression processing.
8. The device according to claim 7, wherein: The output signal comprises a gain to be applied to the input of the spectral noise suppression process.
9. The device according to claim 7, wherein: The output signal comprises an amplitude to be used for the spectral noise suppression process.
10. The device according to claim 1, wherein: The output signal is applied to the future frames of audio signal processing in the frequency domain.
11. The device according to claim 1, wherein: The program code receives a single input and provides a single output.
12. The device according to claim 1, wherein: The program code includes a machine learning program.
13. The device according to claim 12, wherein: The machine learning program includes a neural network circuit.
14. The device according to claim 1, wherein: The same output signal is applied to multiple future frames of audio signal processing.
15. The device according to claim 1, wherein: Based at least in part on the delay requirement, a number of current or previous frames in the acquired audio signal to use for predicting the output signal for a future frame is selected.
16. The device according to claim 1, wherein The apparatus is for use in an audio communication setting.
17. The device according to claim 16, wherein: The audio communication setting is at least one of the following; One-way communication setup; or Two-way communication setup.
18. The device according to claim 1, wherein: The apparatus is at least one of: a phone, a camera, a computing device, a teleconferencing device, a television, a virtual reality device, and an augmented reality device.
19. A method for noise suppression, comprising: acquiring at least one audio signal for a current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames; using program code to predict an output signal for a future frame based at least in part on the at least one audio signal for the current frame or one or more previous frames; as well as The output signal is used in a first audio signal processing to process the future frames of the at least two microphone signals, and the output signal is used in a second audio signal processing to process the future frames of the output of the first audio signal processing to achieve noise suppression.
20. A computer program for noise suppression comprising instructions which, when executed by an apparatus, cause the apparatus to at least: acquiring at least one audio signal for a current frame or one or more previous frames based on at least two microphone signals for the current frame or one or more previous frames; using program code to predict an output signal for a future frame based at least in part on the at least one audio signal for the current frame or one or more previous frames; as well as The output signal is used in a first audio signal processing to process the future frames of the at least two microphone signals, and the output signal is used in a second audio signal processing to process the future frames of the output of the first audio signal processing to achieve noise suppression.