Microphone signal processing

By combining AEC and ambient noise suppression processing in microphone signal processing, the noise suppression level is gradually adjusted, which solves the problem of insufficient echo and noise suppression in the separation equipment and improves the listening quality.

CN121641052APending Publication Date: 2026-03-10NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In existing microphone signal processing, acoustic echo cancellation (AEC) is ineffective in separate audio acquisition and output devices, resulting in insufficient suppression of echoes and ambient noise, and producing audible artifacts.

Method used

When AEC processing is enabled, it is combined with environmental noise suppression processing. By gradually increasing and decreasing the noise suppression level, the signal characteristics are smoothed and the effects of pseudo-sound are reduced.

Benefits of technology

It effectively reduces the adverse effects of echo and environmental noise, providing a more natural listening experience and avoiding sudden fluctuations in falsetto.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641052A_ABST
    Figure CN121641052A_ABST
Patent Text Reader

Abstract

Various example embodiments relate to microphone signal processing. For example, a method is disclosed that includes, in response to a trigger event, enabling an Acoustic Echo Cancellation (AEC) process on at least one acquired microphone signal, where the AEC process is performed for a first time period for providing at least one post-processed acquired microphone signal. The method further includes applying an ambient noise suppression process to the at least one post-processed acquired microphone signal, where a level of ambient noise suppression applied to the at least one post-processed acquired microphone signal increases from a first level to one or more second levels in response to the AEC process being enabled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various example embodiments relate to microphone signal processing. Background Technology

[0002] A user can listen to audio while a nearby microphone is enabled for voice capture. For example, in a communication session between a first user and a second user, the captured audio of the first user (far-end user) can be transmitted to the second user (near-end user) in one or more downlink signals. The received one or more downlink signals can be output as one or more first audio signals by one or more loudspeakers associated with the second user. For example, one or more speakers may include a set of headphones, etc. The second user can operate an audio capture device, such as a smartphone, which includes one or more microphones to capture their own audio for transmission back to the first user as part of the communication session. At least some of the first audio signals from the first user can be captured by one or more microphones of the audio capture device when output by one or more loudspeakers, and therefore the first user may hear an echo of their own voice and / or other feedback that may gradually worsen. Acoustic echo cancellation (AEC) processing methods can be used to eliminate or mitigate these forms of echo. Summary of the Invention

[0003] The scope of protection sought by the various embodiments of the present invention is set forth in the independent claims. Embodiments and features (if any) described in this specification that do not fall within the scope of the independent claims should be interpreted as examples that aid in understanding the various embodiments of the invention.

[0004] According to a first aspect, an apparatus is described, comprising: means for enabling acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal; and means for applying ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein in response to the AEC processing being enabled, the level of ambient noise suppression applied to the at least one post-processed acquired microphone signal increases from a first level to one or more second levels.

[0005] In some example embodiments, the components for applying ambient noise suppression processing can be disabled before AEC processing is enabled, and can be enabled in response to AEC processing being enabled.

[0006] In some example embodiments, the level of environmental noise suppression can be increased from the beginning of the first time period.

[0007] In some example embodiments, the level of ambient noise suppression can be gradually increased at a rate less than the rate at which AEC processing is applied to at least one acquired microphone signal.

[0008] In some example embodiments, the rate at which the level of ambient noise suppression increases can be on the order of seconds, and the rate at which AEC processing increases can be on the order of milliseconds.

[0009] In some example embodiments, the level of ambient noise suppression can be maintained at one or more second levels at least until the end of the first time period, and then the level of ambient noise suppression can be reduced toward the first level.

[0010] In some example embodiments, the level of ambient noise suppression can be gradually reduced at a rate that is less than the rate at which AEC processing of at least one acquired microphone signal is reduced at the end of the first time period.

[0011] In some example embodiments, the rate at which the level of environmental noise suppression decreases can be on the order of seconds, and the rate at which AEC processing decreases can be on the order of milliseconds.

[0012] In some example embodiments, AEC processing may be applied to a first frequency range of at least one acquired microphone signal, and noise suppression processing may be applied to a second frequency range of at least one post-processed acquired microphone signal, wherein the second frequency range is determined based on the first frequency range. In some example embodiments, the second frequency range may be substantially the same as the first frequency range. In some example embodiments, the second frequency range may be wider than the first frequency range and include the first frequency range.

[0013] In some example embodiments, the level of ambient noise suppression processing may be based at least in part on the lowest frequency of at least one acquired microphone signal relative to a predetermined threshold. In some example embodiments, the level of ambient noise suppression may be increased from a first level to one or more second levels only when the lowest frequency of at least one acquired microphone signal is at or below the predetermined threshold. In some example embodiments, one or more second levels may be higher than those for cases where the lowest frequency of at least one acquired microphone signal is above the predetermined threshold.

[0014] The device can be included in the user equipment.

[0015] According to a second aspect, a method is described, comprising: enabling acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal; and applying ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein in response to the AEC processing being enabled, the level of ambient noise suppression applied to the at least one post-processed acquired microphone signal increases from a first level to one or more second levels.

[0016] In some example embodiments, the application of ambient noise suppression processing can be disabled before AEC processing is enabled, and can be enabled in response to AEC processing being enabled.

[0017] In some example embodiments, the level of environmental noise suppression can be increased from the beginning of the first time period.

[0018] In some example embodiments, the level of ambient noise suppression can be gradually increased at a rate less than the rate at which AEC processing is applied to at least one acquired microphone signal.

[0019] In some example embodiments, the rate at which the level of ambient noise suppression increases can be on the order of seconds, and the rate at which AEC processing increases can be on the order of milliseconds.

[0020] In some example embodiments, the level of ambient noise suppression can be maintained at one or more second levels at least until the end of the first time period, and then the level of ambient noise suppression can be reduced toward the first level.

[0021] In some example embodiments, the level of ambient noise suppression can be gradually reduced at a rate less than the rate at which AEC processing of at least one acquired microphone signal is reduced at the end of the first time period.

[0022] In some example embodiments, the rate at which the level of environmental noise suppression decreases can be on the order of seconds, and the rate at which AEC processing decreases can be on the order of milliseconds.

[0023] In some example embodiments, AEC processing may be applied to at least a first frequency range of the acquired microphone signal, and noise suppression processing may be applied to at least a second frequency range of the post-processed acquired microphone signal, wherein the second frequency range may be determined based on the first frequency range. In some example embodiments, the second frequency range may be substantially the same as the first frequency range. In some example embodiments, the second frequency range may be wider than the first frequency range and include the first frequency range.

[0024] In some example embodiments, the level of ambient noise suppression processing may be based at least in part on the lowest frequency of at least one acquired microphone signal relative to a predetermined threshold. In some example embodiments, the level of ambient noise suppression may be increased from a first level to one or more second levels only if the lowest frequency of at least one acquired microphone signal is at or below the predetermined threshold. In some example embodiments, one or more second levels may be higher than those for cases where the lowest frequency of at least one acquired microphone signal is above the predetermined threshold.

[0025] This method can be executed by the user equipment.

[0026] According to a third aspect, a computer program product is described, including a set of instructions configured, when executed on a device, to cause the device to perform a method comprising: enabling acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal; and applying ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein, in response to the AEC processing being enabled, the level of ambient noise suppression applied to the at least one post-processed acquired microphone signal increases from a first level to one or more second levels.

[0027] In some example embodiments, the third aspect may include any other features mentioned in relation to the method of the second aspect.

[0028] According to a fourth aspect, a non-transitory computer-readable medium is described, the non-transitory computer-readable medium including program instructions stored thereon for performing a method comprising: enabling acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal; and applying ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein in response to the AEC processing being enabled, the level of ambient noise suppression applied to at least one post-processed acquired microphone signal increases from a first level to one or more second levels.

[0029] In some example embodiments, the fourth aspect may include any other features mentioned in relation to the method of the second aspect.

[0030] According to a fifth aspect, an apparatus is described, comprising at least one processing core and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processing core, to cause the apparatus to: enable acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal; and apply ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein, in response to the AEC processing being enabled, the level of ambient noise suppression applied to at least one post-processed acquired microphone signal increases from a first level to one or more second levels.

[0031] In some example embodiments, the fifth aspect may include any other features mentioned in relation to the method of the second aspect. Attached Figure Description

[0032] Exemplary embodiments will be described by way of non-limiting examples with reference to the accompanying drawings, wherein:

[0033] Figure 1 This illustrates a communication session between the first user and the second user;

[0034] Figure 2 A front view of a second user is shown during the output of an audio signal from an audio output device;

[0035] Figure 3 This is a flowchart illustrating operations according to one or more example embodiments;

[0036] Figure 4 It shows the method for execution Figure 3 The operating device;

[0037] Figure 5Timing diagrams according to one or more example embodiments are shown;

[0038] Figure 6 The spectrum of the audio signal representing speech and ambient audio is shown;

[0039] Figure 7 The spectrum is shown after applying acoustic echo cancellation (AEC) processing;

[0040] Figure 8 A spectrum diagram is shown after the application of environmental noise suppression according to one or more example embodiments;

[0041] Figure 9A This is a flowchart illustrating operations according to an implementation example based on one or more other example embodiments;

[0042] Figure 9B The left and right speakers of the audio output devices associated with the first and second microphones are shown;

[0043] Figure 10 Example audio signal and microphone signal waveforms are shown;

[0044] Figure 11 Showing the target Figure 10 Example of waveform-related process;

[0045] Figure 12A The alignment of audio signals with microphone signals according to one or more example embodiments is shown;

[0046] Figure 12B The alignment of an audio signal with a microphone signal is illustrated according to one or more other example embodiments;

[0047] Figure 13 Another apparatus that can be configured according to one or more example embodiments is shown;

[0048] Figure 14 Another apparatus is shown that can be configured according to one or more example embodiments;

[0049] Figure 15 Showing more details Figure 14 The device's AEC module;

[0050] Figure 16 The illustration shows functional modules of an apparatus that can be configured according to one or more example embodiments; and

[0051] Figure 17 A non-transitory computer-readable medium program is shown, on which instructions are stored for performing a method according to one or more example embodiments. Detailed Implementation

[0052] Various example embodiments relate to apparatus, methods, and computer programs for microphone signal processing.

[0053] The treatment may involve acoustic echo cancellation (AEC) treatment and ambient noise suppression treatment, wherein at least in part based on the AEC treatment being enabled, the level of ambient noise suppression is increased from a first level to one or more second levels.

[0054] As described herein, AEC processing can include any known methods or algorithms for removing or mitigating acoustic echo components from microphone signals; AEC processing is typically used to cancel far-end speech received from a far-end user, where the far-end speech may be picked up by one or more microphones of a near-end user when output via one or more loudspeakers, such that only or primarily the near-end user's near-end speech is transmitted back to the far-end user. In other examples, AEC processing is not limited to canceling echoes from near-end speech. Other examples can include any form of speech, such as directional, non-reverberant speech and / or other desired audio signals.

[0055] AEC processing may produce unwanted artifacts for reasons explained in detail below.

[0056] On the other hand, ambient noise suppression is a general term encompassing a variety of methods. In these methods, specific types of audio (e.g., speech audio) are classified as desired or wanted signals, while other audio is classified as ambient noise or alternatively as background noise. Ambient audio can include, for example, at least one of unwanted noise, reverberant audio, distant audio, non-speech audio, or non-directional audio. The use of the term "noise" here does not imply that ambient audio must be unwanted in all scenarios, as ambient audio can provide a more natural listening experience. Such methods can involve beamforming, machine learning (ML)-based methods, blind source separation (BSS) methods, etc. The example embodiments are not limited to any particular method.

[0057] Figure 1 Example scenario 100 is shown, in which a first user 102 (remote user) communicates with a second user 104 (near-end user) as part of a communication session (e.g., a voice call). Other possible scenarios or use cases are described later.

[0058] The first user 102 and the second user 104 may be provided with corresponding first user equipment 106 and second user equipment 108. The first user 102 and the second user 104 may also be provided with corresponding first audio output device 110 and second audio output device 112.

[0059] The first user equipment 106 may include one or more microphones for acquiring audio from the first user 102. The use of one or more microphones may generate a corresponding first microphone signal. This corresponding first microphone signal may be encoded and transmitted via network 118 in one or more downlink signals 114 to the second user equipment 108. The second user equipment 108 may cause the output of the received one or more downlink signals 114 via the second audio output device 112. For example, the second user equipment 108 may communicate with the second audio output device 112 via a wired or wireless channel (e.g., Bluetooth, Zigbee, WiFi, etc. in the case of a wireless channel).

[0060] Similarly, the second user equipment 108 may include one or more microphones for acquiring audio from the second user 104. The one or more microphones may generate corresponding second microphone signals. These corresponding second microphone signals may be encoded and transmitted to the first user equipment 106 via network 118 as one or more uplink signals 116. The first user equipment 106 may cause the output of one or more uplink signals received via the first audio output device 110. For example, the first user equipment 106 may communicate with the first audio output device 110 via a wired or wireless link (e.g., Bluetooth, Zigbee, WiFi, etc. in the case of a wireless channel).

[0061] Network 118 may include an Internet Protocol (IP) network or other forms of communication network, such as a Radio Access Network (RAN). The respective air interfaces between the first user equipment 106 and the second user equipment 108 and network 118 may be configured to support a cellular or non-cellular radio access technology (RAT), depending on whether both the first and second user equipment and the network are configured. Examples of cellular RATs include Long Term Evolution (LTE) or 5G New Radio (NR) radio access technology, or beyond 5G, or 6G radio access technology, or other communication technologies.

[0062] The first audio output device 110 and the second audio output device 112 may each include a set of first and second speakers of any suitable form, such as a set of speakers for headphones, earbuds, headsets, or head-mounted devices (such as extended reality (XR) headsets). The term headphones or headphone device will be used hereinafter. The first audio output device 110 and the second audio output device 112 may be of the same type or may be of different types.

[0063] The first user device 106 and the second user device 108 may include any device that includes one or more microphones (or a device connected to one or more remote microphones). The first user device 106 and the second user device 108 may, for example, each include a smartphone, tablet computer, personal computer, laptop computer, wearable computer, Internet of Things (IoT) computer, or digital assistant. The first user device 106 and the second user device 108 may be of the same type or may be of different types.

[0064] Figure 2 This is a front view of the second user 104 during the output of an audio signal from the second audio output device 112. The second user device 108 can communicate with the second audio output device 112 using a wireless channel such as Bluetooth channel 209. The second user device 108 is located at a distance from the second user 104 and is typically in front of the second user 104. The second audio output device 112 includes a headset device comprising a left speaker 202 and a right speaker 204 that output corresponding audio sounds, hereinafter referred to as a first audio signal 206 and a second audio signal 208. The second user device 108 may include a body 205 on which a spaced-out first microphone 212 and a second microphone 214 are provided for capturing audio 210 from the second user 104. The spaced-out first microphone 212 and second microphone 214 generate a first microphone signal and a second microphone signal. In other example embodiments, one microphone or two or more microphones may be present.

[0065] At least some energy of the first and / or second audio signals 206, 208 can be captured by the first and / or second microphones 212, 214 during output. If so, the downlink signal 116 transmitted by the second user equipment 108 will include some energy of the first and / or second audio signals 206, 208. Therefore, when the downlink signal 116 is output by the first audio output device 110, the first user 102 can perceive acoustic echoes or other forms of unwanted audible feedback.

[0066] The scenario 100 described above, in which a second user equipment 108 with one or more microphones 212, 214 is physically separated from a second audio output device 112 providing a left-hand speaker 202 and a right-hand speaker 204, is particularly useful, but not exclusively, for stereo or spatial audio acquisition and output. The known spatial audio codec mentioned by way of example is the Immersive Speech and Audio Services (IVAS) codec, which has been standardized by the 3rd Generation Partnership Project (3GPP) for use in voice services. Regarding spatial audio output, the use of headphone devices or similar devices is generally superior to the output of systems relying on stand-alone speaker systems or those integrated into the user equipment, which tend to produce a “tinny” sound with a significant deficiency in low-frequency reproduction. Moreover, with user equipment speakers, stereo or spatial reproduction is generally not well perceived due to the relatively close proximity of the speakers. In terms of spatial audio acquisition, because in a microphone that is, for example, part of a headphone device, the microphone will be relatively close to the user's head (with acoustic shadowing from the opposite side of the user's head), and because the microphones may be relatively close to each other, the user device microphone may be superior to, for example, the microphone that is part of the headphone device. An unknown distance exists between the microphones, depending on the size of the user's head.

[0067] Therefore, for stereo or spatial audio acquisition and reconstruction, separate audio acquisition and audio output devices are generally preferred. However, the example embodiments are not limited to this use case, and other example embodiments may include apparatuses that include both audio acquisition and audio output devices.

[0068] AEC processing methods are known and typically involve the use of adaptive filters to estimate the acoustic transfer function (including delay) from one or more speakers to one or more microphones, where the acoustic transfer function is used to subtract the adaptively filtered speaker signal from the resulting microphone(s) signal using the delay. Additionally, a residual echo suppression component can suppress residual echo. However, AEC processing methods may not work effectively. This may be at least in part due to poor performance of the adaptive filters. For example, there may be unknown delays between the wireless transmission of signals from the user equipment, for example, via a Bluetooth channel, to the audio output device, and delays associated with the subsequent processing and output of these signals. AEC processing methods may also assume that the audio acquisition and audio output devices are part of the same device using a common clock signal. Nonlinearities may also be introduced due to the relatively low bit rate used for wireless transmission of one or more audio signals to the audio output device and processes such as equalization and / or compression that can be performed by the audio output device. Typically, AEC methods may assume that the sound paths from the first and second speakers to one or more microphones are relatively constant, whereas in the case of separate audio acquisition and audio output devices, these paths can change relatively abruptly and frequently, for example, when the user moves and / or rotates the audio acquisition device.

[0069] Given the limitations mentioned above, such as poor or only approximate time delay estimation, a relatively large amount of suppression is used in AEC processing methods to counteract poor adaptive filter performance. This can even involve enabling AEC processing before and / or after the detection of an echo (i.e., the trigger event used to enable AEC processing). Even greater suppression may be needed in cases where the audio output device includes a set of headphones with relatively large audio leakage and a wide frequency range. This overly aggressive approach to AEC processing can produce audible (larger than normal) artifacts that can be uncomfortable for the listener. For example, artifacts may be due to perceived fluctuations in the ambient audio(s) caused by AEC processing being enabled and disabled at multiple levels.

[0070] The following describes apparatus, methods, and computer programs that can avoid or mitigate at least some of these problems.

[0071] Figure 3 This is a flowchart illustrating operation 300 according to one or more example embodiments. Operation 300 can be performed in hardware, software, firmware, or a combination thereof. For example, operation 300 can be performed individually or jointly by components, wherein these components may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the execution of the operation. Operation 300 can, for example, be performed by... Figure 1The first user equipment 106 and the second user equipment 108 described herein shall be used to perform the operation.

[0072] The first operation 301 may include enabling acoustic echo cancellation (AEC) processing on at least one acquired microphone signal in response to a triggering event, wherein the AEC processing is performed during a first time period to provide at least one post-processed acquired microphone signal.

[0073] AEC processing can involve any conventional methods used for AEC processing. See Figure 9 to [the following text is missing from the original] for further details. Figure 13 Example implementations are described, designed to improve overall performance for near-end user operation of audio acquisition devices including multiple microphones and discrete audio output devices including multiple speakers, where the distance between the microphones and speakers can change frequently and abruptly. However, it should be understood that the example embodiments are not limited to this specific implementation and work well with other commonly used AEC processing methods where the distance between one or more microphones and the distance between one or more speakers are fixed and / or change only by a limited amount, such as in examples of flexible or foldable devices and / or head-mounted devices. The following also describes… Figure 14 and Figure 15 This shows another example implementation of this AEC processing method.

[0074] The second operation 302 may include applying ambient noise suppression processing to at least one post-processed acquired microphone signal, wherein in response to AEC processing being enabled, the level of ambient noise suppression applied to at least one acquired microphone signal is increased from a first level to one or more second levels.

[0075] Environmental noise suppression processing can involve any conventional method used for environmental noise suppression. The reference to one or more second levels clarifies that only one second level may be used, or multiple different second levels may be used, which can be determined based on one or more parameters, such as whether the lowest frequency to which the AEC processing is applied is below a predetermined threshold.

[0076] By applying ambient noise suppression in a controlled manner in response to the activation of AEC processing, it can be seen that artifacts are removed or at least reduced due to the AEC processing. This will be explained below, for example, by referring to [reference needed]. Figures 6 to 8 The spectrum shown effectively smooths the temporal and frequency characteristics of at least one acquired microphone signal in the presence of artifacts and / or the most perceptible parts of the artifacts by applying ambient noise suppression.

[0077] In some example embodiments, at least one of the acquired microphone signals represents at least in part speech audio primarily from a near-end user but potentially also from a far-end user as described above, as well as ambient noise in addition to the speech audio generated by sources farther from the near-end user.

[0078] In some example embodiments, the AEC processing of the first operation 301 can substantially eliminate or remove all far-end audio, such as both speech audio and ambient audio output by an audio output device (such as a set of headphones) associated with at least one microphone; in contrast, the ambient noise suppression processing of the second operation 302 can suppress or remove both near-end and far-end ambient audio.

[0079] The apparatus configured to perform the first operation 301 and the second operation 302 (or any related operations as described herein) may include any electronic device capable of generating or receiving at least one microphone signal. For example, the apparatus may include at least one microphone for acquiring at least one microphone signal, or the at least one microphone signal may be received by the apparatus from an external device including one or more microphones. The apparatus may also include one or more speakers for outputting remote audio, or alternatively, the apparatus may be associated with an external device including one or more loudspeakers for outputting remote audio. For example, the external device may include a set of headphones or the like, wherein the external device communicates with the apparatus via a wired or wireless (e.g., Bluetooth, etc.) link. The apparatus may be included in a user device, such as a smartphone or the like, examples of which are mentioned above.

[0080] Figure 4 A block diagram of an apparatus 400 according to some example embodiments is shown. For example, apparatus 400 may include user equipment, examples of which are described above.

[0081] Device 400 may include an AEC processing module 402 and an ambient noise suppression module 404. These modules may be separate modules as shown in the figure, or, in other example embodiments, may include a single module that performs both processing functions. Device 400 may also include at least one microphone 406 for acquiring audio signals(s) and generating at least one corresponding acquired microphone signal 406A, 406B, which is provided as input to AEC processing module 402. AEC processing module 402 may also receive one or more audio signals 410 that have been or are being output by one or more speakers of audio output device 412 associated with device 400 as input.

[0082] AEC processing module 402 or another processing module (not shown) can be configured to detect a trigger event according to known methods, in particular the presence of one or more echo components in at least one of the acquired microphone signals 406A, 406B. AEC processing module 402 can be enabled in response to the detection of the trigger event and perform AEC processing for a first time period. The first time period can be a finite time period, such as the period during which the echo continues to be detected.

[0083] The AEC processing module 402 can generate at least one post-processed version of the acquired microphone signals 406A, 406B, which is provided as input to the ambient noise suppression module 404 for further processing.

[0084] The ambient noise suppression module 404 can apply ambient noise suppression processing to a post-processed version of at least one acquired microphone signal 406A, 406B.

[0085] In response to the AEC processing module 402, or alternatively its actual AEC processing being enabled, the applied level of ambient noise suppression increases from a first level to one or more second levels. The ambient noise suppression module 404 may receive a control signal from the AEC processing module 402 indicating the enabling, or in other example embodiments, may receive the control signal or a different control signal from another module such as a control module (not shown).

[0086] In some example embodiments, the application of ambient noise suppression processing is disabled before AEC processing is enabled and enabled in response to AEC processing being enabled. In this case, the first level can be zero (no ambient noise suppression applied), and one or more second levels include one or more non-zero levels of ambient noise suppression.

[0087] In other example embodiments, ambient noise suppression processing may have been enabled before AEC processing was enabled. In this case, the first level may be non-zero (a relatively low level of ambient noise suppression has been applied), and one or more second levels may include one or more relatively high levels of suppression.

[0088] One or more second levels may include only one second level or include multiple second levels.

[0089] The signal 414 generated by the ambient noise suppression module 404 can be provided (e.g., via antenna 416) as a signal for remote users (e.g. Figure 1 The uplink signal of the first user 102 in the system. The remote user can be reached via user equipment (e.g., Figure 1The first user equipment 106) receives the uplink signal and outputs audio corresponding to the uplink signal via the first audio output device 110.

[0090] Figure 5 A first timing diagram 502 and a second timing diagram 512 are shown, which are associated with AEC processing and ambient noise suppression processing, respectively, and can be executed by AEC processing module 402 and ambient noise suppression processing module 404, respectively.

[0091] Referring to the first timing diagram 502, at the first time instance t1, a triggering event can be detected, such as the presence of an echo in at least one of the acquired microphone signals 406A and 406B.

[0092] AEC processing can be enabled responsively, for example, by enabling AEC processing module 402 from a disabled state. Full AEC processing (i.e., full echo suppression) may require a finite time period to take effect, as indicated by reference numeral 504, which indicates the rate of increase in the level of echo suppression. This can be a relatively abrupt transition and can be on the order of tens of milliseconds.

[0093] AEC processing is maintained during the first time period TP1 until such a time instance t2 (when it can be determined that AEC processing is no longer needed, for example, in response to the detection that no echo is detected in at least one of the acquired microphone signals 406A, 406B).

[0094] Similar to the above, AEC processing (i.e., echo suppression) may require a limited time period to become completely disabled or reduced to a minimum, as indicated by reference numeral 506 in the attached figure, which indicates the rate at which AEC processing decreases. This can be a relatively abrupt transition, and can be on the order of tens of milliseconds.

[0095] Note that the level of AEC inhibition is shown as constant over the first time interval TP1, but this is not necessarily the case in all examples.

[0096] Referring to the second timing diagram 512, at or shortly thereafter at the first time instance t1, in response to the activation of AEC processing, ambient noise suppression can be applied, wherein the level of ambient noise suppression increases from the first level L1 to the second level L2.

[0097] In the example shown, the ambient noise suppression module 404 may be disabled before it increases to the second level L2. In other example embodiments, the ambient noise suppression module 404 may have been enabled and a relatively low level of ambient noise suppression may have been applied, which is then increased to the second level L2 at or after the beginning of the first time period t1.

[0098] In some example embodiments, the level of ambient noise suppression increases gradually at a rate indicated by reference numeral 514, which is smaller than the rate 504 at which the AEC processing is applied to at least one acquired microphone signal 406A, 406B. For example, the rate 514 at which the level of ambient noise suppression increases can be on the order of seconds, compared to the rate 504 at which the AEC processing increases (which can be on the order of milliseconds).

[0099] In some example embodiments, the level of ambient noise suppression can be maintained at the second level L2 at least until the end of the first time period TP1. And then the level of ambient noise suppression can be reduced toward the first level L1 (i.e., reduced to the first level L1 or reduced to another reduced level).

[0100] In some example embodiments, the level of ambient noise suppression decreases gradually at a rate indicated by reference numeral 516, which is smaller than the rate 506 at which AEC processing of at least one acquired microphone signal 406A, 406B decreases at the end of the first time period TP1. For example, the rate 516 of decreasing the level of ambient noise suppression can be on the order of seconds, compared to the rate 506 of AEC processing decrease (which can be on the order of milliseconds).

[0101] The following is for reference. Figures 6 to 8 Mention the advantages associated with such gradual changes.

[0102] In some example embodiments, ambient noise suppression processing can be performed on the entire or arbitrarily wide frequency range, or on one or more finite frequency ranges that may depend on the AEC processing frequency range. For example, AEC processing can be applied to a first frequency range of at least one acquired microphone signal 406A, 406B, and noise suppression processing can be applied to a second frequency range of at least one acquired microphone signal, wherein the second frequency range is determined based on the first frequency range. For example, the second frequency range can be substantially the same as the first frequency range, or the second range can be a finite frequency range that is wider than the first frequency range and includes the first frequency range. For example, the first frequency range can be 1.5kHz-3.5kHz, and the second frequency range can be 1kHz-4kHz. Limiting the frequency range can require less processing resources / energy, and less ambient audio will be suppressed, thus providing a more natural user experience.

[0103] As can be understood, the audibility of artifacts can be more perceived at lower frequencies because ambient noise is generally louder at lower frequencies. Therefore, in some example embodiments, the level of ambient noise suppression processing can be based at least in part on the lowest frequency (i.e., the lowest frequency requiring AEC processing) of at least one of the acquired microphone signals 406A, 406B (which include echo components) relative to a predetermined threshold (e.g., a relatively low frequency threshold).

[0104] For example, another operation may include determining the lowest frequency of at least one acquired microphone signal 406A, 406B that includes an echo component and determining whether the lowest frequency is at or below a predetermined threshold. In one example, the level of ambient noise suppression is increased from a first level to one or more second levels according to the second operation 302 only if the lowest frequency is at or below the predetermined threshold. In another example, if the lowest frequency is at or below the predetermined threshold, one or more second levels are set to be higher than or increased for the case where the lowest frequency is above the predetermined threshold. For example, one or more second levels may increase as the lowest frequency decreases. In this way, if lower frequency signals require echo cancellation, more ambient noise suppression is applied, thereby counteracting more perceptible artifacts.

[0105] Figures 6 to 8 The corresponding first spectrogram 600, second spectrogram 700, and third spectrogram 800 are shown, which help to understand the advantages of the example embodiments.

[0106] refer to Figure 6 The first spectrogram 600 represents the time and frequency characteristics of at least one acquired microphone signal 406A, 406B. Different portions 602, 604, and 606 of the first spectrogram 600 indicate the corresponding types of acquired audio. For example, the first portion 602 may represent near-field speech, the second portion 604 may represent far-field speech echo components, and the third portion 606 may represent ambient noise. Ambient noise may, for example, represent the audio from a concert.

[0107] refer to Figure 7 The second spectrogram 700 shows the time and frequency characteristics of at least one acquired microphone signal 406A, 406B after AEC processing, wherein the AEC processing is associated with the first section 604 and is graphically indicated by reference numeral 702. For reasons already explained, adaptive filter limitations require overly aggressive AEC processing; therefore, the AEC processing covers a wider frequency range than the second section 604 and uses a relatively strong level of echo cancellation. This results in the AEC processing section 702 being related to… Figure 6The relatively abrupt change at the boundary portion 704 between the ambient noise portions 606 indicated in the diagram. Partly because ambient noise is typically non-transient, these abrupt changes cause audible and transient artifacts perceptible to remote users, and the abrupt changes are easily perceived. This effect is further exacerbated if ambient noise exists between sequential instances of AEC processing, as the ambient noise level will fluctuate significantly.

[0108] refer to Figure 8 According to an example embodiment, the third spectrogram 800 represents the time and frequency characteristics of at least one acquired microphone signal 406A, 406B after AEC processing and after the level of the applied ambient noise suppression processing is increased from a first level L1 to one or more second levels L2. It can be seen that the effect is smooth. Figure 6 The aforementioned abrupt change in at least some of the boundary portions 704 avoids or mitigates audible artifacts.

[0109] In the example described above, the environmental noise suppression treatment is gradually increased and / or gradually decreased at a slower rate than that of the AEC treatment, which helps to further make the transition less noticeable as the effect is gradually (rather than abruptly) introduced and reduced.

[0110] Figures 6 to 8 Assume that ambient noise suppression processing is applied to a relatively wide frequency range (e.g., all audio frequencies) of at least one microphone signal 406A, 406B. Also, as explained above, ambient noise suppression processing may be limited to a smaller frequency range, which may be the same frequency range as the AEC processing applied, or a wider frequency range than the AEC processing frequency range and include the AEC processing frequency range (but not covering all audio frequencies). AEC processing implementation example

[0111] An example implementation of the AEC processing module 402 is now described, which is particularly suitable for situations where near-end user operation includes an audio acquisition device with multiple microphones and a separate audio output device with multiple speakers.

[0112] Figure 9A This is a flowchart illustrating operation 900 according to an implementation example. Operation 900 can be performed in hardware, software, firmware, or a combination thereof. For example, operation 900 can be performed individually or jointly by components, wherein the components may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the execution of the operation.

[0113] The first operation 901 may include receiving one or more microphone signals from corresponding microphones for acquiring corresponding first and second audio signals output by the first and second speakers. The second operation 902 may include associating a first combination of the first and second audio signals with the one or more microphone signals. The third operation 903 may include determining a time delay in which the first combination of the first and second audio signals most similar to the one or more microphone signals. The fourth operation 904 may include aligning the first and second audio signals with the one or more microphone signals based on the time delay. The fifth operation 905 may include attenuating the one or more microphone signals by an attenuation amount to provide at least one post-processed microphone signal, the attenuation amount being determined at least in part based on a second combination of the aligned first audio signal and the aligned second audio signal. The first through fifth operations 905 may correspond to... Figure 3 The first operation 301 is to enable AEC processing. The sixth operation 906 may include applying ambient noise suppression processing to at least one post-processed microphone signal, wherein, in response to AEC processing being enabled, the level of ambient noise suppression increases from a first level to one or more second levels. It should also be noted that the above and following references to the first and second combinations for aligning the first and second audio signals are given as examples and should not be considered limiting. Other methods for aligning the first and second audio signals may be used in other exemplary embodiments.

[0114] In some examples, a first combination of the first audio signal and the second audio signal (hereinafter referred to as the "first combination") may include a weighted sum of the first audio signal and the second audio signal, for example: Y = w1 (first audio signal) + w2 (second audio signal), Where w1 and w2 are the corresponding weights, and their sum can be equal to 1.

[0115] For example, the first combination may include one item from the following (non-exhaustive) list of audio signal combinations y, where italic values ​​indicate the corresponding weights: 0.0 (first audio signal) + 1.0 (second audio signal); 0.1 (first audio signal) + 0.9 (second audio signal); 0.2 (first audio signal) + 0.8 (second audio signal); 0.5 (first audio signal) + 0.5 (second audio signal); 0.8 (first audio signal) + 0.2 (second audio signal); 0.9 (first audio signal) + 0.1 (second audio signal); and 1.0 (first audio signal) + 0.0 (second audio signal). Table 1 - Example Set of Audio Signals

[0116] As will be seen, the first and last items in Table 1 indicate that the first combination Y includes only the second audio signal and only the first audio signal, respectively. The other items indicate the corresponding intermediate weighting of a certain amount of both the first and second audio signals. In some examples, one of the first and second audio signals in the first combination has a smaller gain than the aligned first and second audio signals in the second combination. In some examples, correlation may include performing cross-correlation or a similar similarity function to determine the maximum similarity value. In some instances, the first combination can be determined by correlating each microphone signal x in one or more microphone signals with each audio signal combination y, the first combination being determined as the audio signal combination most similar to at least one of the one or more microphone signals. In other words, the first combination is the pairing of audio signal combination y with microphone signal x that produces the highest maximum similarity or correlation value.

[0117] For example, a correlation can be performed for each of the following pairs (x, y) of microphone signal x and audio signal combination y: x = first microphone signal, y = first audio signal only; x = second microphone signal, y = first audio signal only; x = first microphone signal, y = second audio signal only; x = second microphone signal, y = second audio signal only; x = first microphone signal, y = first audio signal + second audio signal; and x = second microphone signal, y = first audio signal + second audio signal. Table 2 - Examples Related

[0118] The first four items in Table 2 indicate that, for y, only the correlation of one of the first and second audio signals is used, according to items one and seven of Table 1. Items five and six of Table 2 indicate that, for y, the correlation of a specific sum of the first and second audio signals is used, according to items two through six of Table 1.

[0119] After determining the first combination, the time delay can be determined based on the time displacement of the first combination relative to the generation of the highest maximum similarity value of the microphone signal x. The time delay is the time delay mentioned in the third operation 303.

[0120] Figure 9BThe diagram illustrates the relationship between the first microphone signal 212 and the second microphone 214 during the output of the first audio signal 206 and the second audio signal 208. Figure 2 The left speaker 202 and the right speakers 202 and 204.

[0121] The first microphone 212 can collect at least some energy of a first audio signal 206 and / or a second audio signal 208 indicated by a corresponding first path a and a second path b. The first microphone 212 generates a first microphone signal, which may include the collected first audio signal 206 and / or the second audio signal 208. The second microphone 214 can collect at least some energy of the first audio signal 206 and / or the second audio signal 208 indicated by a corresponding third path c and a fourth path d. The second microphone 214 generates a second microphone signal, which may include the collected first audio signal 206 and / or the collected second audio signal 208. In some cases, the first microphone 212 and / or the second microphone 214 may not collect the energy of the first audio signal 206 or the second audio signal 208.

[0122] If the first microphone 212 and / or the second microphone 214 “hear” the first audio signal 206 and / or the second audio signal 208, an acoustic echo may occur. This could be if a proportion of the signals arrive at the microphone(s) at a level higher than that of other sound sources, or at a level up to 10 dB lower than that of other sound sources, or at a level otherwise higher than that of the internal or ambient noise associated with the microphone(s). The lengths of paths a through d may vary significantly and may change abruptly and frequently depending on the position and / or orientation of the second user 104 relative to the second user equipment 108. For example, if the first path a is significantly shorter than the fourth path d, this means that the first microphone 212 may collect (or hear) more energy of the first audio signal 206 than the second audio signal 208. The echo effect is unlikely to be particularly strong (because, in the case of headphone devices, there is typically a small amount of audio leakage outside the user's ear), and therefore attenuation according to the fifth operation 305 is only necessary if the left-hand amplifier 202 and the right-hand amplifier 204 are relatively close (e.g., 1 meter or less) to the second user equipment 108, and there may be no significant sound source near the user equipment. This proximity can be identified based on a high correlation or similarity between the first combination and at least one of the first and second microphone signals.

[0123] According to the second operation 902, the signal pair (x, y) can be expected from Figure 9: x = first microphone signal, y = first audio signal It will have the highest maximum similarity value.

[0124] In other words, the final item in Table 1 can be determined as the first combination (y = 1.0 (first audio signal) + 0.0 (second audio signal)).

[0125] The time delay may include the amount of time displacement of the first audio signal 206 relative to the first microphone signal, as it will produce the highest maximum similarity value.

[0126] Figure 10 Example time-domain waveforms for a first audio signal 206 and a second audio signal 208, and a first microphone signal 1006 and a second microphone signal 1008 are shown. It will be seen that the first microphone signal 1006 is an attenuated version of the first audio signal 206 with a specific time delay d1, and the second microphone signal 1008 is a more attenuated version of the first audio signal with a specific time delay d2, where d2 > d1.

[0127] In this example, neither the first microphone 1006 nor the second microphone 1008 picks up or hears the second audio signal 208, but in other examples, the situation may be different.

[0128] Figure 11 The example demonstrates how to perform cross-correlation in the time domain for only two signal pairs (x, y), namely: x = first microphone signal, y = first audio signal; and x = second microphone signal, y = first audio signal.

[0129] Reference numeral 1102 graphically indicates how cross-correlation can be performed using time window 1104. The length of time window 1104 can be set and thus limited based on an estimated time delay for the data representing the arrival of the first audio signal 206 and the second audio signal 208 at the first microphone 206 and the second microphone 208. The time delay for time window 1104 can, for example, include 3 ms (approximate time it takes for sound to travel 1 meter). This is because it can be assumed that if the first microphone 212 and the second microphone 214 are more than 1 meter away from the left speaker 202 and the right speaker 204, the microphones will not pick up or hear the first audio signal 206 and the second audio signal 208. Additionally, there may be further delays due to the wireless channel (e.g., Bluetooth channel 209) between the second user terminal 108 and the second audio output device 112, and also due to processing and / or buffering performed at the second audio output device. These delays can be longer than the aforementioned 3 ms delay and can be ignored in some cases. Assuming the worst-case scenario, the time delay for time window 1104 can be as high as 400 ms. The time delay can typically be expected to be around 100-200ms.

[0130] Figure 1106 graphically indicates the corresponding first time delay D1 and second time delay D2 when the maximum similarity (cross-correlation) is measured.

[0131] The attached figure, labeled 1108, graphically indicates the position of the similarity or cross-correlation value C(1110, 1112) (which can vary between 0 and 1) and the corresponding maximum similarity (cross-correlation) values ​​Cmax1 and Cmax2.

[0132] Therefore, in this simple example, the signal pair (x, y) includes: x = first microphone signal 1006, y = first audio signal 206 Because Cmax1 > Cmax2, this signal pair produces the highest maximum similarity / correlation value.

[0133] Therefore, the first combination will actually include the final item in Table 1: 1.0 (first audio signal) + 0.0 (second audio signal).

[0134] The time delay used for the purpose of the third operation 903 may include at least the first time delay D1.

[0135] refer to Figure 12A The first audio signal 206 and the second audio signal 208 can be aligned with the first microphone signal 1006 and the second microphone signal 1008 based on a time delay D1.

[0136] The first microphone signal 1006 and the second microphone signal 1008 can be attenuated by an attenuation amount A, which is determined at least in part based on a second combination of the aligned first audio signal 206 and the aligned second audio signal 208.

[0137] The first audio signal 206 and the second audio signal 208 may combine in an unintended manner during their journey to the first microphone 212 and the second microphone 214 (e.g., due to the characteristics of the user's head, the room in which the user is located, and / or the characteristics of the audio output device), which may cause reflections and damping. Therefore, the safest option may be to attenuate all (in this case, the first microphone signal and the second microphone signal) microphone signals 1006, 1008 individually, or at least attenuate those microphone signals in at least one pair whose similarity value is higher than a threshold.

[0138] For the same reason, the attenuation A can be based on the worst-case combination of the first audio signal 206 and the second audio signal 208, for example, based on the summation of the aligned first audio signal 206 and the aligned second audio signal 208.

[0139] Therefore, the second combination may include the sum of the aligned first audio signal 206 and the aligned second audio signal 208, and the attenuation may be based at least in part on this sum.

[0140] In some instances, the sum of the aligned first audio signal 206 and the aligned second audio signal 208 can be a weighted sum, for example: Y = w3 (first audio signal) + w4 (second audio signal) w3 and w4 are the corresponding weights, and their sum can be equal to 1.

[0141] In some examples, the corresponding weights w3 and w4 can both include 0.5. In this case, the second audio signal 208 will have a smaller gain in the first combination than in the second combination.

[0142] In some examples, the corresponding weights w3 and w4 can be in the range of 0.3 to 0.7, so that their sum is 1.0.

[0143] In some examples, in the first combination, at least one of the weights w3 and w4 for one of the audio signals 206 and 208 is less than the corresponding weight for the same audio signal in the second combination.

[0144] In some examples, in the first combination, at least one of the weights w3 and w4 for one of the audio signals 206 and 208 is greater than the corresponding weight for the same audio signal in the second combination.

[0145] In some instances, the corresponding weights w3 and w4 can be based on the correlation between one or more microphone signals 1006 and 1008 and the first combination. In some examples, the greater the correlation, the greater the attenuation.

[0146] In some examples, because the time delay D1 is an estimate, a relatively long time window and / or a smooth envelope can be used instead of quickly and accurately following the shape of the first audio signal 206 and the second audio signal 208 to perform attenuation.

[0147] In some examples, the correlation value C is smoothed over time using previous correlation estimates.

[0148] In some examples, the attenuation A can have a maximum value of 5-20 dB.

[0149] In some examples, the attenuation A is determined on a per-subband basis (e.g., for each subband).

[0150] In some examples, the subband can cover a frequency range of 1-5 kHz.

[0151] refer to Figure 12B In an alternative example, the first audio signal 206 and the second audio signal 208 can be aligned with the first microphone signal 1006 and the second microphone signal 1008 based on corresponding first time delays D1 and D2. For example, the first audio signal 206 and the second audio signal 208 can be aligned with the first microphone signal 1006 based on the first time delay D1, and the first audio signal and the second audio signal can be aligned with the second microphone signal 1008 based on the second time delay D2. The corresponding weights w3 and w4 for the first microphone signal 1006 and the second microphone signal 1008 can be based on corresponding highest similarity (relevance) values ​​(i.e., based on Cmax1 for the first microphone signal 1006 and Cmax2 for the second microphone signal).

[0152] In some examples, attenuation of at least the first microphone signal 1006 and the second microphone signal 1008 can be performed in the frequency domain, and the attenuated first microphone signal and the attenuated second microphone signal can then be converted to the time domain for output.

[0153] In some examples, operations 902 through 905 can be performed in the frequency domain, as will now be described using a general example for determining attenuation A.

[0154] In general, the microphone signal x and the audio signal y can be framed, windowed (e.g., using a 20ms long window with 50% overlap), and transformed to the frequency domain using, for example, a Fast Fourier Transform (FFT). Other transforms and / or filter banks can also be used.

[0155] Signals x and y can be divided into frequency subbands (e.g., 1 / 3 octave, Bark, and / or similar subband division methods).

[0156] Signal X i,j,k and signal Y i,j,k It can be derived that i is the frame index, j is the subband index, and k is the binary number in the given subband. The correlation value C between signal x and signal y can be calculated as: Equation (1A) corresponds to a zero-delay correlation. In some examples, correlations with different delays can be calculated by applying different time frames / data to one of the signals x or y, where the different time frames / data are delayed compared to time frame i. For example:

[0157] Time frames with different delays (0..400ms) were tested to find the delay that gave the highest relevance.

[0158] The correlation value C for the signal pair (x, y) that produces the highest maximum similarity / correlation value can be used.

[0159] According to the third operation 903 and the fourth operation 904, the first audio signal and the second audio signal can be aligned with the first microphone signal and the second microphone signal using the time delay that produces the highest maximum similarity value.

[0160] The aligned first and second audio signals can be combined to create a worst-case-safe energy calculation. For example, the aligned first audio signal energy and the aligned second audio signal energy can be summed for each frame and each frequency band as follows:

[0161] The signal energy of each microphone signal x can be determined as: Where m is the microphone index.

[0162] The correlation value C rarely reaches 0 or 1. Therefore, for each microphone time frame and frequency band, a lookup table can be used (e.g., where the correlation value C is mapped to an attenuation A, which can be a value between 0 and 20 dB, e.g., 5 dB, and directly the maximum attenuation A). m (i,j))) is used to map the relevant values ​​to C to more useful values.

[0163] Maximum attenuation is used if the microphone signal is significantly lower than the worst-case signal energy (e.g., if the difference is 35 dB or greater). In some examples, attenuation is not used otherwise.

[0164] In some examples, when the difference between the microphone signal energy and the worst-case signal energy is not 35 dB, the attenuation used may be less than the maximum attenuation. The 35 dB value is merely an example value, and other values ​​may be used, for example, in the range of 20 to 50 dB.

[0165] According to operation 906 in Figure 9, the attenuation A can be applied to the microphone signal in the frequency domain and converted back to the time domain to apply ambient noise suppression.

[0166] In general, the example embodiments can reduce the perceived echo and also reduce the resulting artifacts through controlled application of ambient noise suppression. The example embodiments can be used in a variety of use cases, including but not limited to: - Voice call echo cancellation; - Audio (and possibly video) is recorded via a user device while the user of the user device listens to other audio (e.g., music) via headphones, wherein such other audio should not be recorded via the user device; and - When a user listens to other audio via headphones, audio voice commands are captured for speech recognition processing. The other audio should not interfere with speech recognition processing.

[0167] Figures 13 to 15 The following are illustrated according to some example embodiments. Figure 4 Alternative or detailed examples of the apparatus.

[0168] Figure 13 A block diagram of apparatus 1300 for AEC processing is shown for the implementation example described above with reference to Figures 9 to 12.

[0169] Apparatus 1300 may include an AEC processing module 1302 and an ambient noise suppression system 1304 based on the above implementation example. The ambient noise suppression system 1304 may include an ambient noise suppression module 1303 and a mixing module 1305. Apparatus 1300 may also include at least one microphone 1306, 1308 for acquiring audio 1311 output by at least one speaker 1312 of an associated audio output device (based on the received audio signal 1313). At least one microphone 1306, 1308 may provide at least one corresponding acquired microphone signal 1306A, 1308A, which is an input to or provided to the AEC processing module 1302. The AEC processing module 1302 may also receive the audio signal 1313 as input from the associated audio output device.

[0170] According to an implementation example, AEC processing module 1302 or another processing module (not shown) can be configured to detect a trigger event, specifically the presence of one or more echo components in at least one of the acquired microphone signals 1306A, 1308A. AEC processing module 1302 can be enabled in response to the detection of the trigger event and perform AEC processing for a first time period. The first time period can be a finite time period, such as the period during which the echo continues to be detected.

[0171] AEC processing module 1302 can generate at least one post-processed version 1324 of the acquired microphone signals 1306A, 1308A, which is an input to ambient noise suppression system 1304 or provided as an input to ambient noise suppression system 1304 for further processing.

[0172] The ambient noise suppression module 1303 can apply ambient noise suppression processing to the received post-processed version 1324 of at least one acquired microphone signal 1306A, 1308A to provide an output signal 1326 with reduced ambient noise.

[0173] The level of ambient noise suppression can be controlled via a mixing module 1305. The mixing module 1305 may include a first gain module 1316, a second gain module 1318, a gain controller 1320, and a mixer 1322. The first gain module 1316 may receive an output signal 1326 with reduced ambient noise and apply a gain "G" based on a control signal from the gain controller 1320. The second gain module 1318 may receive a post-processed version 1324 of at least one acquired microphone signal 1306A, 1306B and apply a gain "1-G" based on a control signal 1315 from the gain controller 1320. The corresponding outputs 1328, 1329 from the first gain module 1316 and the second gain module 1318 can be mixed by the mixer 1322 to provide an output signal to the antenna 1330.

[0174] Gain controller 1320 may receive control signal 1315 from AEC processing module 1302 to indicate whether (and when) AEC processing is enabled and optionally, the amount or level of AEC processing is applied to at least one acquired microphone signal 1306A, 1308A. When control signal 1315 indicates that AEC processing is enabled or the level is increasing, gain controller 1320 may control the value of G and therefore also the value of 1-G, such that the contribution from the first gain module 1316 to mixer 1322 increases to one or more second levels, and the contribution from the second gain module 1318 to mixer decreases accordingly. As described above, this can be performed gradually. For example, gain controller 1320 may increase the value of G from zero (or other minimum value) to one (or other maximum value) at a rate of 0.01 per frame (e.g., every 20ms frame). When control signal 1315 indicates that AEC processing is disabled (or will be disabled) or is being reduced, gain controller 1320 can control the value of G and thus the value of 1-G, such that the contribution from the first gain module 1316 to mixer 1322 decreases toward a first level, and the contribution from the second gain module 1318 to mixer increases accordingly. As described above, this can be performed gradually. For example, gain controller 1320 can reduce the value of G from one (or other maximum value) to zero (or other minimum value) at a rate of 0.002 per frame (e.g., every 20ms frame). In this case, the value of G becomes zero within ten seconds. Other values ​​of G can be used to achieve a similar effect where ambient noise suppression decreases to a minimum over several seconds or (even minutes) and / or ramps up the value of G at a faster rate over several seconds or tens of seconds. In fact, this is a method where the level of ambient noise suppression changes slowly compared to the rate at which AEC processing for echo cancellation of at least one acquired microphone signal 1306A, 1308A is changed. Instead of using the hybrid module 1305, other methods may involve using machine learning (NL) methods to determine the level of environmental noise suppression.

[0175] For the sake of completeness, Figure 14 A block diagram of an apparatus 1400 according to another example embodiment is shown, which addresses a different form of AEC processing module 1402 (different from those described above with reference to Figures 9 to 12) used for AEC processing. Apparatus 1400 includes... Figure 13 The device 1300 contains some components that are the same as or similar to those in the drawing. Similar elements are indicated by similar reference numerals and may be assumed to be identical to those in the drawing. Figure 13 The device operates in the same or similar manner as the 1300. Figure 15A block diagram of an example functional module of a conventional AEC processing module 1402 is shown. The conventional AEC processing module 1402 may include an adaptive filter module 1502, a residual echo suppression module 1504, and a mixer 1506, the details of which are unnecessary or not considered necessary for understanding the example embodiments described herein. Example device

[0176] Figure 16 An example device 1600 capable of supporting at least some embodiments is shown. Device 1600 may include, for example... Figure 4 , Figure 13 or Figure 14 Any of the illustrated devices 400, 1300, or 1400 may include at least a portion of the user equipment of any of the preceding examples. Device 1600 includes a processor 1610, which may include, for example, a single-core or multi-core processor, wherein a single-core processor includes one processing core and a multi-core processor includes more than one processing core. Processor 1610 may typically include a control device. Processor 1610 may include more than one processor. Processor 1610 may be a control device. Processor 1610 may include at least one application-specific integrated circuit (ASIC). Processor 1610 may include at least one field-programmable gate array (FPGA). Processor 1610 may be a component in device 1600 for performing method steps. Processor 1610 may be configured at least partially by computer instructions to perform actions.

[0177] The processor may include or be configured as a circuit system or multiple circuit systems configured to perform phases of the methods according to the embodiments described herein. As used herein, the term "circuit system" may refer to one or more of the following: (a) a hardware circuit implementation only (such as an implementation only in an analog / or digital circuit system) and (b) a combination of hardware circuitry and software, for example (if applicable): (i) a combination of (multiple) analog and / or digital hardware circuitry having software / firmware and (ii) any portion of (multiple) hardware processors having software (including (multiple) digital signal processors, software, (multiple) memories) working together to enable devices such as first user equipment 106 or second user equipment 108, or devices configured to control their operation, to perform various functions) and (c) (multiple) hardware circuitry and / or (multiple) processors, such as (multiple) microprocessors or portions thereof that require software (e.g., firmware) for operation (but may be absent when software is not required for operation).

[0178] This definition of circuit system applies to all instances of the term as used in this application (including any claim). As another example, as used herein, the term circuit system also encompasses implementations of hardware circuitry or processors (or processors) alone, or a portion thereof, and their accompanying software and / or firmware. The term circuit system also encompasses, for example and as applicable to specific claim elements, baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or networking devices.

[0179] Device 1600 may include memory 1620. Memory 1620 may include random access memory and / or permanent memory. Memory 1620 may include at least one RAM chip. Memory 1620 may include, for example, solid-state, magnetic, optical, and / or holographic memory. Memory 1620 may be at least partially accessible by processor 1610. Memory 1620 may be at least partially included in processor 1610. Memory 1620 may be a component for storing information. Memory 1620 may include computer instructions configured to be executed by processor 1610. When computer instructions configured to cause processor 1610 to perform certain actions are stored in memory 1620, and device 1600 as a whole is configured to run under the guidance of processor 1610 using computer instructions from memory 1620, processor 1610 and / or at least one of its processing cores may be considered to be configured to perform said certain actions. Memory 1620 may be at least partially included in processor 1610. The memory 1620 may be at least partially outside the device 1600, but may be accessed by the device 1600.

[0180] Device 1600 may include a transmitter 1630. Device 1600 may include a receiver 1640. The transmitter 1630 and receiver 1640 may be configured to transmit and receive information according to at least one cellular standard or a non-cellular standard, respectively.

[0181] Transmitter 1630 may include more than one transmitter. Receiver 1640 may include more than one receiver. Transmitter 1630 and / or receiver 1640 may be configured to operate, for example, according to Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), 5G / NR, 5G-Advanced (i.e., NR Rel-18, 19 and above), Long Term Evolution (LTE), IS-95, Wireless LAN, WLAN, Ethernet and / or Global Microwave Access Interoperability (WiMAX) standards.

[0182] Device 1600 may include a near-field communication (NFC) transceiver 1650. The NFC transceiver 1650 may support at least one NFC technology, such as NFC, Bluetooth, Wibree, or similar technologies.

[0183] Device 1600 may include a user interface (UI) 1660. UI 1660 may include at least one of a display, a keyboard, a touchscreen, a vibrator arranged to signal to a user by causing device 1600 to vibrate, a speaker, and a microphone. A user may be able to operate device 1600 via UI 1660, for example, to accept incoming telephone calls, initiate telephone or video calls, browse the Internet, manage digital files stored in memory 1620 or manage digital files in the cloud accessible via transmitter 1630 and receiver 1640, or accessible via NFC transceiver 1650, and / or play games.

[0184] Device 1600 may include or be arranged to accept user identification module 1670.

[0185] The user identification module 1670 may include, for example, a subscriber identification module, SIM card, or card that can be installed in the device 1600. The user identification module 1670 may include subscription information identifying the user of the device 1600. The user identification module 1670 may include password information that can be used to verify the identity of the user of the device 1600 and / or to facilitate the encryption of communicated information and to facilitate billing for communications initiated by the user of the device 1600.

[0186] Processor 1610 may be equipped with a transmitter arranged to output information from processor 1610 to other devices included in device 1600 via electrical leads within device 1600. Such a transmitter may include a serial bus transmitter arranged to output information to memory 1620 for storage, for example, via at least one electrical lead. As an alternative to a serial bus, the transmitter may include a parallel bus transmitter.

[0187] Similarly, processor 1610 may include a receiver arranged to receive information from other devices included in device 1600 via electrical leads within device 1600. Such a receiver may include a serial bus receiver arranged, for example, via at least one electrical lead from receiver 1640 for processing in processor 1610. Alternatively, the receiver may include a parallel bus receiver.

[0188] Device 1600 may include Figure 16Other devices not shown. For example, in the case where device 1600 includes a smartphone, it may include at least one digital camera. Some devices 1600 may include a rear camera and a front camera, wherein the rear camera may be designed for digital photography and the front camera for video calling. Device 1600 may include a fingerprint sensor arranged to at least partially authenticate the user of device 1600. In some embodiments, device 1600 lacks at least one of the above-described devices. For example, some devices 1600 may lack an NFC transceiver 1650 and / or a user identification module 1670.

[0189] Processor 1610, memory 1620, transmitter 1630, receiver 1640, NFC transceiver 1650, UI 1660, and / or user identity module 1670 can be interconnected in various ways via electrical leads within device 1600. For example, each of the aforementioned devices can be connected to a main bus within device 1600 to allow the devices to exchange information. However, as those skilled in the art will understand, this is merely an example, and various ways of interconnecting at least two of the foregoing devices may be chosen depending on the embodiment without departing from the scope of the invention.

[0190] Figure 17 A non-transitory medium 1700 according to some embodiments is shown. The non-transitory medium 1700 is a computer-readable storage medium. It may be, for example, a CD, DVD, USB stick, Blu-ray disc, etc. The non-transitory medium 1700 stores computer program instructions that cause a device to perform a method, such as any of the aforementioned processes disclosed with reference to the flowcharts and related features in this specification.

[0191] The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details, such as examples of length, width, shape, etc., have been provided in the foregoing description to provide a thorough understanding of embodiments of the invention. However, those skilled in the art will recognize that the invention can be practiced without one or more specific details or using other methods, components, materials, etc. In other instances, well-known structures, materials, or operations have not been shown or described in detail to avoid obscuring aspects of the invention.

[0192] While the foregoing examples are illustrative of the principles of embodiments in one or more specific applications, it will be apparent to those skilled in the art that many modifications can be made in form, use, and implementation details without practicing the invention and without departing from its principles and concepts. Therefore, the invention is not intended to be limiting except by the claims set forth below.

[0193] The verbs “comprising” and “including” are used in this document as open-ended restrictions, neither excluding nor requiring the presence of any unrecorded features. Unless otherwise expressly stated, the features recited in the dependent claims may be freely combined with each other. Furthermore, it should be understood that the use of “an” or “a” (i.e., the singular form) throughout this document does not exclude the plural form.

Claims

1. An apparatus for microphone signal processing, comprising: at least one processor; at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: in response to a trigger event, enable acoustic echo cancellation (AEC) processing for at least one captured microphone signal, wherein the AEC processing is performed for a first time period to provide at least one post-processed captured microphone signal; and apply ambient noise suppression processing to the at least one post-processed captured microphone signal, wherein a level of ambient noise suppression applied to the at least one post-processed captured microphone signal increases from a first level to one or more second levels in response to the AEC processing being enabled.

2. The apparatus of claim 1, wherein the application of the ambient noise suppression processing is disabled prior to the AEC processing being enabled and is enabled in response to the AEC processing being enabled, or the application of the ambient noise suppression processing is enabled prior to the AEC processing being enabled.

3. The apparatus of claim 1, wherein the level of ambient noise suppression increases from the beginning of the first time period.

4. The apparatus of claim 1, wherein the level of ambient noise suppression gradually increases at a rate that is less than a rate at which the AEC processing is applied to the at least one captured microphone signal.

5. The apparatus of claim 4, wherein the rate at which the level of ambient noise suppression increases is on the order of seconds and the rate at which the AEC processing increases is on the order of milliseconds.

6. The apparatus of claim 1, wherein the level of ambient noise suppression is maintained at the one or more second levels at least until the end of the first time period and then the level of ambient noise suppression decreases toward the first level.

7. The apparatus of claim 6, wherein the level of ambient noise suppression gradually decreases at a rate that is less than a rate at which the AEC processing decreases for the at least one captured microphone signal at the end of the first time period.

8. The apparatus of claim 7, wherein the rate at which the level of ambient noise suppression decreases is on the order of seconds and the rate at which the AEC processing decreases is on the order of milliseconds.

9. The apparatus of claim 1, wherein the AEC processing is applied to a first frequency range of the at least one captured microphone signal, and the noise suppression processing is applied to a second frequency range of the at least one post-processed captured microphone signal, wherein the second frequency range is determined based on the first frequency range.

10. The apparatus of claim 9, wherein the second frequency range is substantially the same as the first frequency range, or the second frequency range is wider than and includes the first frequency range.

11. The apparatus of claim 1, wherein ​ The level of the ambient noise suppression is based at least in part on a lowest frequency of the at least one captured microphone signal relative to a predetermined threshold.

12. The apparatus of claim 11, wherein The level of the ambient noise suppression is increased from the first level to the one or more second levels only if the lowest frequency of the at least one captured microphone signal is at or below the predetermined threshold.

13. The apparatus of claim 11, wherein The one or more second levels are higher in the case that the lowest frequency of the at least one captured microphone signal is at or below the predetermined threshold than in the case that the lowest frequency of the at least one captured microphone signal is above the predetermined threshold.

14. A method for microphone signal processing, comprising: in response to a trigger event, enabling an acoustic echo cancellation (AEC) process for at least one captured microphone signal, wherein the AEC process is performed for a first time period for providing at least one post-processed captured microphone signal; and applying an ambient noise suppression process to the at least one post-processed captured microphone signal, wherein in response to the AEC process being enabled, a level of ambient noise suppression applied to the at least one post-processed captured microphone signal is increased from a first level to one or more second levels.

15. The method of claim 14, wherein the application of the ambient noise suppression process is disabled prior to the AEC process being enabled and enabled in response to the AEC process being enabled, or the application of the ambient noise suppression process is enabled prior to the AEC process being enabled.

16. The method of claim 14, wherein the level of the ambient noise suppression is increased from the beginning of the first time period.

17. The method of claim 14, wherein the level of the ambient noise suppression is gradually increased at a rate that is less than a rate at which the AEC process is applied to the at least one captured microphone signal.

18. The method of claim 14, wherein the AEC process is applied to a first frequency range of the at least one captured microphone signal, and the noise suppression process is applied to a second frequency range of the at least one post-processed captured microphone signal, wherein the second frequency range is determined based on the first frequency range.

19. The method of claim 14, wherein the level of the ambient noise suppression is based at least in part on a lowest frequency of the at least one captured microphone signal relative to a predetermined threshold.

20. A non-transitory computer-readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: in response to a trigger event, enabling an acoustic echo cancellation (AEC) process for at least one captured microphone signal, wherein the AEC process is performed for a first time period for providing at least one post-processed captured microphone signal; and ​ applying an ambient noise suppression process to the at least one post-processed collected microphone signal, wherein in response to the AEC process being enabled, a level of ambient noise suppression applied to the at least one post-processed collected microphone signal is increased from a first level to one or more second levels.