Electronic apparatus, method performed by the same, storage medium and computer program product
The use of dual microphones in electronic devices to detect acoustic feedback path changes for rapid and precise howling suppression addresses the slow and inaccurate detection issues in existing technologies, enhancing user experience by preventing howling.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-08-20
- Publication Date
- 2026-07-23
Smart Images

Figure KR2025012633_23072026_PF_FP_ABST
Abstract
Description
ELECTRONIC APPARATUS, METHOD PERFORMED BY THE SAME, STORAGE MEDIUM AND COMPUTER PROGRAM PRODUCT
[0001] The present disclosure relates to a field of signal processing, and specifically to a method performed by an electronic apparatus, the electronic apparatus, a computer-readable storage medium, and a computer program product.
[0002] The function of sound amplification is an important function in audio control that may be integrated into many apparatuses (such as, various types of earphones (e.g., a True Wireless Stereo (TWS) earphone, etc.), a hearing aid, etc.). Due to the phenomenon of self-amplification and sound wave superposition produced by sound in the propagation process, sharp piercing sound may be generated in electronic apparatuses such as an earphone and a hearing aid, such a phenomenon is referred to as howling. The occurrence of the howling may affect the normal operation of these apparatuses and lead to a poor listening experience of the user, and therefore, suppression of the howling is required in these apparatuses. However, current howling suppression methods suffer from the problems such as slow speed of howling detection and low accuracy of the howling detection, etc, resulting in poor howling suppression. In view of this, a better howling suppression technology is required.
[0003] According to a first aspect of embodiments of the present disclosure, it is provided a method performed by an electronic apparatus, the electronic apparatus comprising a first microphone and a second microphone, the method comprising: acquiring a signal difference between the first microphone and the second microphone; determining whether an acoustic feedback path of at least one of the first microphone and the second microphone changes based on the signal difference; and performing howling suppression in a case of determining that the acoustic feedback path changes.
[0004] In an embodiment, the acquiring of the signal difference between the first microphone and the second microphone comprises: acquiring a difference between a first acoustic feedback path of the first microphone and a second acoustic feedback path of the second microphone, as the signal difference.
[0005] In an embodiment, the electronic apparatus further comprises a speaker, the acquiring of the difference between the first acoustic feedback path of the first microphone and the second acoustic feedback path of the second microphone comprising: acquiring the difference between the first acoustic feedback path and the second acoustic feedback path based on a first microphone signal collected by the first microphone, a second microphone signal collected by the second microphone, a speaker signal played by the speaker, and a signal amplification gain corresponding to the speaker.
[0006] In an embodiment, the acquiring of the difference between the first acoustic feedback path and the second acoustic feedback path based on the first microphone signal collected by the first microphone, the second microphone signal collected by the second microphone, the speaker signal played by the speaker, and the signal amplification gain corresponding to the speaker comprises: determining the difference between the first acoustic feedback path and the second acoustic feedback path at each frequency point within a preset time period based on energy information corresponding to the each frequency point of the first microphone signal and the second microphone signal within the preset time period, the signal amplification gain, and energy information of the speaker signal.
[0007] In an embodiment, the method further comprising: outputting signals collected by the first microphone and the second microphone if the signal difference is less than a first threshold; wherein the determining whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference comprises: determining that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, in a case of the signal difference being not less than the first threshold.
[0008] In an embodiment, the outputting of the signals collected by the first microphone and the second microphone if the signal difference is less than the first threshold, comprises: determining that neither of first microphone and the second microphone howls and outputting the signals collected by the first microphone and the second microphone, if the difference between the first acoustic feedback path and the second acoustic feedback path at each frequency point within a preset time period is less than the first threshold, wherein the determining that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, in a case of the signal difference being not less than the first threshold, comprises: determining that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, if the difference between the first acoustic feedback path and the second acoustic feedback path at at least one frequency point within the preset time period is not less than the first threshold.
[0009] In an embodiment, the determining whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference comprises: extracting a long-time difference feature and a short-time difference feature based on the difference between the first acoustic feedback path and the second acoustic feedback path; determining whether the acoustic feedback path changes based on the long-time difference feature and the short-time difference feature.
[0010] In an embodiment, the performing of the howling suppression comprises: performing the howling suppression based on whether the first microphone and the second microphone howl simultaneously, in a case that the first microphone and the second microphone do not howl simultaneously, based on a signal corresponding to a first frequency point at which howling occurs among microphone signals collected by a microphone of the first microphone and the second microphone that does not howl, performing the howling suppression on a signal corresponding to the first frequency point among microphone signals collected by a microphone of the first microphone and the second microphone that howls; in a case that the first microphone and the second microphone howl simultaneously, performing the howling suppression by performing wave limiting processing based on a second frequency point at which the howling occurs simultaneously in a first microphone signal and a second microphone signal.
[0011] In an embodiment, the performing of the howling suppression based on whether the first microphone and the second microphone howl simultaneously comprises: determining the first frequency point at which the howling occurs based on the signal difference, and performing fusion processing on the first microphone signal and the second microphone signal to obtain processed microphone signal based on the first microphone signal collected by the first microphone, the second microphone signal collected by the second microphone, and the first frequency point, wherein, in the fusion processing, the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that howls is replaced with the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that does not howl; determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal; in a case of determining that the first microphone and the second microphone do not howl simultaneously, outputting the processed microphone signal as an output signal after performing the howling suppression; in a case of determining that the first microphone and the second microphone howl simultaneously, performing, based on the second frequency point, the wave limiting processing on the processed microphone signal to obtain the output signal after performing the howling suppression.
[0012] In an embodiment, the determining of the first frequency point at which the howling occurs based on the signal difference comprises: determining a frequency point at which the difference between the first acoustic feedback path and the second acoustic feedback path is not less than a second threshold as the first frequency point at which the howling occurs for the microphone of the first microphone and the second microphone that howls.
[0013] In an embodiment, the performing of the fusion processing on the first microphone signal and the second microphone signal to obtain the processed microphone signal comprises: encoding the first microphone signal and the second microphone signal respectively in time domain to obtain encoded features respectively corresponding to the first microphone signal and the second microphone signal; fusing the obtained encoded features based on the first frequency point, wherein, in the fusion, encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that howls are replaced with encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that does not howl; decoding the fused encoded features to obtain the processed microphone signal.
[0014] In an embodiment, the microphone of the first microphone and the second microphone that howls is determined based on the first acoustic feedback path and the second acoustic feedback path.
[0015] In an embodiment, the determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal, comprises: extracting signal features of the first microphone signal and the second microphone signal; determining whether the first microphone and the second microphone howl simultaneously and the second frequency point using an artificial intelligence network, based on the signal features.
[0016] In an embodiment, the signal features comprise at least one of energy growth features, a signal similarity feature, and signal envelope features of the first microphone signal and the second microphone signal.
[0017] In an embodiment, the signal features comprise energy growth features of the first microphone signal and the second microphone signal, wherein the extracting of the signal features of the first microphone signal and the second microphone signal comprises: calculating frequency domain energy of each of the first microphone signal and the second microphone signal; based on the frequency domain energy, extracting the energy growth features of the first microphone signal and the second microphone signal, respectively, wherein the determining whether the first microphone and the second microphone howl simultaneously and the second frequency point, based on the signal features, comprises: determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching a predetermined energy growth feature, wherein the predetermined energy growth feature corresponds to an energy growth feature when predetermined howling occurs; determining the second frequency point based on the energy growth features.
[0018] In an embodiment, the signal features further comprise a signal similarity feature of the first microphone signal and the second microphone signal, wherein the extracting of the signal features of the first microphone signal and the second microphone signal further comprises: extracting a change trend of a signal similarity of the first microphone signal and the second microphone signal as the signal similarity feature, wherein the determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching the predetermined energy growth feature, comprises: determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching the predetermined energy growth feature and the signal similarity of the first microphone signal and the second microphone signal decreasing.
[0019] In an embodiment, the determining of the second frequency point based on the energy growth features comprises: determining a frequency point as the second frequency point, if a change of energy growth degree of the first microphone and the second microphone at the frequency point is increasing over time and a similarity of the energy growth degree of the first microphone and the second microphone is below a predetermined similarity threshold.
[0020] In an embodiment, the performing, based on the second frequency point, the wave limiting processing on the processed microphone signal to obtain the output signal after performing the howling suppression comprises: performing, utilizing a wave limiter designed based on the second frequency point, the wave limiting processing on the processed microphone signal to obtain the output signal after performing the howling suppression.
[0021] According to a second aspect of embodiments of the present disclosure, it is provided an electronic apparatus, comprising: a memory; and a processor coupled to the memory and configured to perform the methods described above.
[0022] According to a third aspect of embodiments of the present disclosure, it is provided a computer-readable storage medium having stored thereon a computer program or instructions which, when executed by at least one processor, cause(s) the at least one processor to perform the methods described above.
[0023] According to a fourth aspect of embodiments of the present disclosure, it is provided a computer program product comprising a computer program which, when executed by a processor, implements the methods described above.
[0024] According to the technical solutions of the embodiments of the present disclosure, the electronic apparatus acquires a signal difference between the first microphone and the second microphone, determines whether an acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, and performs howling suppression in a case of determining that the acoustic feedback path changes, and since the essence of the howling occurrence is the change of the acoustic feedback path, through the above method, the occurrence of the howling may be detected more quickly and accurately, and the howling suppression may be performed in a timely manner in a case that the howling occurrence is detected, so that the howling is not heard by the user, thereby improving the listening experience of the user.
[0025] It should be understood that the above general description and the detailed descriptions that follow are merely exemplary and explanatory and do not limit the present disclosure.
[0026] The accompanying drawings herein are incorporated into and form part of the specification, illustrate embodiments consistent with the disclosure, which are used in conjunction with the specification to explain the principles of the disclosure and do not constitute an undue limitation of the disclosure.
[0027] FIG. 1 is a schematic diagram illustrating an acoustic feedback path.
[0028] FIGS. 2 and 3 are schematic diagrams illustrating examples of howling generation.
[0029] FIG. 4 is a schematic diagram illustrating a principle of howling generation.
[0030] FIG. 5 is a schematic diagram illustrating time of howling suppression.
[0031] FIG. 6 is a flowchart of a method performed by an electronic apparatus according to an embodiment of the present disclosure.
[0032] FIG. 7 is a schematic diagram illustrating an example of the method shown in FIG. 6 according to an embodiment of the present disclosure.
[0033] FIG. 8 is a schematic diagram illustrating acoustic feedback paths for dual microphones according to an embodiment of the present disclosure.
[0034] FIG. 9 is a schematic diagram illustrating a difference between acoustic feedback paths of dual microphones according to an embodiment of the present disclosure.
[0035] FIG. 10 is a schematic diagram illustrating single microphone howling detection based on path change according to an embodiment of the present disclosure.
[0036] FIG. 11 is a schematic diagram illustrating a difference between acoustic feedback paths changes with the acoustic feedback path.
[0037] FIG. 12 is a schematic diagram illustrating single microphone howling suppression.
[0038] FIG. 13A is a schematic diagram illustrating difference values of acoustic feedback paths in open environment.
[0039] FIG. 13B is a schematic diagram determining a first frequency point at which howling occurs.
[0040] FIG. 14 is a schematic diagram illustrating dual microphone signal fusion.
[0041] FIG. 15 is a schematic diagram illustrating a signal envelope feature of a microphone that howls.
[0042] FIG. 16 is a schematic diagram illustrating energy growth features of the dual microphone signals.
[0043] FIG. 17 is a schematic diagram illustrating dual microphone howling detection based on energy growth features.
[0044] FIG. 18 is a block diagram illustrating an electronic apparatus according to an embodiment of the present disclosure.
[0045] FIG. 19 is a schematic diagram illustrating a structure of an electronic apparatus according to an embodiment of the present disclosure.
[0046] The following description with reference to the accompanying drawings is provided to aid in a thorough understanding of various embodiments of the present disclosure as defined by claims and equivalents thereof. This description includes various specific details to aid in understanding but should only be considered exemplary. Accordingly, those ordinary skills in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known features and structures may be omitted for the sake of clarity and brevity.
[0047] The terms and phrases used in the claims and the following description are not limited to dictionary meaning thereof, but are used only by the inventor to enable a clear and consistent understanding of the present disclosure. Accordingly, it should be apparent to those skilled in the art that, the following description of the various embodiments of the present disclosure is provided for an illustrative purpose only and is not intended to a purpose of limiting the present disclosure as defined by the appended claims and equivalents thereof.
[0048] It should be understood that, "a", "an" and "the" in a singular form may also include a plural reference, unless the context clearly indicates otherwise. Thus, for example, a reference to a "part surface" includes a reference to one or more such surfaces. When it refers to one element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to a connection relationship between the one element and the other element established through an intermediate element. In addition, "connected" or "coupled" as used herein may include wirelessly connected or wirelessly coupled.
[0049] The term "include" or "may include" refers to the presence of a function, operation, or component of the corresponding disclosure that may be used in the various embodiments of the present disclosure, and does not limit the presence of one or more additional functions, operations, or features. In addition, the terms "include" or "have" may be interpreted to denote certain features, figures, steps, operations, constituent elements, components, or combinations thereof, but should not be interpreted to exclude the possibility of the presence of one or more other features, figures, steps, operations, constituent elements, components, or combinations thereof.
[0050] The term "or" as used in the various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing a plurality of (two or more) items, the plurality of items may refer to one, more, or all of the plurality of items if a relationship among the plurality of items is not explicitly defined. For example, for the description "a parameter A comprises A1, A2, A3", it may be implemented as parameter A comprising A1, A2 or A3, or as parameter A comprising at least two of the three items of the parameter A1, A2, A3.
[0051] All terms (including technical or scientific terms) used in the present disclosure have the same meaning as understood by those skilled in the art to which the present disclosure belongs, unless defined differently. Common terms as defined in dictionaries are interpreted to have a meaning consistent with the context in the relevant technology art and should not be interpreted in an idealized or overly formalistic manner, unless expressly so defined in the present disclosure.
[0052] At least part of the functions in a device or electronic apparatus provided in the embodiments of the present disclosure may be implemented through an AI model, such as, at least one of a plurality of modules of the device or electronic apparatus may be implemented through the AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor.
[0053] The processor may include one or more processors. At this time, the one or more processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, or may be a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).
[0054] The one or more processors control processing of input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0055] Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or an AI model of a desired characteristic is made. The learning may be performed in a device or electronic apparatus itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.
[0056] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a neural network calculation by calculating between the input data of this layer (such as, a calculation result of the previous layer and / or the input data of the AI model) and the plurality of weight values of the current layer. Examples of neural networks include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial networks (GAN), and a deep Q-network.
[0057] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0058] The methods provided in the present disclosure may involve one or more of technical fields such as speech, language, image, video, or data intelligence.
[0059] In an embodiment, when involving the field of speech or language, in the method according to the present disclosure executed by electronic apparatus, a speech signal, which is an analog signal, may be received via speech input devices (e.g., a microphone), and the speech part is converted into computer readable text using an automatic speech recognition (ASR) model. The user's intent of utterance may be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technique for recognizing and applying / processing human language / text and includes, e.g., natural language processing, machine translation, dialog system, question answering, or speech recognition / synthesis.
[0060] In an embodiment, when involving the field of image or video, in the method according to the present disclosure executed by electronic apparatus, output data may be obtained by using image data as input data for an artificial intelligence model. The method of the present disclosure may involve the field of visual understanding in the artificial intelligence technology, and the visual understanding is a technique for recognizing and processing things as does human vision and includes, e.g., object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.
[0061] In an embodiment, when involving the field of data intelligence processing, in the method according to the present disclosure executed by electronic apparatus, in the reasoning or predicting stage, an artificial intelligence model can be used to perform predictions by using real-time input data. Processors of the electronic apparatus may perform a pre-processing operation on the data to convert into a form appropriate for use as an input for the artificial intelligence model. Reasoning and prediction is a technique of logically reasoning and predicting by determining information and includes, e.g., knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.
[0062] In the present application, the artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.
[0063] Below, the technical solutions of the embodiments of the disclosure and the technical effects produced by the technical solutions of the disclosure will be explained by describing several optional embodiments. It should be noted that, the following embodiments may be referred to, imitated or combined with each other, and the same term, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0064] For ease of understanding, the following first provides a brief introduction of the relevant background knowledge involved in the technical solutions of the present disclosure.
[0065] An acoustic feedback path is used to represent a channel through which a sound signal propagates in the air, which is used to describe the propagation process of a sound signal from a sound-emitting end (e.g., a speaker 110) to a receiving end (e.g., a microphone 120). the acoustic feedback path may characterize a transmission response corresponding to a sound signal from the sound-emitting end (e.g., a speaker 110) that is collected at the receiving end (e.g., the microphone 120). FIG. 1 is a schematic diagram illustrating an acoustic feedback path. The sound signal is emitted from the speaker 110 in a ray-like manner, and is transmitted to the microphone 120 through processes such as wall reflections. The reflection processes are related to the size of the obstacle and the distance from the sound source, and emitting waves generated by these different reflection processes are superposed to form the acoustic feedback path.
[0066] Howling refers to a phenomenon of self-amplification and sound wave superposition produced by sound in the propagation process, which may generate sharp piercing sound in a sound reinforcement system, an earphone, a hearing aid and other products, affecting the normal operation of these apparatuses. FIGS. 2 and 3 are schematic diagrams of examples of howling generation. FIG. 2 is a schematic diagram of howling generation in a conference room sound amplification system. As shown in FIG. 2, since the speaker 210 and the microphone 230 are in the same physical space, the signal received by the microphone 230 is transmitted to the speaker 210 for the playback after being amplified by the audio amplifier 220, then collected again by the microphone 230. It is cycled like that and repeated for many times, which forms positive feedback at certain frequencies, causes self-excited oscillations of the acoustic signal, which results in the howling. Personal sound amplification products (PSAPs), such as an earphone, also have microphones 230, amplifiers 220 and speakers 210, which also produce positive feedback and generate the howling. There are many scenarios in life that generate the howling, so the howling is also an issue that must be considered in PSAP products. For example, when the earphone is in amplification gain, the howling may occur if the earphone is blocked by an object. FIG. 3 is a schematic diagram illustrating an example of howling generation in an earphone 310. As shown in FIG. 3, the howling may be generated when the earphone 310 is blocked by an object (e.g., a hand 320). In addition to this, the howling may be generated when the earphone 310 is covered by a hat, a scarf, a body while being hugged, or a cell phone while making a phone call.
[0067] The howling occurs in many scenarios in life, however, current howling suppression technologies suffer from problems such as slow speed of howling detection and low accuracy of the howling detection, etc, resulting in poor howling suppression.
[0068] For example, the current howling suppression technologies have at least one of the following problems:
[0069] 1. Unable to suppress the howling, causing damage to the user's hearing and affecting the use feeling:
[0070] In the current howling suppression technologies, it is determined whether the howling occurs through the magnitude comparison between a ratio of peak energy to average energy of the sound signal and a threshold. If the average energy of the sound signal is large, the howling can be detected only when the howling is completely generated (i.e., when the peak value is large), and the detection speed is slow. For example, when a music signal and a howling signal occur at the same time, only when energy of the howling signal is much larger than energy of the music signal, the howling occurrence can be detected, but at this time, the howling signal has been too large, which may cause damage to the user's hearing.
[0071] 2. Easy for misdetection, generating noise and interfering with the normal use of the user:
[0072] For the current howling suppression technologies, if they want to detect the howling before it reaches an audible threshold, this can be only achieved by lowering the detection threshold, which may misidentify a normal voice signal as the howling signal, resulting in howling-like noise. For example, when an external signal is a harmonic signal such as music, the music signal may be incorrectly detected as the howling signal and processed, resulting in distortion of the music signal.
[0073] In this regard, the present disclosure proposes a howling suppression scheme based on a difference of dual microphones from the principle of the howling generation, which may more quickly and accurately determine whether howling may be generated from the condition of the howling generation, so that the howling may be suppressed in a timely manner before the audible threshold of the howling and the user cannot hear the howling.
[0074] Studies find that the principle of the howling generation is: as shown in FIG. 4, taking the earphone as an example, the microphone in the earphone receives the external signal S(n) for amplification gain, and then broadcasts Y(n) through the speaker. For example, since the earphone is not completely closed when worn, the signal of the speaker may leak into the microphone, i.e., x(n)=S(n)+Y(n)*H(n) for closed-loop repeated amplification, and the feedback system becomes unstable if |H·G|≥1 is satisfied. When using PSAP, usually, the amplification gain G is a constant, which shows that the condition of the howling generation is that the acoustic feedback path H changes. For example, when the earphone is blocked by an object such as a hand, H changes, leading to the howling generation.
[0075] For example, as shown in FIG. 5, the acoustic feedback path changes at the time t1, the howling increases enough to be heard by the user at the time t2, and the whole process takes about 10 milliseconds. At the time t2, the howling signal has not increased to the audible threshold. If the howling is suppressed before the time t2, the user will not hear the howling. During this period of time from t1 to t2, detecting and suppressing the trend of the howling generation may fundamentally prevent the user from hearing the howling, however, the current howling suppression methods detect the howling only at time t3 when the howling is completely generated and the user has already heard the piercing howling. In this regard, the present disclosure proposes that, according to the principle of the howling generation, the path change is the essential cause of the howling generation. Therefore, if it is desired to detect the howling before the time t2, whether the acoustic feedback path changes may be determined. In the case that the acoustic feedback path changes, the howling suppression is performed. The most direct way to estimate the change of the acoustic feedback is to estimate the acoustic path in real time and observe its change on the time axis.However, in the audio amplification process, the external audio signal and the acoustic environment of the earphone are unknown. It is impossible to observe its change by directly estimating the acoustic feedback path. In this regard, in consideration that the hardware structure of a general sound amplification apparatus usually has at least two microphones, the present disclosure proposes that whether the acoustic feedback path changes may be determined based on a signal difference between the microphones, so as to exclude the influence of unknown ambient sound.
[0076] In accordance with the above inventive concept, the present disclosure provides a method performed by an electronic apparatus. FIG. 6 is a flowchart of a method performed by an electronic apparatus according to an embodiment of the present disclosure. According to embodiments, the electronic apparatus may be an earphone apparatus, a hearing aid apparatus, and the like, but is not limited thereto. According to embodiments, the electronic apparatus may include at least two microphones, such as a first microphone and a second microphone. According to embodiments, the above at least two microphones are closer to each other, and a similarity of the collected microphone signals is higher, such as in the TWS earphone, the hearing aid apparatus, and the like.
[0077] Referring to FIG. 6, at step S610, a signal difference between the first microphone and the second microphone is acquired. At step S620, whether an acoustic feedback path of at least one of the first microphone and the second microphone changes is determined based on the signal difference. At step S630, howling suppression is performed in a case of determining that the acoustic feedback path changes. Since the essence of the howling occurrence is the change of the acoustic feedback path, through the above method, the occurrence of the howling may be detected more quickly and accurately. The howling suppression may be performed in a timely manner in the case that the howling occurrence is detected, so that the user cannot hear the howling, thereby improving the listening experience of the user.
[0078] In the following, various steps in the method performed by the electronic apparatus according to embodiments of the present disclosure will be described in connection with examples. FIG. 7 is a schematic diagram illustrating an example of the method shown in FIG. 6 according to an embodiment of the present disclosure. In the following, the method shown in FIG. 6 will be described in detail in connection with the example illustrated in FIG. 7.
[0079] In an embodiment, step S610 may include: acquiring a difference between a first acoustic feedback path of the first microphone and a second acoustic feedback path of the second microphone, as the signal difference. Although in the following, the difference between the first acoustic feedback path of the first microphone and the second acoustic feedback path of the second microphone is taken as the signal difference between the first microphone and the second microphone, the signal difference between the first microphone and the second microphone is not limited thereto, and any form of the signal difference may be acquired.
[0080] Since the electronic apparatus receives both the ambient sound and repeatedly amplified howling components, it is difficult to directly detect the change of the acoustic feedback path of single microphone. The acoustic feedback paths of the two microphones are different, and whether the acoustic feedback path changes may be determined based on the difference between the acoustic feedback paths of the two microphones.
[0081] In step S710, single microphone howling may be detected based on path change.
[0082] In an embodiment, a signal of the first microphone (e.g., microphone signal 1), a signal of the second microphone (e.g., microphone signal 2) and a signal of the speaker (e.g., speaker signal) may be acquired. Each of the signals may be transformed into frequency-domain signal. The difference between the acoustic feedback paths of the two microphones may be calculated based on the frequency domain signal of the dual microphone and speaker. The difference may be acquired through trained AI model. When the difference is small, it may be considered that the feedback path is definitely not changed, and the at least one signal of first microphone and the second microphone is directly amplified. When the difference is large, it may be considered that howling may occur.
[0083] In an embodiment, the analysis of howling frequency points may be carried out in the frequency domain. If the howling frequency points are replaced in the frequency domain, it may be also necessary to carry out time-frequency transformation, which will lead to a large delay in signal generation. The generation of pure audio signals may be carried out in the time domain through the network. It is necessary to learn the characteristics of the frequency point without howling in the second microphone in the time domain through the network, and fuse the characteristics into the time domain signal of the first microphone and the second microphone.
[0084] In step S720, dual microphone signals may be fused.
[0085] In an embodiment, the time-domain signal of the first microphone (e.g., microphone signal 1) and the time-domain signal of the second microphone (e.g., microphone signal 2) may be acquired. The first frequency point at which howling occurs may be acquired. According to the first frequency point, the time-domain signals may be fused and the howling signal is replaced with a clean signal. The operation does not need to carry out time-frequency conversion. Therefore, the delay is small. It can achieve the same effect of frequency domain replacement.
[0086] In step S730, dual microphone howling may be detected based on signal features.
[0087] In an embodiment, under normal circumstances, the signals of the two microphones may be mainly environmental sounds. The components of the two signals may be basically same. The signal similarity of two signals may be high. The signal energy may increases together when the howling occurs. However, the howling characteristic may be different because the acoustic path is different. When the similarity between the two signals decreases and energy growth conforms to the howling growth pattern, it may be a howling state.
[0088] In an embodiment, dual microphone howling may be detected based on the signal energy growth characteristics and signal similarity. First, the time-frequency features of the two microphone signals may be extracted respectively. It may be judged whether they are howling according to their envelope shape and signal energy growth. The extracted signal features may be used as the input of CNN network for training. The network may judge the signal similarity and outputs the frequency of howling. The corresponding trap (e.g., notch filter) may be designed according to the howling frequency output of the network, and the microphone signal may be suppressed by howlingFIG. 8 is a schematic diagram illustrating acoustic feedback paths of dual microphones according to an embodiment of the present disclosure, which illustrates a frequency domain diagram of the acoustic feedback paths H1 and H2 of the two microphones in an open environment. As shown in FIG. 8, the acoustic feedback paths of the two microphones are different due to the different locations of the two microphones.
[0089] FIG. 9 is a schematic diagram illustrating a difference between acoustic feedback paths of dual microphones according to an embodiment of the present disclosure. As shown in FIG. 9, a difference value between the two acoustic feedback paths H1 and H2 is small when the acoustic feedback paths do not change. The difference value between the two acoustic feedback paths H1 and H2 is large when the acoustic feedback path changes. It is possible to determine whether the acoustic feedback path changes based on the difference between the acoustic feedback paths. It is possible to suppress the howling before the howling increase to the audible threshold, so that the user cannot hear the howling.
[0090] According to embodiments, the electronic apparatus may further include a speaker. The difference between the first acoustic feedback path and the second acoustic feedback path may be acquired based on a first microphone signal collected by the first microphone, a second microphone signal collected by the second microphone, a speaker signal played by the speaker, and a signal amplification gain corresponding to the speaker. In an embodiment, the difference between the first acoustic feedback path and the second acoustic feedback path at each frequency point within a preset time period may be determined based on energy information corresponding to the each frequency point of the first microphone signal and the second microphone signal within the preset time period, the signal amplification gain, and energy information of the speaker signal. As an example, the difference between the first acoustic feedback path and the second acoustic feedback path may be acquired using an artificial intelligence network based on the first microphone signal, the second microphone signal, the speaker signal and the signal amplification gain.
[0091] As shown in FIG. 7, in step S710, single microphone howling detection based on path change may be performed based on the microphone signal 1, the microphone signal 2, the speaker signal and the signal amplification gain. For example, the difference between the first acoustic feedback path and the second acoustic feedback path may first be acquired based on frequency domain signal X1 corresponding to the microphone signal 1, frequency domain signal X2 corresponding to the microphone signal 2, frequency domain signal Y corresponding to the speaker signal, and signal amplification gain G. Then, it is determined, based on the difference, whether the acoustic feedback path changes and a first frequency point at which the howling occurs. For example, the difference between the first acoustic feedback path and the second acoustic feedback path may be a magnitude difference between the first acoustic feedback path and the second acoustic feedback path. In the present disclosure, the "first frequency point at which the howling occurs" may be a frequency point at which the howling may occur, rather than necessarily a frequency point at which the howling does occur.
[0092] FIG. 10 is a schematic diagram illustrating single microphone howling detection based on path change S710 according to an embodiment of the present disclosure. Assuming X is a frequency domain signal corresponding to the microphone signal, Y is a frequency domain signal corresponding to the speaker signal, H is the acoustic feedback path, S denotes a frequency domain signal corresponding to the ambient acoustic signal, and G denotes a signal amplification gain corresponding to the speaker, X = S + YHG, and the frequency domain signals corresponding to two microphone signals may be represented as , . As shown in FIG. 10, according to the frequency-domain signals of the two microphones (signal X1 and signal X2) and the frequency-domain signal of the speaker (signal Y), as well as the signal amplification gain G, a path difference calculation may be performed to obtain the amplitude difference between the two acoustic feedback paths. Specifically, , .
[0093] In an embodiment, the method shown in FIG. 6 may further include: outputting signals collected by the first microphone and the second microphone if the signal difference is less than a first threshold. For example, if the difference between the first acoustic feedback path and the second acoustic feedback path is less than the first threshold, it may be considered that the howling does not occur. In the case, the howling suppression may not be performed and the signals collected by the first microphone and the second microphone may be directly output. As an example, it is determined that neither of first microphone and the second microphone howls and the signals collected by the first microphone and the second microphone are output, if the difference between the first acoustic feedback path and the second acoustic feedback path at each frequency point within a preset time period is less than the first threshold. If the signal difference is not less than the first threshold, it is considered that the howling may occur. It is determined, based on the signal difference, whether the acoustic feedback path changes, so as to accurately determine whether the howling does occur. In an embodiment, step S620 may include: determining that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, in a case of the signal difference being not less than the first threshold. For example, it is determined that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, if the difference between the first acoustic feedback path and the second acoustic feedback path at at least one frequency point within the preset time period is not less than the first threshold. In an embodiment, instead of making a determination of whether the signal difference being less than the first threshold, it is also possible to determine, directly based on the signal difference, whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes, so as to determine whether the howling occurs.
[0094] For example, as shown in FIG. 10, if the amplitude difference between the two acoustic feedback paths is less than the first threshold T, at this time, the no howling flag = True, it is considered that no howling is generated. The signals collected by the microphones may be directly output. And if is not less than the first threshold T, at this time, the no-howling flag = False, it is considered that the howling may be generated. It may be further determined, based on , whether the acoustic feedback path changes to accurately determine whether the howling occurs.
[0095] According to embodiments, the determining whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference may include: extracting a long-time difference feature and a short-time difference feature based on the difference between the first acoustic feedback path and the second acoustic feedback path; determining whether the acoustic feedback path changes based on the long-time difference feature and the short-time difference feature. For example, the long-time difference feature and the short-time difference feature may be extracted using a first artificial intelligence network, based on the difference between the first acoustic feedback path and the second acoustic feedback path. According to embodiments, the long-time difference feature may be a steady-state feature in the case that the change of the acoustic feedback path does not occur, which is also referred to as a steady-state feature in the following. The short-time difference feature may be a transient feature in the case that the change of the acoustic feedback path occurs, which is also referred to as a transient feature in the following.
[0096] For example, as shown in FIG. 10, after acquiring the magnitude difference of the two acoustic feedback paths, the first artificial intelligence network may be utilized to extract the steady-state feature of in the case that the change of the acoustic feedback path does not occur and the transient feature of in the case that the change of the acoustic feedback path occurs. As an example, the first artificial intelligence network may be a two-channel neural network (e.g., a convolutional neural network or a recurrent neural network).
[0097] FIG. 11 is a schematic diagram illustrating a difference between acoustic feedback paths changes with the acoustic feedback path. In the case that the acoustic feedback path does not change, a difference value between H1 and H2 does not change much, as shown by the light colored blocks in FIG. 11. When the acoustic feedback path begins to change, for example, when the earphone is covered by a hand, the difference value between H1 and H2 begins to increase, as shown by the dark colored blocks in FIG. 11. In order to quickly determine whether the acoustic feedback path changes, the steady-state feature when the difference value of the acoustic feedback paths changes over time and the transient feature when the path changes may be extracted. Since the change of the acoustic feedback path is a process rather than a mutation behavior, the steady-state feature and the transient feature may be extracted. It is determined jointly, based on the steady-state feature and the transient feature, whether the acoustic feedback path changes. For example, as shown in FIG. 10, a flag indicating whether the path changes may be output by performing a path decision based on the steady-state feature and the transient feature. For example, the artificial intelligence network may be pre-trained so that the artificial intelligence network learns a relationship between the steady-state feature and the transient feature and whether the acoustic feedback path changes. Later, after extracting the steady-state feature and the transient feature, the steady-state feature and the transient feature may be input into the pre-trained artificial intelligence network to determine whether the acoustic feedback path changes. Through the above method, whether the acoustic feedback path changes may be determined more quickly and accurately.
[0098] Referring back to FIG. 6, at step S630, howling suppression is performed in a case of determining that the acoustic feedback path changes. According to embodiments, the performing of the howling suppression may include: performing the howling suppression in different manners according to whether the first microphone and the second microphone howl simultaneously. According to embodiments, in a case that the first microphone and the second microphone do not howl simultaneously, based on a signal corresponding to a first frequency point at which howling occurs among microphone signals collected by a microphone of the first microphone and the second microphone that does not howl, the howling suppression is performed on a signal corresponding to the first frequency point among microphone signals collected by a microphone of the first microphone and the second microphone that howls; in a case that the first microphone and the second microphone howl simultaneously, the howling suppression is performed by performing wave limiting processing based on a second frequency point at which the howling occurs simultaneously in a first microphone signal and a second microphone signal. FIG. 12 is a schematic diagram illustrating single microphone howling suppression. For example, as shown in FIG. 12, it is possible to detect whether the acoustic feedback path changes based on the difference between the acoustic feedback path H1(n) of microphone 1 and the acoustic feedback path H2(n) of microphone 2. It is possible to perform the howling suppression on the microphone 2 by the signal replacement in time in the case of determining that the microphone 2 howls.
[0099] According to embodiments, the performing of the howling suppression in different manners according to whether the first microphone and the second microphone howl simultaneously may include: determining the first frequency point at which the howling occurs based on the signal difference, and performing fusion processing on the first microphone signal and the second microphone signal to obtain processed microphone signal based on the first microphone signal collected by the first microphone, the second microphone signal collected by the second microphone, and the first frequency point, wherein, in the fusion processing, the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that howls is replaced with the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that does not howl; determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal; in a case of determining that the first microphone and the second microphone do not howl simultaneously, outputting the processed microphone signal as an output signal after performing the howling suppression; in a case of determining that the first microphone and the second microphone howl simultaneously, performing, based on the second frequency point, the wave limiting processing on the processed microphone signal to obtain the output signal after performing the howling suppression. In the present disclosure, the "second frequency point at which the howling occurs" may be a frequency point at which the howling may occur simultaneously, rather than necessarily a frequency point at which howling does occur simultaneously.
[0100] In an embodiment, the determining of the first frequency point at which the howling occurs based on the signal difference includes: determining a frequency point at which the difference between the first acoustic feedback path and the second acoustic feedback path is not less than a second threshold as the first frequency point at which the howling occurs for the microphone of the first microphone and the second microphone that howls.
[0101] For example, as shown in FIG. 10, the first frequency point at which the howling occurs may be determined based on the difference |H1|-|H2| between the first acoustic feedback path and the second acoustic feedback path, in the case of determining that the acoustic path changes. FIG. 13A is a schematic diagram illustrating difference values of acoustic feedback paths in open environment. FIG. 13B is a schematic diagram of determining a first frequency point at which howling occurs. As shown in FIG. 13A and FIG. 13B, the horizontal axis is the frequency point and the vertical axis is the amplitude of the acoustic feedback path. As can be seen from FIG. 13A, a difference value between the acoustic feedback paths of the microphone 1 and the microphone 2 is small in the case that no change of the acoustic feedback path occurs in an open environment. As can be seen from FIG. 13B, The difference value 13 between the acoustic feedback paths of the microphone 1 and the microphone 2 is large in the case that the change of the acoustic feedback path occurs, for example, when the earphone is covered with a hand. According to the difference values 13 of the acoustic feedback path at different frequency points, the first frequency point at which the howling occurs may be determined. As shown in FIG. 13B, the difference value 13 of the acoustic feedback paths is larger at a frequency point at which the howling occurs, and thus, for example, the frequency point at which the difference between the acoustic feedback paths is not less than the second threshold may be determined as the first frequency point at which the howling occurs. For frequency at which howling may occur, the signal corresponding to the frequency of the microphone that howls is replaced with the signal corresponding to the frequency of the microphone that does not howl.
[0102] Referring back to FIG. 7, after determining the first frequency point at which the howling occurs, dual microphone fusion may be performed to obtain the processed microphone signals based on the microphone signal 1, the microphone signal 2, and the first frequency point at which the howling occurs in step S720.Through the fusion processing, the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that howls is replaced with the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that does not howl. According to embodiments, the microphone of the first microphone and the second microphone that howls may be determined based on the first acoustic feedback path and the second acoustic feedback path. For example, the microphone of the first microphone and the second microphone that howls may be determined based on a magnitude of the first acoustic feedback path and a magnitude of the second acoustic feedback path. For example, if the magnitude of the first acoustic feedback path is larger than the magnitude of the second acoustic feedback path, the microphone that howls is determined to be the first microphone, and if the magnitude of the second acoustic feedback path is larger than the magnitude of the first acoustic feedback path, the microphone that howls is determined to be the second microphone. The premise of the dual microphone fusion is to assume that only one microphone howls, i.e., the microphone 1 and the microphone 2 do not howl simultaneously, so that if it is determined that the microphone 1 howls, the signal corresponding to the first frequency point at which the howling occurs of the microphone 1 may be replaced with the signal corresponding to the first frequency point of the microphone 2. If it is determined that the microphone 2 howls, the signal corresponding to the first frequency point of the microphone 2 may be replaced with the signal corresponding to the first frequency point of the microphone 1.
[0103] Usually, the analysis of a howling frequency point is carried out in the frequency domain. If the signal corresponding to the howling frequency point is replaced in the frequency domain, time-frequency transformation is further needed, which will result in a large delay in obtaining the processed microphone signal. In this regard, according to embodiments of the present disclosure, the processed microphone signal is obtained by implementing the replacement of the signal corresponding to the howling frequency point in the time domain. For example, the performing of the fusion processing on the first microphone signal and the second microphone signal to obtain the processed microphone signal includes: encoding the first microphone signal and the second microphone signal respectively in time domain to obtain encoded features respectively corresponding to the first microphone signal and the second microphone signal; fusing the obtained encoded features based on the first frequency point, wherein, in the fusion, encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that howls are replaced with encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that does not howl.
[0104] FIG. 14 is a schematic diagram illustrating dual microphone signal fusion. As shown in FIG. 14, the microphone signal 1 and the microphone signal 2 may be encoded respectively using encoders 1410 and 1420 to obtain the encoded features. After the encoded features are fused according to the first frequency point that howls, by decoding which with a decoder 1430, the processed microphone signal in the time domain may be obtained. Specifically, in the fusion of the encoded features, the encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that howls are replaced with the encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that does not howl. By the replacement of the encoded features, it is realized that the signal corresponding to the first frequency point among the microphone signals collected by the microphone that howls is replaced with the signal corresponding to the first frequency point among the microphone signals collected by the microphone that does not howl. Since this operation does not require the time-frequency transformation, the delay in obtaining the processed microphone signal is small. The listening experience of the user is not affected. The same effect of frequency domain replacement may be achieved.
[0105] The single microphone howling suppression may be achieved by the signal replacement. However, the dual microphone fusion is based on that only single microphone howls, but there is a possibility that the dual microphones howl simultaneously. In this regard, in order to perform the howling suppression more accurately, as shown in FIG. 7, in the case of determining the acoustic feedback path changes, dual microphone howling detection may be performed. In an embodiment, in the step S730, the dual microphone howling detection based on signal features is performed to determine whether the first microphone and the second microphone howl simultaneously and a second frequency point at which the howling occurs.
[0106] According to embodiments, the determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal, may include: extracting signal features of the first microphone signal and the second microphone signal; determining whether the first microphone and the second microphone howl simultaneously and the second frequency point using an artificial intelligence network, based on the signal features. According to embodiments, the signal features may include at least one of energy growth features, a signal similarity feature, and signal envelope features of the first microphone signal and the second microphone signal, but is not limited thereto.
[0107] According to the principle of the howling generation shown in FIG. 4, x(n)=S(n)+Y(n)*H(n). When no howling occurs, the feedback signal Y(n)*H(n) in the two microphone signals is small and S(n) is the same external signal, so that the dual microphone signals x1(n) and x2(n) are very similar. When the howling occurs, the similarity of the dual microphone signals decreases and the signal energy and envelope grow exponentially.
[0108] FIG. 15 is a schematic diagram illustrating a signal envelope feature of a microphone that howls. As shown in FIG. 15, the signal envelope grows exponentially when the howling occurs, and whether the microphone howls may be determined according to the signal envelope. For example, if the signal envelopes of two microphones have an exponential growth trend, it is determined that the two microphones howl simultaneously.
[0109] According to FIG. 4 shows the principle of the howling generation, the speaker signal Y is generated by the external signal S through the gain amplifier module. The speaker signal Y and the feedback path H are coupled to generate the feedback signal Y * H. The feedback signal is also collected by the microphone. The feedback signal is amplified by the amplifier module, and it is cycled as shown in FIG. 4 and amplified repeatedly.
[0110] For example, time 1:
[0111] Time 2:
[0112] Time 3:
[0113] ... ...
[0114] The output of the previous time is put into the input of the next time, the output signal of the speaker at time n is obtained as:
[0115] Time n:
[0116] According to the output expression of the speaker, signal Y(n) contains a exponential gain modulated by the feedback path H, it can be seen that the energy growth feature of the signal when howling occurs is similar to the exponential growth distribution, and whether the microphone howls may be determined based on the energy growth feature.
[0117] FIG. 16 is a schematic diagram illustrating energy growth features of dual microphone signals. As shown in FIG. 16, the energy growth features of both microphones 1 and 2 are similar to the exponential growth distribution, in this case, it may be determined that the two microphones howl simultaneously.
[0118] According to embodiments, the extracting of the signal features of the first microphone signal and the second microphone signal may include: calculating frequency domain energy of each of the first microphone signal and the second microphone signal; based on the frequency domain energy, extracting the energy growth features of the first microphone signal and the second microphone signal, respectively. The determining whether the first microphone and the second microphone howl simultaneously and the second frequency point, based on the signal features, includes: determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching a predetermined energy growth feature, wherein the predetermined energy growth feature corresponds to an energy growth feature when predetermined howling occurs; determining the second frequency point based on the energy growth features. For example, the energy growth feature when the predetermined howling occurs may be the exponential growth, as described above. For example, the determining of the second frequency point based on the energy growth features includes: determining a frequency point as the second frequency point, if a change of energy growth degree of the first microphone and the second microphone at the frequency point is increasing over time and a similarity of the energy growth degree of the first microphone and the second microphone is below a predetermined similarity threshold.
[0119] FIG. 17 is a schematic diagram illustrating dual microphone howling detection based on energy growth features. As shown in FIG. 17, at step S1710, a frequency domain energy calculation is performed for the microphone signal 1 to obtain energy E1. At step S1720, energy growth feature extraction is performed based on the energy E1 to obtain the energy growth feature GR1 of the microphone signal 1. Similarly, at step S1730, the frequency domain energy calculation is performed for the microphone signal 2 to obtain energy E2. At step S1740, the energy growth feature extraction is performed based on the energy E2 to obtain the energy growth feature GR2 of the microphone signal 2. For example, an artificial intelligence network may be utilized to extract the energy growth feature of the microphone signal based on the calculated energy. Whether the microphone 1 and the microphone 2 howl simultaneously may be determined based on the energy growth feature of the microphone signal 1 and the energy growth feature of the microphone 2. For example, if the energy growth features of both the microphone signal 1 and the microphone signal 2 match the exponential growth feature, it is determined that both the microphone 1 and the microphone 2 howl simultaneously. A frequency point at which the howling occurs may be determined based on the energy growth features of the microphone signal 1 and the microphone signal 2. For example, a neural network (e.g., a convolutional neural network) may be pre-trained with the signal features, which may provide a determination result of whether the microphone 1 and the microphone 2 howl simultaneously and the frequency point at which the howling occurs based on the signal features of the microphone 1 and the microphone 2. After obtaining the trained neural network, the energy growth features of the microphone signal 1 and the microphone signal 2 may be input into the neural network to determine whether the microphone 1 and the microphone 2 howl simultaneously and the frequency point at which the howling occurs. As shown in FIG. 17, after determining that the two microphones howl simultaneously and determining the frequency point at which the howling occurs based on the energy growth features, the wave limiter 1760 may be utilized to suppress the signal at the frequency point at which the howling occurs. The output signal after performing the howling suppression may be obtained. For example, a corresponding wave limiter may be designed based on the frequency point at which the howling occurs.
[0120] In addition to using the energy growth feature, whether the two microphones howl simultaneously may be determined based on the signal similarity. For example, according to the above microphone signal expression, it can be seen that the signal composition of the microphone 1 may be expressed as , where the external audio signal S is superimposed with the feedback signal YH1G; and the signal composition of the microphone 2 is , where the external audio signal S is superimposed with the feedback signal YH2G. For the two microphone signals X1and X2, when no howling occurs, the feedback signal YHG is weak in energy, and the microphone signals mainly consist of the external audio signal S. In contrast, for the two microphone signals X1and X2, when howling occurs, the influence of the feedback signal YH1G and YH2G may be increased. The similarity of the two microphone signals X1and X2may be decreased. Therefore, whether the two microphones howl simultaneously may be determined based on the signal similarity feature of the two microphone signals.
[0121] According to embodiments, the signal features may further include the signal similarity feature of the first microphone signal and the second microphone signal in addition to the energy growth features. For example, the extracting of the signal features of the first microphone signal and the second microphone signal further includes: extracting a change trend of a signal similarity of the first microphone signal and the second microphone signal as the signal similarity feature. For example, an artificial intelligence network may be used to learn the change trend of similarity of the two microphone signals, and accordingly determine whether the two microphones howl simultaneously. In an embodiment, the determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching the predetermined energy growth feature, may include: determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching the predetermined energy growth feature and the signal similarity of the first microphone signal and the second microphone signal decreasing. In other words, whether the two microphones howl simultaneously may be determined jointly based on both the energy growth features and the signal similarity feature of the two microphone signals, and in this way, whether the two microphones howl simultaneously may be determined more accurately.
[0122] The dual microphone howling detection according to embodiments of the present disclosure may detect the howling in time when the howling increases to the audible threshold, and then perform the howling suppression, so that the howling is not heard by the user.
[0123] Referring back to FIG. 7, the dual microphone howling detection based on the signal features S730 may determine not only whether the first microphone and the second microphone howl simultaneously, but also determine the second frequency point at which the howling occurs. Here, the second frequency point may be the same as or different from the first frequency point.
[0124] If it is determined that the first microphone and the second microphone do not howl simultaneously, the processed microphone signal obtained by the dual microphone fusion (i.e., a signal obtained by performing the single microphone howling suppression) may be directly output as the output signal after performing the howling suppression. On the contrary, if it is determined that the first microphone and the second microphone howl simultaneously, the wave limiting processing is performed on the processed microphone signal based on the second frequency point to obtain the output signal after performing the howling suppression. According to embodiments, the wave limiting processing may be performed on the processed microphone signal, utilizing a wave limiter designed based on the second frequency point, to obtain the output signal after performing the howling suppression. The signal at the second frequency point may be eliminated or suppressed by performing the wave limiting processing on the processed microphone signal with the wave limiter designed based on the second frequency point.
[0125] Above, the method performed by an electronic apparatus according to embodiments of the present disclosure has been described in connection with examples. According to the method, the electronic apparatus may detect the occurrence of the howling more quickly and accurately. The electronic apparatus may perform the howling suppression in a timely manner in the case that the howling occurrence is detected. The howling is not heard by the user, which improves the listening experience of the user.
[0126] The method according to embodiments of the present disclosure may be applied in many scenarios that require the sound amplification. For example, it may be applied in a hearing aid scenario, in which it may help, by amplifying the sound, the user clearly hear the voice of a distant speaker and improve the quality of communication. By using the method according to embodiments of the present disclosure, the occurrence of the howling may be detected accurately and timely so as to perform the howling suppression in a timely manner, which prevents the user from hearing the howling that affects the listening experience. For example, firstly, a PSAP mode may be turned on in a cell phone. Secondly, a hearing test of the user may be performed, and a personalized amplification function may be provided to the user based on the test result. The hearing test may be performed only when the PSAP function is used for the first time. Subsequently, the method according to embodiments of the present disclosure is enabled to perform the howling suppression. For example, the method may output the signal after the howling suppression based on the microphone signal and the speaker signal on the electronic apparatus. By enabling the method, the howling may be made inaudible to the user while providing the user with a sufficiently high gain to amplify a small sound at a distance.
[0127] Although an application scenario of the method of the embodiments of the present disclosure is illustrated above only by taking the hearing aid scenario as an example, the application scenario of the method of the present disclosure is not limited to the hearing aid scenario, and may be applied to any scenario in which the howling may be generated.
[0128] Above, the method performed by an electronic apparatus according to embodiments of the present disclosure and its application have been described. In the following, the electronic apparatus according to embodiments of the present disclosure is briefly described.
[0129] FIG. 18 is a block diagram illustrating an electronic apparatus according to an embodiment of the present disclosure. Referring to FIG. 18, the electronic apparatus 1800 may include a memory 1810 and a processor 1820, wherein the processor 1820 is coupled to the memory 1810 and configured to perform any of the methods described above. For example, the electronic apparatus may include at least two microphones, which may be earphones, hearing aids, and the like, but are not limited thereto.
[0130] In embodiments of the present disclosure, there is also provided an electronic apparatus that includes at least one processor. In an embodiment, there is also provided an electronic apparatus that further includes at least one transceiver and / or at least one memory coupled to the at least one processor, wherein, the at least one processor is configured to perform the steps of the method provided in any alternative embodiment of the present disclosure.
[0131] FIG. 19 illustrates a schematic diagram of a structure of an electronic apparatus applicable to an exemplary embodiment of the present application. As shown in FIG. 19, the electronic apparatus 4000 shown in FIG. 19 includes: a processor 4001 and a memory 4003. Wherein the processor 4001 and the memory 4003 are coupled, e.g., through a bus 4002. In an embodiment, the electronic apparatus 4000 may further include a transceiver 4004 which may be used for data interaction between the electronic apparatus and other electronic apparatuses, such as transmitting of data and / or receiving of data. It should be noted that, each of the processor 4001, the memory 4003, and the transceiver 4004 is not limited to one in a practice application, and the structure of the electronic apparatus 4000 does not constitute a limitation of the embodiments of the present disclosure. In an embodiment, the electronic apparatus may be the first network node, the second network node, or the third network node.
[0132] The processor 4001 may be a Central Processing Unit (CPU), general purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, transistor logic device, hardware part, or any combination thereof. It may implement or perform various exemplary logic boxes, modules, and circuits described in conjunction with the disclosed contents of the present disclosure. The processor 4001 may also be a combination that implements computing functions, such as a combination containing one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0133] The bus 4002 may include a pathway to transfer information between the above components. The bus 4002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, and the like. The bus 4002 may be classed as an address bus, a data bus, a control bus, and the like. For ease of representation, only one bold line is shown in FIG. 19, but it does not mean that there is only one bus or one type of bus.
[0134] The memory 4003 may be a Read Only Memory (ROM) or other types of static storage apparatuses that can store static information and instructions, a Random Access Memory (RAM) or other types of dynamic storage apparatuses that can store information and instructions, may be an Electrically Erasable Programmable Read Only Memory (EEPROM), Compact Disc Read Only Memory (CD-ROM) or other optical disc storages, an optical disc storage (including a compressed disc, laser disc, optical disc, digital universal disc, Blu-ray disc, etc.), a disk storage medium, other magnetic storage apparatuses, or any other medium that can be used to carry or store computer programs and can be read by a computer, it is not limited herein.
[0135] The memory 4003 is used to store computer programs or executable instructions for performing the embodiments of the present disclosure, and is controlled for execution by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the preceding method of the embodiments.
[0136] Moreover, various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer readable program code. The phrase "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.
[0137] An embodiment of the present disclosure provides a computer readable storage medium storing computer programs or instructions, the computer programs or instructions, when being executed by at least one processor may perform or implement the steps in the preceding method of the embodiments and corresponding contents.
[0138] An embodiment of the present disclosure provides a computer program product including computer programs, the computer programs, when being executed by a processor, may implement the steps shown in the preceding method of the embodiments and corresponding contents.The terms "first", "second", "third", "fourth", "1", "2" and the like (if exists) in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequence. It should be understood that, data used as such may be interchanged in appropriate situations, so that the embodiments of the present disclosure described here may be implemented in an order other than the illustration or text description.
[0139] It should be understood that, although each operation step is indicated by an arrow in the flowcharts of the embodiments of the present disclosure, an implementation order of these steps is not limited to an order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in the flowcharts may be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include a plurality of sub steps or stages, based on an actual implementation scenario. Some or all of these sub steps or stages may be executed at the same time, and each sub step or stage in these sub steps or stages may also be executed at different times. In scenarios with different execution times, an execution order of these sub steps or stages may be flexibly configured according to a requirement, which is not limited by the embodiment of the present disclosure.
[0140] The above text and accompanying drawings are provided as examples only to assist readers in understanding the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the content disclosed herein, it is apparent to those skilled in the art that, changes can be made to the illustrated embodiments and examples without departing from the scope of the present disclosure, and other similar implementation methods based on the technical concepts of the present disclosure also belongs to a protection scope of the embodiments of the present disclosure.
Claims
1.A method performed by an electronic apparatus, the electronic apparatus comprising a first microphone and a second microphone, the method comprising:acquiring a signal difference between the first microphone and the second microphone;determining whether an acoustic feedback path of at least one of the first microphone and the second microphone changes based on the signal difference; andperforming howling suppression in a case of determining that the acoustic feedback path changes.2.The method of claim 1, wherein the acquiring of the signal difference between the first microphone and the second microphone comprises:acquiring a difference between a first acoustic feedback path of the first microphone and a second acoustic feedback path of the second microphone, as the signal difference.3.The method of claim 2, wherein the electronic apparatus further comprises a speaker,wherein the acquiring of the difference between the first acoustic feedback path of the first microphone and the second acoustic feedback path of the second microphone comprising:acquiring the difference between the first acoustic feedback path and the second acoustic feedback path based on a first microphone signal collected by the first microphone, a second microphone signal collected by the second microphone, a speaker signal played by the speaker, and a signal amplification gain corresponding to the speaker.4.The method of any one of claims 2 to 3, further comprising:outputting signals collected by the first microphone and the second microphone if the signal difference is less than a first threshold;wherein the determining whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference comprises:determining that the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference, in a case of the signal difference being not less than the first threshold.5.The method of any one of claims 2 to 4, wherein the determining whether the acoustic feedback path of the at least one of the first microphone and the second microphone changes based on the signal difference comprises:extracting a long-time difference feature and a short-time difference feature based on the difference between the first acoustic feedback path and the second acoustic feedback path;determining whether the acoustic feedback path changes based on the long-time difference feature and the short-time difference feature.6.The method of any one of claims 1 to 5, wherein the performing of the howling suppression comprises:performing the howling suppression based on whether the first microphone and the second microphone howl simultaneously,in a case that the first microphone and the second microphone do not howl simultaneously, based on a signal corresponding to a first frequency point at which howling occurs among microphone signals collected by a microphone of the first microphone and the second microphone that does not howl, performing the howling suppression on a signal corresponding to the first frequency point among microphone signals collected by a microphone of the first microphone and the second microphone that howls;in a case that the first microphone and the second microphone howl simultaneously, performing the howling suppression by performing wave limiting processing based on a second frequency point at which the howling occurs simultaneously in a first microphone signal and a second microphone signal.7.The method of claim 6, wherein the performing of the howling suppression in different manners according to whether the first microphone and the second microphone howl simultaneously comprises:determining the first frequency point at which the howling occurs based on the signal difference,performing fusion processing on the first microphone signal and the second microphone signal to obtain processed microphone signal based on the first microphone signal collected by the first microphone, the second microphone signal collected by the second microphone, and the first frequency point, wherein, in the fusion processing, the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that howls is replaced with the signal corresponding to the first frequency point among the microphone signals collected by the microphone of the first microphone and the second microphone that does not howl;determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal;in a case of determining that the first microphone and the second microphone do not howl simultaneously, outputting the processed microphone signal as an output signal after performing the howling suppression;in a case of determining that the first microphone and the second microphone howl simultaneously, performing, based on the second frequency point, the wave limiting processing on the processed microphone signal to obtain the output signal after performing the howling suppression.8.The method of claim 7, wherein the determining of the first frequency point at which the howling occurs based on the signal difference comprises:determining a frequency point at which the difference between the first acoustic feedback path and the second acoustic feedback path is not less than a second threshold as the first frequency point at which the howling occurs for the microphone of the first microphone and the second microphone that howls.9.The method of any one of claims 7 and 8, wherein the performing of the fusion processing on the first microphone signal and the second microphone signal to obtain the processed microphone signal comprises:encoding the first microphone signal and the second microphone signal respectively in time domain to obtain encoded features respectively corresponding to the first microphone signal and the second microphone signal;fusing the obtained encoded features based on the first frequency point, wherein, in the fusion, encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that howls are replaced with encoded features corresponding to the first frequency point of the microphone signals collected by the microphone that does not howl;decoding the fused encoded features to obtain the processed microphone signal.10.The method of any one of claims 7 to 9, wherein the microphone of the first microphone and the second microphone that howls is determined based on the first acoustic feedback path and the second acoustic feedback path.11.The method of any one of claims 7 to 10, wherein the determining whether the first microphone and the second microphone howl simultaneously and the second frequency point at which the howling occurs, based on the first microphone signal and the second microphone signal, comprises:extracting signal features of the first microphone signal and the second microphone signal;determining whether the first microphone and the second microphone howl simultaneously and the second frequency point using an artificial intelligence network, based on the signal features.12.The method of claim 11, wherein the signal features comprise at least one of energy growth features, a signal similarity feature, and signal envelope features of the first microphone signal and the second microphone signal.13.The method of claim 11, wherein the signal features comprise energy growth features of the first microphone signal and the second microphone signal,wherein the extracting of the signal features of the first microphone signal and the second microphone signal comprises:calculating frequency domain energy of the first microphone signal and the second microphone signal; andbased on the frequency domain energy, extracting the energy growth features of the first microphone signal and the second microphone signal, respectively,wherein the determining whether the first microphone and the second microphone howl simultaneously and the second frequency point, based on the signal features, comprises:determining that the first microphone and the second microphone howl simultaneously, based on the energy growth features of the first microphone signal and the second microphone signal both matching a determined energy growth feature, wherein the determined energy growth feature corresponds to an energy growth feature when howling occurs; anddetermining the second frequency point based on the energy growth features.14.An electronic apparatus comprising:memory storing a program or at least one instruction; andat least one processor coupled to the memory,wherein the at least one processor executes the program or at least one instruction to cause the electronic apparatus to perform the method of any one of claims 1 to 13.15.A computer-readable storage medium having stored thereon a computer program or instructions which, when executed by at least one processor, cause(s) the at least one processor to perform the method of any one of claims 1 to 13.