Whistling Sound Suppression Method, Device, Headphone and Storage Medium

By using neural network models in the headset to detect ear canal audio signals to predict howling events, and using a second filter group to filter the ambient audio signals when the howling is detected, the problem of howling sounds occur in the transparent mode of the headset is solved, and the user experience is improved.

CN114513723BActive Publication Date: 2025-06-10BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210166826.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-06-10
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

Existing headphones are prone to whistling in transparent mode, resulting in poor user experience.

Method used

By acquiring the ear canal audio signal and determining the predicted audio signal using a preset neural network model, if a howling event is detected for the predicted audio signal, the second filter bank is used to filter the subsequent ambient audio signal to suppress the howling sound.

Benefits of technology

Before the howling occurs, the second filter group can be activated to effectively avoid the howling sound, realize the transparent mode without howling, and improve the user's user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114513723B_ABST
    Figure CN114513723B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, earphone and storage medium for suppressing whistling sound. The method for suppressing whistling sound includes: obtaining an environmental audio signal; filtering the environmental audio signal according to a preset first filter bank to obtain a first audio signal; controlling a speaker to play the first audio signal; obtaining an ear canal audio signal; determining a predicted audio signal according to the ear canal audio signal and a preset neural network model; if a whistling event is detected in the predicted audio signal, filtering a subsequently obtained environmental audio signal according to a preset second filter bank to obtain a second audio signal. In this method, the second filter bank can be started in advance before the whistling occurs, so as to better avoid the whistling sound, realize a clear-through mode without whistling sound, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of earphones, and in particular, to a method and device for suppressing whistling sounds, an earphone, and a storage medium. Background Art

[0002] In the field of audio information, there are various earphones for collecting and outputting sound signals. Among them, there are also earphones applied in the transparent mode. The transparent mode means that the earphone collects ambient sound, filters the ambient sound and then outputs it, and superimposes the sound leaking into the ear, so that the human ear receives the complete ambient sound.

[0003] When a user wears an earphone and talks to others, the transparent mode can be switched, which is equivalent to the effect of taking off the earphone, and a clear conversation with the other party can be realized. With the rapid popularization of earphones with the transparent mode, the frequency and duration of user use of earphones are both increasing. The transparent transmission of ambient sound is also being studied in the direction of more accurate and natural listening experience. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a method and device for suppressing whistling sounds, an earphone, and a storage medium.

[0005] According to the first aspect of the embodiments of the present disclosure, there is provided a method for suppressing whistling sounds, which is applied to an earphone. The earphone includes a speaker, a feedforward microphone, and a feedback microphone. The method includes:

[0006] Obtain an ambient audio signal, where the ambient audio signal is a sound signal in the environment around the earphone collected by the feedforward microphone;

[0007] Filter the ambient audio signal according to a preset first filter bank to obtain a first audio signal;

[0008] Control the speaker to play the first audio signal;

[0009] Obtain an ear canal audio signal, where the ear canal audio signal is a sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal;

[0010] Determine a predicted audio signal according to the ear canal audio signal and a preset neural network model;

[0011] If a whistling event is detected in the predicted audio signal, filter the subsequently obtained ambient audio signal according to a preset second filter bank to obtain a second audio signal.

[0012] Optionally, the method further includes:

[0013] The amplitudes of the first frequency response curves corresponding to the coefficients of the second filter bank are all smaller than the amplitudes of the second frequency response curves corresponding to the coefficients of the corresponding first filter bank.

[0014] Optionally, the ear canal audio signal includes the audio signals of m frames up to the current frame, and the predicted audio signal includes the audio signals of n frames, where m and n are integers greater than or equal to 1 respectively.

[0015] Optionally, each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1.

[0016] Determining the predicted audio signal according to the ear canal audio signal and a preset neural network model includes:

[0017] Determining the i sub-audio signals in the predicted audio signal according to the ear canal audio signal and the preset neural network model to determine the predicted audio signal;

[0018] Wherein, when determining the u-th sub-audio signal of the predicted audio signal:

[0019] Removing the first u - 1 sub-audio signals from the ear canal audio signal to obtain a first input signal, where u is an integer greater than or equal to 1 and less than or equal to i;

[0020] Determining the first u - 1 sub-audio signals of the predicted audio signal as a second input signal;

[0021] Determining an input audio signal according to the first input signal and the second input signal;

[0022] Inputting the input audio signal into the neural network model to determine the u-th sub-audio signal of the predicted audio signal.

[0023] Optionally, the neural network model is obtained by the following method:

[0024] Constructing a plurality of training sample pairs, each of the training sample pairs including m * i + 1 sub-audio signal samples, where the first m * i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair;

[0025] Training an original network model according to the plurality of training sample pairs to determine the neural network model.

[0026] Optionally, each sub-audio signal includes at least one audio sampling point.

[0027] According to a second aspect of the embodiments of the present disclosure, a howling suppression device is provided, which is applied to an earphone. The earphone includes a speaker, a feedforward microphone, and a feedback microphone. The device includes:

[0028] An acquisition module, configured to acquire an environmental audio signal, where the environmental audio signal is a sound signal in the environment around the earphone collected by the feedforward microphone;

[0029] A determination module, configured to filter the environmental audio signal according to a preset first filter bank to obtain a first audio signal;

[0030] A control module, configured to control the speaker to play the first audio signal;

[0031] The acquisition module is further configured to acquire an ear canal audio signal, where the ear canal audio signal is a sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal;

[0032] The determination module is further configured to determine a predicted audio signal according to the ear canal audio signal and a preset neural network model;

[0033] It is further configured to, if a howling event is detected in the predicted audio signal, filter a subsequently acquired environmental audio signal according to a preset second filter bank to obtain a second audio signal.

[0034] Optionally,

[0035] The amplitudes of the first frequency response curves corresponding to the coefficients of the second filter bank are all smaller than the amplitudes of the second frequency response curves corresponding to the coefficients of the first filter bank.

[0036] Optionally, the ear canal audio signal includes m frames of audio signals up to the current frame, and the predicted audio signal includes n frames of audio signals, where m and n are integers greater than or equal to 1 respectively.

[0037] Optionally, each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1,

[0038] The determination module is further configured to:

[0039] Determine i sub-audio signals in the predicted audio signal according to the ear canal audio signal and a preset neural network model to determine the predicted audio signal;

[0040] Wherein, when determining the u-th sub-audio signal of the predicted audio signal:

[0041] Remove the first u-1 sub-audio signals from the ear canal audio signal to obtain a first input signal, where u is an integer greater than or equal to 1 and less than or equal to i;

[0042] Determine the first u-1 sub-audio signals of the predicted audio signal as a second input signal;

[0043] Determine an input audio signal based on the first input signal and the second input signal;

[0044] Input the input audio signal into the neural network model to determine the u-th sub-audio signal of the predicted audio signal.

[0045] Optionally, the neural network model is obtained by the following method:

[0046] Construct a plurality of training sample pairs, each of the training sample pairs including m*i+1 sub-audio signal samples, wherein the first m*i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair;

[0047] Train the original network model according to the plurality of training sample pairs to determine the neural network model.

[0048] Optionally, each sub-audio signal includes at least one audio sampling point.

[0049] According to a third aspect of the embodiments of the present disclosure, there is provided a headset, the headset including:

[0050] A processor;

[0051] A memory for storing processor-executable instructions;

[0052] Wherein the processor is configured to execute the method as described in the first aspect.

[0053] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a headset, enabling the headset to execute the method as described in the first aspect.

[0054] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In this method, the subsequent predicted audio signal can be determined based on the ear canal audio signal. After it is determined that there is a howling event in the predicted audio signal, the subsequent environmental audio signal can be filtered to suppress the howling sound in the subsequent environmental audio signal. In this method, the second filter bank can be started before the howling occurs to better avoid the howling sound, achieve a howling-free transparent mode, and improve the user experience.

[0055] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings

[0056] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments in accordance with the present invention, and are used together with the specification to explain the principles of the present invention.

[0057] Figure 1 is a flowchart of a howling suppression method shown according to an exemplary embodiment.

[0058] Figure 1a is a schematic diagram of an original frequency response curve and an original processed frequency response curve shown according to an exemplary embodiment.

[0059] Figure 1b is a schematic diagram of a differential frequency response curve shown according to an exemplary embodiment.

[0060] Figure 1c is a schematic diagram of a differential frequency response curve, a first frequency response curve, and a second frequency response curve shown according to an exemplary embodiment.

[0061] Figure 2 is a block diagram of a howling suppression device shown according to an exemplary embodiment.

[0062] Figure 3 is a block diagram of a headset shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0063] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0064] For the design of a headset with a transparent mode, generally, the headset is measured in a laboratory to design the filter coefficients in the transparent mode. However, in actual production, due to the MIC error and the assembly difference of the structural cavity, for the same filter parameters, the effect of the transparent mode often has certain differences, which may lead to the mismatch of the filter coefficients, and thus howling occurs after the transparent mode is turned on. Herein, MIC is an abbreviation of Microphone, referring to a microphone. The scientific name of the microphone is a microphone, which is a simple device for picking up and transmitting sound, and can convert the sound signal into an electrical signal, commonly known as a microphone.

[0065] In the related art, generally, the ear canal audio signal collected by the feedback microphone is first subjected to howling detection. If a howling event is detected, the gain is adjusted, and then the adjusted audio signal is subjected to howling detection again. If a howling event is detected, the gain is adjusted again, and the adjusted audio signal is output. In this method, the howling event can be detected only when there is already a howling event in the collected ear canal audio signal, and then processing can be performed. However, at this time, the howling sound has already been generated and caused a bad experience to the user, resulting in a poor user experience.

[0066] The present disclosure provides a howling sound suppression method applied to headphones. In this method, the subsequent predicted audio signal can be determined based on the ear canal audio signal. After it is determined that there is a howling event in the predicted audio signal, the subsequent ambient audio signal can be filtered to suppress the howling sound in the subsequent ambient audio signal. In this method, the second filter bank can be started before the howling occurs to better avoid the howling sound, achieve a howling-free transparent mode, and improve the user experience.

[0067] In an exemplary embodiment, a howling sound suppression method applied to headphones is provided. The headphones include a speaker, a feedforward microphone, and a feedback microphone. Refer to Figure 1 As shown, the method includes:

[0068] S110. Obtain an ambient audio signal;

[0069] S120. Filter the ambient audio signal according to a preset first filter bank to obtain a first audio signal;

[0070] S130. Control the speaker to play the first audio signal;

[0071] S140. Obtain an ear canal audio signal;

[0072] S150. Determine a predicted audio signal according to the ear canal audio signal and a preset neural network model;

[0073] S160. If a howling event is detected in the predicted audio signal, filter the subsequently obtained ambient audio signal according to a preset second filter bank to obtain a second audio signal.

[0074] In step S110, the ambient audio signal is a sound signal in the environment around the headphones collected by the feedforward microphone.

[0075] Wherein, the user can turn on the transparent mode of the headphones through the corresponding function button, or can also turn on the transparent mode by voice control, which is not limited here. When the headphones are in the transparent mode, after the feedforward microphone collects the sound signal in the environment around the headphones, it can be transmitted to the processor of the headphones so that the processor can obtain the ambient audio signal.

[0076] In step S120, the first filter bank is used to filter the ambient audio signal to better achieve the transparent experience of the headset. After the first filter bank filters the ambient audio signal, the first audio signal can be obtained, and then the first audio signal is transmitted to the processor so that the processor can obtain the first audio signal.

[0077] In step S130, after the processor of the headset obtains the first audio signal, it can transmit it to the speaker of the headset, and the speaker can play the first audio signal so that the user can obtain a transparent experience.

[0078] In step S140, the in-ear canal audio signal is the sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal.

[0079] Among them, after the speaker plays the first audio signal, the first audio signal can propagate in the ear canal, the feedback microphone of the headset can collect the sound signal propagating in the ear canal, and the feedback microphone can obtain the in-ear canal audio signal. After the feedback microphone collects the in-ear canal audio signal, it can transmit it to the processor of the headset so that the processor can obtain the in-ear canal audio signal.

[0080] In step S150, the neural network model can be preset in the headset. It can be set before the headset leaves the factory or after the headset leaves the factory. And after the neural network model is set, it can be modified subsequently to better meet the different needs of users and further improve the user experience.

[0081] Among them, the in-ear canal audio signal can be input into the neural network model, and the neural network model can output the predicted audio signal. Of course, the predicted audio signal can also be determined by other means, which is not limited herein.

[0082] In some embodiments, the processor of the headset presets a neural network model. In the processor, the in-ear canal audio signal collected by the feedback microphone can be used as the input audio signal of the neural network model, and then the input audio signal is input into the neural network model. After the neural network model processes the input audio signal, it can output the predicted audio signal, so that the processor can determine the subsequent predicted audio signal.

[0083] In step S160, howling, in essence, is a kind of feedback sound, which is mainly caused by the self-excitation of energy due to problems such as too close a distance between the sound source and the sound amplification device. For example, when a microphone and a speaker are used simultaneously, the sound played back by the audio device can be transmitted through space to the microphone and the sound energy emitted by the speaker is large enough, and the sound pickup sensitivity of the microphone is high enough, and so on. Howling is highly harmful, not only making the user experience worse, but more seriously, it is likely to damage the headphones and harm the user's hearing.

[0084] Among them, methods such as frequency-domain peak detection and energy detection can be used to determine whether there is howling in the predicted audio signal.

[0085] In some embodiments, if it is determined that the local peak-valley difference within the preset local frequency band range where the full-band peak point of the predicted audio signal is located satisfies the first condition, and it is determined that the amplitude change situation of the predicted audio signal satisfies the second condition, then it is determined that there is howling in the predicted audio signal.

[0086] Among them, the full-band peak point is: the frequency point with the largest amplitude within the full frequency band; the local peak-valley difference is: the amplitude difference between the frequency point with the largest amplitude and the frequency point with the smallest amplitude in the predicted audio signal within the preset local frequency band range.

[0087] Among them, the full frequency band range can be 0 - 24 kHz; the preset local frequency band range can be set according to the actual situation. In some embodiments, the preset local frequency band range can be: ±1000 Hz.

[0088] Among them, the frequency-domain characteristics and time-domain characteristics of the predicted audio signal can be obtained, and according to the frequency-domain characteristics of the predicted audio signal, it is determined whether the local peak-valley difference within the preset local frequency band range where the full-band peak point of the predicted audio signal is located satisfies the first condition; when the local peak-valley difference within the preset local frequency band range where the full-band peak point of the predicted audio signal is located satisfies the first condition, further according to the time-domain characteristics of the predicted audio signal, it is determined whether the amplitude change situation of the predicted audio signal satisfies the second condition. If the predicted audio signal satisfies the second condition, then it is determined that there is howling in the predicted audio signal.

[0089] Among them, the first condition may include:

[0090] The local peak-valley difference is greater than the first threshold;

[0091] In the embodiments of the present disclosure, by comparing the local peak-valley difference and the first threshold, if the local peak-valley difference is greater than the first threshold, it is determined that the local peak-valley difference within the preset local frequency band range where the full-band peak point of the predicted audio signal is located satisfies the first condition.

[0092] Among them, the value range of the first threshold can be: 25 dB - 35 dB; in some embodiments, the first threshold can be 30 dB.

[0093] It should be noted that in the spectrogram of the howling sound, the amplitude corresponding to a single and fixed howling frequency point is much larger than the amplitudes of other frequency points in the audio signal; therefore, if the local peak-valley difference within the preset local frequency band where the full-band peak point of the predicted audio signal is located is greater than the first threshold, it is determined that the first condition is satisfied.

[0094] Among them, the second condition may include:

[0095] The changing trend of the amplitude of the predicted audio signal is that the amplitude is gradually increasing.

[0096] In the embodiments of the present disclosure, the changing trend of the amplitude of the predicted audio signal can be determined according to the time-domain characteristics of the predicted audio signal. If the changing trend of the amplitude of the predicted audio signal is that the amplitude is gradually increasing (i.e., showing a growth trend), it is determined that the changing situation of the amplitude of the predicted audio signal satisfies the second condition.

[0097] In practical applications, the following method can be used to determine whether the changing trend of the amplitude of the predicted audio signal is that the amplitude is gradually increasing: calculate the amplitude energy of each frame of data, and statistically analyze the amplitude energy data of multiple frames of data; perform linear regression calculation based on the amplitude energy data of multiple frames of data, and determine whether the slope is greater than 0 according to the linear regression result; if the slope is greater than 0, it indicates that the changing trend of the amplitude of the predicted audio signal is that the amplitude is gradually increasing; otherwise, it indicates that the changing trend of the amplitude of the predicted audio signal is that the amplitude is gradually decreasing.

[0098] Among them, the second condition may also include:

[0099] The changing trend of the amplitude of the predicted audio signal within the preset time range is that the amplitude is gradually increasing.

[0100] Among them, the preset time range can be set according to actual needs.

[0101] It should be noted that the time-domain waveform of the howling sound is a sine wave with a relatively constant frequency, and its amplitude will increase rapidly with the passage of time. When it exceeds the power amplifier amplification region and enters the saturation region and cut-off region, a clipping phenomenon occurs. Therefore, the amplitude of the howling sound shows a growth trend within a certain time range.

[0102] Among them, the method of acoustic event detection can also be used to determine whether there is a howling sound. That is, the howling sound is regarded as an acoustic event, denoted as a howling event, and then the method of acoustic event detection is used to judge whether there is a howling event. If it is judged that there is a howling event, it means that a howling sound will occur when playing the to-be-played audio signal corresponding to the predicted audio signal; if it is judged that there is no howling event, it means that a howling sound will not occur when playing the to-be-played audio signal corresponding to the predicted audio signal. Exemplarily, in a manner based on a deep learning model, a convolutional neural network can be used to judge whether there is a howling event in the predicted audio signal.

[0103] Of course, it is also possible to determine whether there is a howling event in the predicted audio signal by a method, and no limitation is made thereto.

[0104] Among them, two filter banks are set in the earphone, denoted as the second filter bank and the first filter bank respectively. The coefficients of the second filter bank and the first filter bank (i.e., filter bank coefficients) are different, so as to achieve different filtering processes. The number of filters in both the first filter bank and the second filter bank can be 6. Both the first filter bank and the second filter bank include 6 cascaded filters. Both the second filter and the first filter include gain values. The gain value of each first filter is less than the gain value of the corresponding second filter, which can enable the second filter bank to not only filter out the audio signal corresponding to the environmental sound leaking into the ear bypassing the earphone in the environmental audio signal, but also filter the interference signal causing the earphone to howl in the first audio signal, thereby suppressing the earphone from generating a howling sound.

[0105] Among them, the coefficient of the first filter bank can be denoted as the first filtering coefficient, and the coefficient of the second filter bank can be denoted as the second filtering coefficient.

[0106] The frequency response curve corresponding to the second filtering coefficient can be denoted as the second frequency response curve. After the environmental audio signal is filtered by the second filter bank, a first audio signal is obtained, and the frequency response curve of the first audio signal can be denoted as the first processed frequency response curve. The difference between the original frequency response curve and the first processed frequency response curve is the second frequency response curve.

[0107] The frequency response curve corresponding to the first filtering coefficient can be denoted as the first frequency response curve, and the frequency response curve of the environmental audio signal can be denoted as the original frequency response curve. After the environmental audio signal is filtered by the first filter bank, a second audio signal is obtained, and the frequency response curve of the second audio signal can be denoted as the second processed frequency response curve. The difference between the original frequency response curve and the second processed frequency response curve is the first frequency response curve.

[0108] Among them, at any frequency, the amplitude of the second frequency response curve is smaller than that of the first frequency response curve. In this earphone, the first filter bank can only perform transparent filtering on the ambient audio signal to realize the function of the transparent mode. The second filter bank can perform transparent filtering and howling filtering on the ambient audio signal, not only can realize the function of the transparent mode, but also can effectively suppress the occurrence of howling, further improving the user experience.

[0109] Exemplarily, the first frequency response curve and the second frequency response curve can be determined in the following manner.

[0110] Refer to Figures 1a to 1c As shown, before each earphone is put on the market, it is necessary to first measure the acoustic characteristics of the prototype in an anechoic chamber. Through an artificial head, the original audio signal when the ear is empty can be collected, so as to obtain the original frequency response curve (refer to the curve A shown in Figure 1a ). Wear the earphone on the artificial head, so as to collect the processed audio signal after passive noise reduction when wearing the earphone. The frequency response curve of this processed audio signal is recorded as the original processed frequency response curve (refer to the curve B shown in Figure 1a ). By comparing curve A and curve B, the differential frequency response curve is obtained (refer to the curve C shown in Figure 1b and 1c ). Curve C represents the difference between curve A and curve B.

[0111] In this example, 6 cascaded second-order IIR filters can be used to approximate curve C (the frequency range of general concern is 1 kHz to 6 kHz). The exemplary steps are as follows: First, each IIR filter has a random initialization value (initial filter coefficient), then randomly update the frequency, gain value, and Q value, so as to update the filter coefficient, and then calculate the curve D corresponding to the updated filter coefficient (for example, refer to the curve shown in Figure 1c ). Compare the difference between curve D and curve C. If the difference between curve D and curve C is smaller than the previous difference, then continue to update the frequency, gain value, and Q value based on the current filter coefficient. And so on, perform multiple iterations until the difference between curve D and curve C stabilizes, so as to determine that the stabilized curve D is the first frequency response curve, and these 6 cascaded second-order IIR filters constitute the first filter bank.

[0112] Based on curve D, a filter bank with a decreasing average amplitude is designed, denoted as the second filter bank. Among them, the second filter coefficient of the second filter bank is different from the first filter coefficient, so that at any frequency, the amplitude of the second frequency response curve (refer to the curve E shown in Figure 1c ) corresponding to the second filter coefficient is smaller than the amplitude of the first frequency response curve. It should be noted that generally, the gain value of the second filter bank is smaller than the gain value of the first filter bank.

[0113] In some embodiments, the gain value of each first filter is 1 / 3 of the gain value of the corresponding second filter. 1 / 3 is an empirical value obtained through multiple experiments in this application.

[0114] In this application, when the number of filters in the first filter bank and the second filter bank changes, the gain values, frequency values, and Q values corresponding to each filter can be flexibly adjusted.

[0115] In some embodiments, the frequency value of each of the first filters is equal to the frequency value of the corresponding second filter, and the Q value of each of the first filters is equal to the Q value of the corresponding second filter.

[0116] It should be noted that the Q value represents the quality factor. Q value = center frequency ÷ filter bandwidth. The larger the Q value, the narrower the filter bandwidth, and the smaller the Q value, the wider the filter bandwidth.

[0117] In this embodiment, the filtering bandwidths of the filters in the first filter bank are basically the same as the filtering bandwidths of the corresponding filters in the second filter bank. For example, the bandwidth of the sixth filter in the first filter bank is the same as the bandwidth of the sixth filter in the second filter bank, the bandwidth of the fifth filter in the first filter bank is the same as the bandwidth of the fifth filter in the second filter bank, and so on, so that the first filter bank and the second filter bank have the same filtering bandwidth for audio signals with the same center frequency, which is beneficial to the processing of environmental audio signals in the same bandwidth.

[0118] Among them, before the headphones leave the factory, the first filter coefficient and the second filter coefficient are burned into the headphone storage component, or the filter coefficients can be updated to the headphones through subsequent upgrades. The storage component can be a read-only memory (ROM) or a flash memory. When the processor of the headphones needs to use the first filter coefficient or the second filter coefficient, it can directly extract them from the storage component.

[0119] In this step, since a howling event is detected in the pre-stored audio signal, it indicates that there is howling in the subsequent environmental audio signal, and then the second filter bank can be controlled to filter the subsequent audio signal, which can not only ensure the transparency effect of the transparent mode but also suppress the howling sound in the audio signal to be played.

[0120] Among them, "subsequent" refers to the time point after obtaining the detection result of the howling event.

[0121] For example, the first filter bank filters the first environmental audio signal to obtain a first audio signal. When the speaker plays the first audio signal, the feedback microphone collects the sound signal in the ear canal to obtain an ear canal audio signal. A predicted audio signal is determined based on the ear canal audio signal. If it is determined that a howling event is detected in the predicted audio signal, the second filter bank is used to filter other environmental audio signals after the first environmental audio signal.

[0122] In some embodiments, the ear canal audio signal includes the audio signals of m frames up to the current frame, and the predicted audio signal includes the audio signals of n frames, where m and n are integers greater than or equal to 1 respectively.

[0123] Among them, the feedback microphone can collect the sound signal in the ear canal frame by frame to determine the audio signal of each frame. Among them, the audio signals of m frames up to the current frame can be determined as the ear canal audio signal. The earphone can determine the audio signals of n frames after the current frame based on the audio signals of m frames, and determine the determined n-frame audio signals as the predicted audio signal.

[0124] Among them, if a howling event is detected in the predicted audio signal, the second filter bank can be controlled to filter the subsequent n-frame environmental audio signals to obtain a second audio signal, and then the speaker is controlled to play the second audio signal to ensure the user's clear experience and avoid howling.

[0125] Among them, n can be less than or equal to m. It can be understood that the larger m is and the smaller n is, the more sub-audio signals the ear canal audio signal includes and the fewer sub-audio signals the predicted audio signal includes. Based on the characteristics of the neural network model, the prediction result is more accurate.

[0126] It should be noted that if it is determined that no howling event is detected in the predicted audio signal, it means that no howling will occur subsequently, and the first filter bank can be continuously controlled to filter the subsequent environmental audio signals to ensure the clear effect of the clear mode.

[0127] In this method, the subsequent predicted audio signal can be determined based on the ear canal audio signal. After it is determined that a howling event exists in the predicted audio signal, the subsequent environmental audio signals can be filtered to suppress the howling sound in the subsequent environmental audio signals. In this method, the second filter bank can be started before the howling occurs to better avoid the howling sound, realize a howling-free clear mode, and improve the user experience.

[0128] In an exemplary embodiment, a method for suppressing whistling is provided, which is applied to an earphone. The earphone includes a speaker, a feedforward microphone, and a feedback microphone. In this method, each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1. That is, each frame of ear canal audio signal includes i sub-audio signals, each frame of predicted audio signal also includes i sub-audio signals, and each frame of ambient audio signal also includes i sub-audio signals.

[0129] It should be noted that when the feedback microphone collects the ear canal audio signal, the number of sampling points included in each frame of audio signal can be denoted as h, where h is an integer greater than or equal to 1. That is, each frame of audio signal can collect h audio sampling points. That is, h is related to the sampling frequency of the feedback microphone. For example, h can be 48. Each sub-audio signal may include at least one audio sampling point. In each frame of audio signal, the i sub-audio signals may include h audio sampling points.

[0130] In this method, determining the predicted audio signal according to the ear canal audio signal and a preset neural network model may include:

[0131] S210. Determine the i sub-audio signals in the predicted audio signal according to the ear canal audio signal and the preset neural network model to determine the predicted audio signal.

[0132] Among them, the i sub-audio signals can be determined one by one, or all i sub-audio signals can be directly determined.

[0133] In some embodiments, the ear canal audio signal may include 4 frames of audio signals, each frame of audio signal may include 48 sub-audio signals, and each sub-audio signal includes one audio sampling point. In this embodiment, the 4 * 48 sub-audio signals up to the current frame can be directly input into the neural network model, and the neural network model can output 48 predicted sub-audio signals, and these 48 predicted sub-audio signals can form the predicted audio signal.

[0134] In this method, each frame of audio signal may include h audio sampling points, and the h audio sampling points are divided into i sub-audio signals, and then based on m * i sub-audio signals, the i sub-audio signals of the next frame are predicted to obtain the predicted audio signal, which can improve the prediction accuracy and further improve the user experience.

[0135] In an exemplary embodiment, a method for suppressing whistling is provided, which is applied to an earphone. The earphone includes a speaker, a feedforward microphone, and a feedback microphone. Among them, when determining the u-th sub-audio signal of the predicted audio signal, this method may include:

[0136] S310. Remove the first u - 1 sub - audio signals from the ear canal audio signal to obtain a first input signal, where u is an integer greater than or equal to 1 and less than or equal to i.

[0137] S320. Determine the first u - 1 sub - audio signals of the predicted audio signal as a second input signal.

[0138] S330. Determine an input audio signal based on the first input signal and the second input signal.

[0139] S340. Input the input audio signal into a neural network model to determine the u - th sub - audio signal of the predicted audio signal.

[0140] Among them, the order of the sub - audio signals in the ear canal audio signal can be determined according to the sampling order of the feedback microphone.

[0141] When determining the u - th sub - audio signal of the predicted audio signal, the first u - 1 sub - audio signals among the m * i sub - audio signals in the ear canal audio signal can be removed, and the remaining sub - audio signals form the first input signal. And the first u - 1 sub - audio signals that have been determined in the predicted audio signal are determined as the second input signal. Then, the first input signal and the second input signal are sequentially combined to form the input audio signal. Finally, the input audio signal is input into the neural network model, and the neural network model can output the u - th sub - audio signal of the predicted audio signal.

[0142] It should be noted that when determining the first sub - audio signal of the predicted audio signal, the first 0 sub - audio signals can be removed from the ear canal audio signal to obtain the first input signal, and the first 0 sub - audio signals of the predicted audio signal can be determined as the second input signal. Then, the first input signal and the second input signal can be determined as the input audio signal. That is, when determining the first sub - audio signal of the predicted audio signal, the ear canal audio signal can be directly used as the input audio signal.

[0143] In some embodiments, m can be 4, and i and h are 48 respectively, that is, the ear canal audio signal includes 4 frames of audio signals, each frame of audio signal includes 48 sub - audio signals, and each sub - audio signal includes 1 audio sampling point. In this embodiment, when the headset is in the transparent mode, the feedback microphone can collect the audio signals of each frame in the ear canal at a frequency of 48 times per frame. That is, 48 audio sampling points are collected for each frame of audio signal. The feedback microphone can transmit the collected audio sampling points to the processor of the headset.

[0144] When selecting a filter bank for the environmental audio signal to process the next frame, the processor may determine the audio signals of the current frame and the 3 frames before the current frame as the ear canal audio signals, and then determine them as the input audio signals, which can be denoted as the first input audio signals. That is, the consecutive 4 * 48 sub-audio signals up to the current frame are determined as the first input audio signals.

[0145] The processor may input the above first input audio signals into the neural network model, and the neural network model can output 1 predicted sub-audio signal, and the processor can determine the first sub-audio signal of the predicted audio signal for the next frame.

[0146] After determining the first sub-audio signal of the predicted audio signal, the processor may remove the first sub-audio signal in the first input audio signals to obtain the first input signal, then determine the first sub-audio signal in the predicted audio signal as the second input signal, and add the second input signal to the rear side of the first input signal to form a new input audio signal, which can be denoted as the second input audio signal. This second input audio signal includes the subsequent (4 * 48 - 1) sub-audio signals in the ear canal audio signals and the first previous sub-audio signal in the predicted audio signal.

[0147] Then, the processor may input the second input audio signal into the neural network model, and the neural network model can output 1 predicted sub-audio signal, which can be determined as the next sub-audio signal in the predicted audio signal. That is, this sub-audio signal can be determined as the second sub-audio signal in the predicted audio signal.

[0148] After determining the second sub-audio signal of the predicted audio signal, the processor may remove the first sub-audio signal in the second input audio signal to obtain the first input signal, then determine the above second predicted sub-audio signal as the second input signal, and add the second input signal after the first input signal to form a new input audio signal, which can be denoted as the third input audio signal. This third input audio signal includes the subsequent (4 * 48 - 2) sub-audio signals in the ear canal audio signals and the first 2 sub-audio signals in the predicted audio signal.

[0149] Then, the processor may input the third input audio signal into the neural network model, and the neural network model can output 1 predicted sub-audio signal, which can be determined as the next sub-audio signal in the predicted audio signal. That is, this sub-audio signal can be determined as the third sub-audio signal in the predicted audio signal.

[0150] And so on, until the i-th sub-audio signal in the predicted audio signal is determined. At this point, the i sub-audio signals in the predicted audio signal can be determined, and thus the entire predicted audio signal is obtained.

[0151] In this method, when determining the predicted audio signal, the next sub-audio signal in the predicted audio signal is determined based on the sub-audio signal of the determined predicted audio signal, so as to better ensure the reliability of the predicted audio signal and improve the reliability of this method.

[0152] In an exemplary embodiment, a method for suppressing whistling sound is provided, which is applied to an earphone. The earphone includes a speaker, a feedforward microphone, and a feedback microphone. In this method, the neural network model can be obtained in the following manner:

[0153] S410. Construct a plurality of training sample pairs, each training sample pair includes m*i + 1 sub-audio signal samples. Among them, the first m*i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair;

[0154] S420. Train the original network model according to the plurality of training sample pairs to determine the neural network model.

[0155] In step S410, the type of the sub-audio signal sample in this step is the same as that of the sub-audio signal in step S310. That is, if the number of audio sampling points included in the sub-audio signal is the same as the number of audio sampling point samples included in the sub-audio signal sample.

[0156] Exemplarily, the audio signal in the ear canal can be collected by the feedback microphone at a sampling frequency of h audio sampling points per frame, and then each frame of the collected audio signal is divided into i sub-audio signals as a sub-audio signal sample. A training sample pair is determined according to the collected continuous m + 1 frames of audio signal samples. Among them, the first m*i sub-audio signals are determined as the input sample, and the last 1 sub-audio signal sample is determined as the output sample.

[0157] It should be noted that in this step, the sub-audio signal samples can be collected by experiments or downloaded from the network, and there is no limitation on this.

[0158] In step S420, the original network model can include an LSTM (Long short-term memory) network model.

[0159] In this step, a plurality of training sample pairs can be used to train the LSTM (Long short-term memory) network model to obtain the neural network model.

[0160] This method can obtain an excellent neural network model. Through this neural network model, the next sub-audio signal can be more accurately determined based on continuous m*i sub-audio signals, so that the next frame of audio signal can be more accurately determined according to continuous m frames of audio signals, thereby improving the reliability of this method, better avoiding howling, and enhancing the user experience.

[0161] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. This device is used to implement the above method. Exemplarily, as shown in Figure 2 This device may include an acquisition module 101, a determination module 102, and a control module 103. During the process of implementing the above method,

[0162] The acquisition module 101 is used to acquire an environmental audio signal, where the environmental audio signal is the sound signal in the environment around the headset collected by the feedforward microphone;

[0163] The determination module 102 is used to filter the environmental audio signal according to a preset first filter bank to obtain a first audio signal;

[0164] The control module 103 is used to control the speaker to play the first audio signal;

[0165] The acquisition module 101 is further used to acquire an ear canal audio signal, where the ear canal audio signal is the sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal;

[0166] The determination module 102 is further used to determine a predicted audio signal according to the ear canal audio signal and a preset neural network model;

[0167] It is further used to, if a howling event is detected in the predicted audio signal, filter the subsequently acquired environmental audio signal according to a preset second filter bank to obtain a second audio signal.

[0168] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. In this device, the amplitudes of the first frequency response curves corresponding to the coefficients of the second filter bank are all smaller than the amplitudes of the second frequency response curves corresponding to the coefficients of the corresponding first filter bank.

[0169] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. In this device, the ear canal audio signal includes m frames of audio signals up to the current frame, and the predicted audio signal includes n frames of audio signals, where m and n are integers greater than or equal to 1 respectively.

[0170] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. In this device, each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1.

[0171] Refer to Figure 2 As shown, the determination module 102 is further configured to:

[0172] Determine i sub-audio signals in the predicted audio signal according to the ear canal audio signal and a preset neural network model to determine the predicted audio signal.

[0173] Wherein, when determining the u-th sub-audio signal of the predicted audio signal:

[0174] Remove the first u - 1 sub-audio signals from the ear canal audio signal to obtain a first input signal, where u is an integer greater than or equal to 1 and less than or equal to i.

[0175] Determine the first u - 1 sub-audio signals of the predicted audio signal as the second input signal.

[0176] Determine the input audio signal according to the first input signal and the second input signal.

[0177] Input the input audio signal into the neural network model to determine the u-th sub-audio signal of the predicted audio signal.

[0178] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. In this device, the neural network model is obtained in the following manner:

[0179] Construct a plurality of training sample pairs. Each training sample pair includes m * i + 1 sub-audio signal samples. Among them, the first m * i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair.

[0180] Train the original network model according to the plurality of training sample pairs to determine the neural network model.

[0181] In an exemplary embodiment, a howling suppression device is provided, which is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. In this device, each sub-audio signal includes at least one audio sampling point.

[0182] In an exemplary embodiment, a headset is provided. The headset includes a speaker, a feedforward microphone, and a feedback microphone. The headset may include a second filter bank and a first filter bank. The headset can be a wireless headset or a wired headset, and this is not limited.

[0183] Reference Figure 3 As shown, the earphone 400 may further include one or more of the following components: a processing component 402, a memory 404, a power component 406, a multimedia component 408, an audio signal component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.

[0184] The processing component 402 generally controls the overall operation of the earphone 400, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 402 may include one or more modules to facilitate the interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.

[0185] The memory 404 is configured to store various types of data to support the operation of the earphone 400. Examples of such data include instructions for any application or method operating on the earphone 400, contact data, phone book data, messages, pictures, videos, etc. The memory 404 may be implemented by any type of volatile or non-volatile storage earphone or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0186] The power component 406 provides power to various components of the earphone 400. The power component 406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the earphone 400.

[0187] The multimedia component 408 includes a screen that provides an output interface between the earphone 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 408 includes a front camera module and / or a rear camera module. When the earphone 400 is in an operating mode, such as a shooting mode or a video mode, the front camera module and / or the rear camera module can receive external multimedia data. Each of the front camera module and the rear camera module can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0188] The audio signal component 410 is configured to output and / or input audio signals. For example, the audio signal component 410 includes a microphone (MIC) that is configured to receive external audio signals when the earphone 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio signal component 410 further includes a speaker for outputting audio signals.

[0189] The I / O interface 412 provides an interface between the processing component 402 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0190] The sensor component 414 includes one or more sensors for providing a status assessment of various aspects of the earphone 400. For example, the sensor component 414 can detect the on / off state of the earphone 400, the relative positioning of components, such as the display and keypad of the earphone 400. The sensor component 414 can also detect a change in the position of the earphone 400 or a component of the earphone 400, the presence or absence of user contact with the earphone 400, the orientation or acceleration / deceleration of the earphone 400, and the temperature change of the earphone 400. The sensor component 414 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 414 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 414 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0191] The communication component 416 is configured to facilitate wired or wireless communication between the earphone 400 and other earphones. The earphone 700 can access a communication standard-based wireless network, such as WiFi, 2G, 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0192] In an exemplary embodiment, the earphone 400 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0193] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 404 including instructions, is also provided. The above instructions can be executed by the processor 420 of the earphone 400 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the storage medium are executed by the processor of the earphone, the earphone is enabled to execute the method shown in the above embodiment.

[0194] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0195] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0196] It should be understood that the present invention is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for suppressing whistling sound, applied to an earphone, the earphone including a speaker, a feedforward microphone, and a feedback microphone, characterized in that, the method includes: acquiring an environmental audio signal, where the environmental audio signal is a sound signal in the environment around the earphone collected by the feedforward microphone; filtering the environmental audio signal according to a preset first filter bank to obtain a first audio signal; controlling the speaker to play the first audio signal; acquiring an ear canal audio signal, where the ear canal audio signal is a sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal; determining a predicted audio signal according to the ear canal audio signal and a preset neural network model; if a whistling event is detected in the predicted audio signal, filtering a subsequently acquired environmental audio signal according to a preset second filter bank to obtain a second audio signal; each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1, the determining the predicted audio signal according to the ear canal audio signal and the preset neural network model includes: determining the i sub-audio signals in the predicted audio signal according to the ear canal audio signal and the preset neural network model to determine the predicted audio signal; wherein, when determining the u-th sub-audio signal of the predicted audio signal: removing the first u - 1 sub-audio signals from the ear canal audio signal to obtain a first input signal, u being an integer greater than or equal to 1 and less than or equal to i; determining the first u - 1 sub-audio signals of the predicted audio signal as a second input signal; determining an input audio signal according to the first input signal and the second input signal; inputting the input audio signal into the neural network model to determine the u-th sub-audio signal of the predicted audio signal.

2. The method according to claim 1, characterized in that, the method further includes: the amplitudes of the first frequency response curves corresponding to the coefficients of the second filter bank are all smaller than the amplitudes of the second frequency response curves corresponding to the coefficients of the first filter bank.

3. The method according to claim 1 or 2, characterized in that, the ear canal audio signal includes m frames of audio signals up to the current frame, and the predicted audio signal includes n frames of audio signals, where m and n are respectively integers greater than or equal to 1.

4. The method according to claim 3, characterized in that, the neural network model is obtained in the following manner: constructing a plurality of training sample pairs, each training sample pair including m * i + 1 sub-audio signal samples, where the first m * i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair; training an original network model according to the plurality of training sample pairs to determine the neural network model.

5. The method according to claim 3, characterized in that, each sub-audio signal includes at least one audio sampling point.

6. A howling suppression device is applied to a headset. The headset includes a speaker, a feedforward microphone, and a feedback microphone. Characterized in that, The device includes: An acquisition module for acquiring an environmental audio signal, where the environmental audio signal is a sound signal in the environment around the headset collected by the feedforward microphone; A determination module for filtering the environmental audio signal according to a preset first filter bank to obtain a first audio signal; A control module for controlling the speaker to play the first audio signal; The acquisition module is further configured to acquire an ear canal audio signal, where the ear canal audio signal is a sound signal collected by the feedback microphone when the first audio signal is played by the speaker and propagates in the ear canal; The determination module is further configured to determine a predicted audio signal according to the ear canal audio signal and a preset neural network model; It is further configured to, if a howling event is detected in the predicted audio signal, filter a subsequently acquired environmental audio signal according to a preset second filter bank to obtain a second audio signal; Each frame of audio signal includes i sub-audio signals, where i is an integer greater than or equal to 1. The determination module is further configured to: Determine i sub-audio signals in the predicted audio signal according to the ear canal audio signal and a preset neural network model to determine the predicted audio signal; Wherein, when determining the u-th sub-audio signal of the predicted audio signal: Remove the first u - 1 sub-audio signals from the ear canal audio signal to obtain a first input signal, where u is an integer greater than or equal to 1 and less than or equal to i; Determine the first u - 1 sub-audio signals of the predicted audio signal as a second input signal; Determine an input audio signal according to the first input signal and the second input signal; Input the input audio signal into the neural network model to determine the u-th sub-audio signal of the predicted audio signal.

7. The device according to claim 6, Characterized in that, The amplitude of the first frequency response curve corresponding to the coefficients of the second filter bank is less than the amplitude of the second frequency response curve corresponding to the coefficients of the corresponding first filter bank.

8. The device according to claim 6 or 7, Characterized in that, The ear canal audio signal includes m frames of audio signals up to the current frame, and the predicted audio signal includes n frames of audio signals, where m and n are respectively integers greater than or equal to 1.

9. The device according to claim 8, Characterized in that, The neural network model is obtained by the following method: Construct a plurality of training sample pairs, each training sample pair including m * i + 1 sub-audio signal samples, where the first m * i sub-audio signal samples constitute the input sample of the training sample pair, and the last 1 sub-audio signal sample constitutes the output sample of the training sample pair; Train an original network model according to the plurality of training sample pairs to determine the neural network model.

10. The device according to claim 8, Characterized in that, Each sub-audio signal includes at least one audio sampling point.

11. A headset, Characterized in that, The earphone includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the method according to any one of claims 1-5.

12. A non-transitory computer-readable storage medium, Characterized in that, When the instructions in the storage medium are executed by the processor of the earphone, the earphone is enabled to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Squeaking suppression method and device, earphone and storage medium

    CN113596665A