Noise removal methods, headsets, devices, storage media, and computer program products
Patent Information
- Application Number
- KR1020257020661
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-28
- Filing Date
- 2023-06-28
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2043-06-28
Smart Images

Figure 112025069225241-PCT00017_ABST
Abstract
Description
Technology Field
[0001] The present application claims priority to Chinese patent application number 202211505288.9, filed on November 28, 2022, under the title “Noise removal method, headset, device, storage medium and computer program product,” the entirety of which is incorporated by reference herein.
[0002] This application relates to the field of audio processing technology, in particular to noise removal methods, headsets, devices, storage media, and computer program products. Background Technology
[0003] When a user wears a headset to listen to audio signals such as music or voice, ambient noise affects the clarity of the audio signal; if the ambient noise is severe, the user may not even be able to hear the headset's audio signal clearly. Therefore, active noise cancellation must be implemented in the headset to minimize the ambient noise heard by the wearer.
[0004] Active noise cancellation in headsets presents many challenges. Ambient noise is variable and irregular. Furthermore, the degree of environmental noise leakage into the ear canal is related to the fit between the headset and the wearer's ear. However, since the size and shape of the ear canal vary from person to person, and even when multiple people wear the same headset, the degree of noise leakage differs due to varying degrees of fit. The fit may also differ if the same user wears the headset multiple times. Additionally, current active noise cancellation is fundamentally based on downlink signals; active noise cancellation cannot be performed in the absence of a downlink signal. Therefore, methods to enhance the effectiveness of active noise cancellation in headsets to minimize the impact of ambient noise on the wearer are currently a major area of research.
[0005] The present application provides a noise cancellation method, a headset, a device, a storage medium, and a computer program product. By eliminating dependency on downlink signals, adaptive noise cancellation can be performed by determining target noise cancellation parameters even in the absence of downlink signals. The technical solution is as follows.
[0006] According to a first aspect, a noise removal method applied to a headset is provided. The headset comprises at least one first reference microphone, one error microphone, at least one speaker, and one first feedforward (FF) filter. The method comprises: determining a target noise removal parameter based on a reference signal collected by at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient—wherein the target noise removal parameter includes a filter coefficient of the first FF filter—and performing noise removal through a target speaker among at least one speaker based on the target noise removal parameter.
[0007] In the present application, the target noise removal parameter can be determined based on a reference signal collected by at least one first reference microphone, an error signal collected by an error microphone, and an initial noise removal coefficient. This eliminates dependency on the downlink signal, so that adaptive noise removal can be performed by determining the target noise removal parameter even in the absence of a downlink signal.
[0008] According to the noise removal method provided in the embodiments of the present application, the target noise removal parameters may be determined on a frame-by-frame basis. That is, a group of target noise removal parameters is determined in each frame. Of course, the target noise removal parameters may, alternatively, be determined on other time units. For example, a group of target noise removal parameters is determined every two frames. Below, frames are used as the unit of description.
[0009] If the headset includes a first FF filter, the target noise cancellation parameter includes the k-th frame filter coefficient of the first FF filter, where k is an integer greater than 1. In some cases, the headset may also include a feedback (FB) filter. In this case, the target noise cancellation parameter further includes the k-th frame filter coefficient of the FB filter. Additionally, if the headset further includes a downlink compensation filter, multiple groups of target noise cancellation parameters further include the k-th frame filter coefficient of the downlink compensation filter. Furthermore, if k is greater than 1, the target noise cancellation level may be additionally determined. Therefore, four parts are explained separately below.
[0010] (1) Determine the k-th frame filter coefficient of the first FF filter.
[0011] When k is 1, the initial filter coefficient of the first FF filter is determined as the k-th frame filter coefficient of the first FF filter, that is, the first frame filter coefficient of the first FF filter is determined as the initial filter coefficient of the first FF filter, or the k-th frame filter coefficient of the first FF filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the first FF filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the target noise removal level. That is, the k-th frame filter coefficient of the first FF filter is determined according to the adaptation method. This determination process is called the adaptation process and is also referred to as the iterative process.
[0012] The initial noise reduction coefficient includes the initial filter coefficient of the first FF filter, and the initial filter coefficient of the first FF filter may be predetermined and may or may not be zero. This is not limited to the embodiments of the present application. The initial noise reduction level may be a preset level, which is a level at which noise reduction can be performed normally using the corresponding noise reduction coefficient without causing stability issues. Of course, the initial noise reduction level may, alternatively, be a level determined based on an announcement sound, such as "noise reduction on" or "ding-dong," transmitted by the user terminal when noise reduction begins. The noise reduction coefficient corresponding to the level can better adapt to the current human ear and wearing position, and can reach a convergence state more quickly by performing adaptation iterations based on the noise reduction coefficient corresponding to the level. This is also not limited to the embodiments of the present application.
[0013] An implementation process for determining the k-th frame filter coefficient of a first FF filter based on a (k-1)th frame reference signal collected by at least one first reference microphone, a (k-1)th frame error signal collected by an error microphone, and a target noise removal level comprises: a step of determining the (k-1)th frame filter coefficient of a target auxiliary path (SP) based on a mapping relationship between the noise removal level and the filter coefficient of the SP—wherein the target SP is a path from the target speaker to the error microphone—and a step of determining the k-th frame filter coefficient of the first FF filter based on a (k-1)th frame reference signal collected by at least one first reference microphone, a (k-1)th frame error signal collected by an error microphone, and the (k-1)th frame filter coefficient of the target SP.
[0014] If the headset includes a first FF filter, the headset may additionally include an FB filter or not include an FB filter. Additionally, the headset may further include at least one second FF filter, and the filter coefficients of each second FF filter are fixed at the same noise removal level. In some cases, the method for determining the k-th frame filter coefficient of the first FF filter is different. That method is described separately below.
[0015] In the first case, the headset does not include an FB filter and at least one second FF filter. In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by an error microphone, and the (k-1)-th frame filter coefficient of the target SP. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0016] When the k-th frame frequency response information of the first FF filter is determined, the residual error can be determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone and the (k-1)-th frame error signal collected by an error microphone, and the k-th frame frequency response information of the first FF filter can be determined based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the target SP, and the residual error.
[0017] In some embodiments, at least one first reference microphone includes one reference microphone. That is, the first FF filter corresponds to one reference microphone. In this case, the residual error is determined based on the (k-1)th frame reference signal collected by the reference microphone and the (k-1)th frame error signal collected by the error microphone.
[0018] In some other embodiments, at least one first reference microphone includes at least two reference microphones. That is, the first FF filter corresponds to at least two reference microphones. In this case, audio mixing is performed on the (k-1)th frame reference signal collected by at least two reference microphones to obtain the (k-1)th frame mixed reference signal. The residual error is determined based on the (k-1)th frame mixed reference signal and the (k-1)th frame error signal collected by the error microphone. In this way, the signal-to-noise ratio of the reference signal can be improved.
[0019] An implementation process for determining the k-th frame filter coefficient of a first FF filter based on the k-th frame frequency response information of a first FF filter includes the step of setting a loss function between the filter coefficient variable of the first FF filter and the k-th frame frequency response information of the first FF filter. The value of the filter coefficient variable is determined based on the loss function according to the gradient descent method, and the k-th frame filter coefficient of the first FF filter is determined based on the value of the filter coefficient variable. That is, a loss function is set between the filter coefficient variable of the first FF filter and the k-th frame frequency response information of the first FF filter. Since the optimal value of the variable is determined according to the gradient descent method, the k-th frame filter coefficient of the first FF filter is determined based on the optimal value of the variable.
[0020] In each frame, the filter coefficients of the first FF filter are determined according to the gradient descent method. Once the filter coefficients of the first FF filter are determined in each frame, one value of the loss function is determined. When the value of the loss function reaches a minimum threshold, it is determined that the filter coefficients of the first FF filter have reached the convergence stability condition. For example, in the case of the k-th frame filter coefficient of the first FF filter, if the value of the loss function between the filter coefficient variable and the k-th frame frequency response information of the first FF filter reaches the minimum threshold, it is determined that the k-th frame filter coefficient of the first FF filter has reached the convergence stability condition. If the value of the loss function does not reach the minimum threshold, it is determined that the k-th frame filter coefficient of the first FF filter has not reached the convergence stability condition. The minimum threshold is preset and may be adjusted according to other requirements depending on the case.
[0021] Optionally, the filter coefficients of each FF filter include at least one biquad filter coefficient and one gain. Variables corresponding to the biquad filter coefficient include the filter type, cutoff frequency, and quality factor. Of course, in actual applications, the filter coefficients of each FF filter may include more or fewer other parameters. This is not limited to the embodiments of this application.
[0022] In some cases, background noise, or the noise floor problem, may occur in quiet environments. For example, semi-open headsets are more likely to have background noise problems than in-ear headsets in quiet environments. Additionally, powerful noise cancellation is not required in quiet environments, and some people may experience discomfort when powerful noise cancellation is performed in such environments. Furthermore, a higher noise cancellation intensity indicates a stronger sense of sound pressure for the user. Therefore, when determining the values of filter coefficient variables according to the gradient descent method, the subjective experience effect of adaptive noise cancellation can be enhanced by dynamically adjusting the target noise cancellation amplitude according to the ambient volume, so that the k-th frame filter coefficient of the first FF filter is determined based on the target noise cancellation amplitude. That is, the target noise cancellation amplitude is determined based on the ambient volume of the (k-1)-th frame and the ambient volume of the t frames prior to the (k-1)-th frame, where t is greater than or equal to 1 and less than k-1. The value of the filter coefficient variable is determined based on the loss function according to the target noise removal amplitude and the gradient descent method, and the k-th frame filter coefficient of the first FF filter is determined based on the value of the filter coefficient variable.
[0023] In the second case, the headset includes one additional FB filter but does not include at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the FB filter.
[0024] Here, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the FB filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0025] When the k-th frame frequency response information of the first FF filter is determined, the residual error can be determined based on the (k-1)th frame reference signal collected by at least one first reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter can be determined based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, the residual error, and the (k-1)th frame filter coefficient of the FB filter.
[0026] In the third case, the headset does not include an FB filter, but includes at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame frequency response information of at least one second FF filter.
[0027] In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target SP, and the (k-1)th frame frequency response information of at least one second FF filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0028] When the k-th frame frequency response information of the first FF filter is determined, the residual error may be determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone and the (k-1)-th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter may be determined based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the target SP, the residual error, and the (k-1)-th frame response information of at least one second FF filter.
[0029] In the fourth case, the headset includes an FB filter and at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, the (k-1)-th frame filter coefficient of the FB filter, and the (k-1)-th frame frequency response information of at least one second FF filter.
[0030] In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target SP, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame frequency response information of at least one second FF filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0031] When the k-th frame frequency response information of the first FF filter is determined, the residual error may be determined based on the (k-1)th frame reference signal collected by at least one first reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter may be determined based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, the residual error, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame frequency response information of at least one second FF filter.
[0032] In the process of determining the k-th frame frequency response information of the aforementioned first FF filter, regardless of whether the headset includes an FB filter, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the target SP is determined based on the target noise removal level by querying the mapping relationship between the noise removal level and the filter coefficient of the SP. Specifically, the (k-1)-th frame filter coefficient of the target SP is an estimated value, and by determining the k-th frame frequency response information of the first FF filter based on the estimated value, dependency on the actual value of the target SP can be eliminated, and filter coefficient adaptation of the FF filter can be implemented even in the absence of a downlink signal.
[0033] (2) Determine the k-th frame filter coefficient of the FB filter.
[0034] When k is 1, the initial filter coefficient of the FB filter is determined as the k-th frame filter coefficient of the FB filter, that is, the first frame filter coefficient of the FB filter is determined as the initial filter coefficient of the FB filter, or the k-th frame filter coefficient of the FB filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the FB filter coefficient. When K is greater than 1, the k-th frame filter coefficient of the FB filter may be determined based on the target noise removal level. The initial noise removal coefficient includes the initial filter coefficient of the FB filter, and the initial filter coefficient may or may not be 0. This is not limited to the embodiments of the present application.
[0035] When k is greater than 1, the k-th frame filter coefficient of the FB filter can be determined in the following two ways.
[0036] In the first method, the k-th frame filter coefficient of the FB filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the FB filter coefficient.
[0037] Since the mapping relationship between the noise removal level and the FB filter coefficients is stored in advance, determining the k-th frame filter coefficient of the FB filter in the first method is stable, simple to operate, and highly efficient.
[0038] In the second method, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the FB filter, and the target noise removal level.
[0039] The first frame filter coefficient of the FB filter can be determined based on the initial noise removal level by querying the mapping relationship between the noise removal level and the FB filter coefficient. Thus, when k is greater than or equal to 1, the k-th frame filter coefficient of the FB filter can be determined in two ways: (1) The k-th frame filter coefficient of the FB filter is determined by querying the mapping relationship between the noise removal level and the FB filter coefficient. (2) When k is 1, the k-th frame filter coefficient of the FB filter is determined by querying the mapping relationship between the noise removal level and the FB filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal level.
[0040] An implementation process for determining the k-th frame filter coefficient of an FB filter based on the (k-1)th frame error signal collected by an error microphone, the (k-1)th frame filter coefficient of an FB filter, and the target noise removal level comprises: the step of determining the (k-1)th frame filter coefficient of a target SP based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficient of the SP—wherein the target SP is the path from the target speaker to the error microphone—and the step of determining the k-th frame filter coefficient of an FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of an FB filter, and the (k-1)th frame filter coefficient of a target SP.
[0041] In the second method described above, the method of querying the mapping relationship between the noise removal level and the FB filter coefficients is combined with an adaptive method to improve the noise removal effect and ensure controllable stability with low complexity.
[0042] It should be noted that in the embodiments of the present application, the k-th frame filter coefficient of the FB filter can be determined in the two ways described above, and alternatively, the k-th frame filter coefficient of the FB filter can be determined in another way. For example, regardless of whether k is greater than 1 or equal to 1, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal level. This is not limited to the embodiments of the present application.
[0043] (3) Determine the k-th frame filter coefficient of the downlink compensation filter.
[0044] When k is 1, the initial filter coefficient of the downlink compensation filter is determined as the k-th frame filter coefficient of the downlink compensation filter, or the k-th frame filter coefficient of the downlink compensation filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the downlink compensation filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the downlink compensation filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the downlink compensation filter coefficient.
[0045] The mapping relationship between noise reduction levels and downlink compensation filter coefficients involves multiple noise reduction levels, and a mapping relationship exists between each noise reduction level and the filter coefficients of the downlink compensation filter; furthermore, the mapping relationships between different noise reduction levels and the filter coefficients of the downlink compensation filter may differ. Therefore, after the target noise reduction level is determined, the corresponding downlink compensation filter coefficients can be obtained from the mapping relationship between the noise reduction level and the downlink compensation filter coefficients based on the target noise reduction level, and the obtained downlink compensation filter coefficients are used as the k-th frame filter coefficients of the downlink compensation filter.
[0046] (4) Determine the target noise reduction level.
[0047] The noise reduction level of the (k-1)th frame is determined, and the noise reduction levels of m frames prior to the (k-1)th frame are obtained, where m is greater than or equal to 1 and less than k-1. The target noise reduction level is determined based on the noise reduction level of the (k-1)th frame and the noise reduction levels of the m frames.
[0048] In the (k-1)th frame, a valid downlink signal may or may not exist, the environment may or may not be quiet, or, of course, an abnormal signal may exist. Depending on the case, the method for determining the noise cancellation level of the (k-1)th frame varies, and this is explained separately below.
[0049] In the first case, in the (k-1)th frame, there is no effective downlink signal and the environment is not quiet. In this case, the noise removal level of the (k-1)th frame is determined based on the reference filter coefficients of the first FF filter and the mapping relationship between the noise removal level and the frequency response information of the first FF filter. When k is 2, the reference filter coefficients are the initial filter coefficients of the first FF filter, or when k is greater than 2, the reference filter coefficients are the filter coefficients of the first FF filter that last satisfied the convergence stability condition before the kth frame, or the (k-1)th frame filter coefficients of the first FF filter.
[0050] In some embodiments, the reference frequency response information of the first FF filter is determined based on the reference filter coefficients of the first FF filter. The noise removal level corresponding to the reference frequency response information of the first FF filter is determined based on the mapping relationship between the noise removal level and the frequency response information of the first FF filter to obtain the (k-1)th frame noise removal level.
[0051] In the second case, an effective downlink signal exists in the (k-1)th frame. In this case, the noise cancellation level of the (k-1)th frame is determined based on the effective downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone.
[0052] In light of the above description, if the headset is in a downlink active state and is not in a downlink intermittent period, it is determined that an effective downlink signal exists in the (k-1)th frame. In this case, based on the (k-1)th frame effective downlink signal, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone, the effective downlink signal can be extracted from the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame noise removal level can be determined based on the extracted effective downlink signal.
[0053] In the third case, there is no valid downlink signal in the (k-1)th frame and the environment is quiet, or there is an abnormal noise signal in the (k-1)th frame. In this case, the noise cancellation level of the (k-3)th frame is determined to be the noise cancellation level of the (k-1)th frame. That is, the noise cancellation level is not changed.
[0054] In the (k-1)th frame, if there is no valid downlink signal and the surrounding environment is quiet, the noise is essentially unchanged. In this case, the noise cancellation level may not change. If an abnormal noise signal is present in the (k-1)th frame, the noise cancellation level remains unchanged to perform robust control and prevent the divergence of the noise cancellation level.
[0055] After the noise reduction level of the (k-1)th frame is determined in the three cases above, the target noise reduction level can be determined by integrating the noise reduction level of the (k-1)th frame with the noise reduction levels of the m frames prior to the (k-1)th frame.
[0056] The noise reduction level in m frames may be the noise reduction level in any m frames prior to the (k-1)th frame, or the noise reduction level in the m frames prior to the (k-1)th frame and closest to the (k-1)th frame. This is not limited to the embodiments of the present application. Additionally, there are multiple implementations for determining a target noise reduction level based on the noise reduction level in the (k-1)th frame and the noise reduction levels in the m frames prior to the (k-1)th frame. For example, the noise reduction effect may be evaluated according to a relevant algorithm to determine the noise reduction probability corresponding to the (k-1)th frame noise reduction level and the noise reduction probability corresponding to the noise reduction level in m frames, and the noise reduction level with the highest noise reduction probability may be determined as the target noise reduction level. Alternatively, the target noise reduction level may be obtained by determining the arithmetic mean or weighted mean of the noise reduction level in the (k-1)th frame and the noise reduction levels in m frames. Alternatively, the noise reduction level that appears most frequently among the (k-1)th frame noise reduction level and the noise reduction levels in m frames is determined as the target noise reduction level, etc.
[0057] Target inverse phase noise is generated based on target noise removal parameters, and noise removal is performed through the target speaker among at least one speaker based on the target inverse phase noise.
[0058] If the target noise removal parameter includes the k-th frame filter coefficient of the first FF filter, the target inverse phase noise includes feedforward inverse phase noise. In this case, the k-th frame reference signal collected by at least one first reference microphone can be processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise.
[0059] In light of the above description, at least one first reference microphone may include one reference microphone or at least two reference microphones. If at least one first reference microphone includes one reference microphone, the k-th frame reference signal collected by the reference microphone is directly processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise. If at least one first reference microphone includes at least two reference microphones, audio mixing is performed on the k-th frame reference signal collected by at least two reference microphones to obtain a k-th frame mixed reference signal, and then the k-th frame mixed reference signal is processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise.
[0060] Optionally, the headset may further include at least one second FF filter. In this case, the k-th frame filter coefficient of at least one second FF filter may be determined. Noise removal is performed through a target speaker based on target noise removal parameters and the k-th frame filter coefficient of at least one second FF filter. That is, the k-th frame reference signal collected by at least one first reference microphone is processed based on the k-th frame filter coefficient of the first FF filter to obtain a first feedforward inverse phase noise. The k-th frame reference signal collected by at least one first reference microphone is processed based on the k-th frame filter coefficient of at least one second FF filter to obtain at least one second feedforward inverse phase noise.
[0061] If k is 1, the initial filter coefficient of at least one second FF filter is determined as the k-th frame filter coefficient of at least one second FF filter, that is, the first frame filter coefficient of at least one second FF filter is the initial filter coefficient of the corresponding second FF filter, or the k-th frame filter coefficient of at least one second FF filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the second FF filter coefficient. If k is greater than 1, the k-th frame filter coefficient of at least one second FF filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the second FF filter coefficient. The initial noise removal coefficient includes the initial filter coefficient of at least one second FF filter, and the initial filter coefficient may be 0 or not 0. This is not limited to the embodiments of the present application.
[0062] The foregoing description is provided using an example in which the first FF filter and at least one second FF filter both correspond to at least one first reference microphone. In actual applications, the first FF filter and at least one second FF filter may correspond to different reference microphones. For example, a headset further includes a plurality of second reference microphones, the first FF filter corresponds to at least one first reference microphone, and each second FF filter corresponds to at least one of the plurality of second reference microphones. In this case, the k-th frame reference signal collected by at least one first reference microphone can be processed based on the k-th frame filter coefficient of the first FF filter to obtain the first feedforward inverse phase noise. The k-th frame reference signal collected by at least one second reference microphone corresponding to each second FF filter can be processed based on the k-th frame filter coefficient of each second FF filter to obtain at least one second feedforward inverse phase noise.
[0063] If the headset additionally includes an FB filter, the target inverse phase noise further includes feedback inverse phase noise. That is, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal. Specifically, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal based on the k-th frame filter coefficients of the downlink compensation filter. Then, after inversion is performed on the k-th frame downlink signal obtained through downlink compensation, audio mixing is performed on the inverted k-th frame downlink signal and the k-th frame error signal collected by the error microphone to obtain the k-th frame noise signal collected by the error microphone. The k-th frame noise signal collected by the error microphone is processed based on the k-th frame filter coefficients of the FB filter to obtain feedback inverse phase noise.
[0064] Downlink compensation can be used to prevent sound quality degradation of the downlink signal by removing all downlink signals from the error signal collected by the error microphone and performing noise removal only on the residual noise signal through an FB filter. Additionally, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal to remove downlink signals from all speakers in the error microphone, thereby preventing sound quality degradation of the full-band downlink signal.
[0065] In light of the above description, when the target noise removal parameter is determined on a frame-by-frame basis, since a single frame may include one sample point or multiple sample points, when the target inverse phase noise is generated, a target inverse phase noise group may be generated at each sample point or a target inverse phase noise group may be generated in a single frame.
[0066] In an embodiment of the present application, when the target noise removal parameter is determined, frequency division is not performed on the downlink signal; that is, the target noise removal parameter is determined based on the full-band downlink signal. In this manner, after the target inverse phase noise is generated based on the target noise removal parameter, the frequency band of the target inverse phase noise includes the sound generation frequency band of at least one speaker; that is, the frequency band of the target inverse phase noise is the entire frequency band.
[0067] After the target inverse phase noise is generated, the target inverse phase noise is mixed with the k-th frame downlink signal to be played through the target speaker, and then the mixed signal is played through the target speaker to achieve noise removal.
[0068] At least one speaker may include a single speaker or a plurality of speakers. If at least one speaker includes a single speaker, the speaker may be a full-band speaker. If at least one speaker includes a plurality of speakers, some of the plurality of speakers may be high-band speakers and others may be low-band speakers. Alternatively, some of the plurality of first speakers may be full-band speakers and others may be non-full-band speakers. In other words, the sound generation frequency bands of the plurality of speakers may differ from one another. Alternatively, all of the plurality of speakers may be full-band speakers. Alternatively, all of the plurality of speakers may be non-full-band speakers. If all of the plurality of speakers are full-band speakers, the k-th frame downlink signal played through the plurality of speakers is the k-th frame downlink signal transmitted from the user terminal. If not all of the plurality of speakers are full-band speakers, frequency division must be performed on the k-th frame downlink signal transmitted by the user terminal based on the sound generation frequency band of each speaker to obtain the k-th frame downlink signal to be played through each speaker.
[0069] In an embodiment of the present application, noise cancellation is performed through a target speaker among at least one speaker. If the at least one speaker includes one speaker, that speaker is the target speaker. If the at least one speaker includes a plurality of speakers and the plurality of speakers includes a first speaker and a second speaker to which digital frequency division is performed, the target speaker is the first speaker, and the second speaker does not participate in noise cancellation. However, the second speaker may participate in downlink compensation (i.e., downlink compensation is performed on a downlink signal transmitted by a user terminal, wherein the downlink signal is a full-band audio signal including an audio signal in the sound generation frequency band of the second speaker). In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker or a mid-band speaker or a low-band speaker. Optionally, the second speaker may not participate in downlink compensation. In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker.
[0070] Optionally, at least one speaker may include a first speaker and a second speaker on which analog frequency division is performed as an alternative. In this case, the target speakers are the first speaker and the second speaker, that is, both the first speaker and the second speaker participate in noise cancellation.
[0071] If the target speakers are the first and second speakers, the first and second speakers may be an analog frequency division combination of the two speakers. That is, the first and second speakers are driven using the same DAC and PA, and the combination of the first and second speakers can be considered as a single speaker.
[0072] The process of determining the target noise removal parameter according to the aforementioned adaptation method requires a specific amount of time. If a single frame contains multiple sample points and the duration of the single frame is long, the duration for determining the target noise removal parameter is shorter than the duration of the single frame. Therefore, based on the relevant data of the (k-1)th frame, calculations can be performed during a portion of the duration of the k-th frame to obtain the target noise removal parameter of the k-th frame, and active noise removal can be performed during another portion of the duration of the k-th frame based on the target noise removal parameter. However, if a single frame contains only one sample point, or if a single frame contains multiple sample points and the duration of the single frame is short, the duration for determining the target noise removal parameter may be equal to the duration of the single frame. In this case, to obtain the target noise removal parameter, calculations may need to be performed over the entire duration of the k-th frame based on the relevant data of the (k-1)th frame. In this case, the target noise reduction parameter can be determined as the target noise reduction parameter of the (k+1)th frame, and active noise reduction can be performed during the period of the (k+1)th frame based on the target noise reduction parameter of the (k+1)th frame. The preceding content is explained using the latter case as an example.
[0073] According to a second aspect, a headset is provided. The headset includes at least one first reference microphone, one error microphone, at least one speaker, one first feedforward FF filter, and one noise removal processor.
[0074] The noise removal processor is configured to implement the steps of the method according to the first aspect.
[0075] Optionally, at least one speaker includes a first speaker and a second speaker on which digital frequency division is performed, the target speaker is the first speaker, and the second speaker does not participate in noise cancellation.
[0076] Optionally, the second speaker participates in downlink compensation.
[0077] Optionally, the second speaker does not participate in downlink compensation, and the second speaker is a high-frequency speaker.
[0078] Optionally, at least one speaker includes a first speaker and a second speaker on which analog frequency division is performed, and the target speaker is the first speaker and the second speaker.
[0079] Optionally, the headset further includes at least one second FF filter, and the filter coefficient of the second FF filter is fixed at the same noise removal level.
[0080] Optionally, the headset further includes a plurality of second reference microphones, the first FF filter corresponds to at least one first reference microphone, and each of the at least one second FF filter corresponds to at least one of the plurality of second reference microphones.
[0081] According to a third aspect, a noise removal device is provided. The noise removal device has the function of implementing the operation of the noise removal method of the first aspect. The noise removal device includes one or more modules, and one or more modules are configured to implement the noise removal method provided in the first aspect.
[0082] According to the fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions. When instructions are executed on a computer, the computer can perform the noise removal method described in the first aspect.
[0083] According to the fifth aspect, a computer program product including instructions is provided. When the instructions are executed on the computer, the computer can perform the noise removal method described in the first aspect.
[0084] The technical effects obtained in the second through fifth aspects are similar to the technical effects obtained by the corresponding technical means of the first aspect. Details are not described further in this specification. Brief explanation of the drawing
[0085] FIG. 1 is a diagram of a system architecture related to a noise removal method according to an embodiment of the present application. FIG. 2 is a flowchart of a noise removal method according to an embodiment of the present application. FIG. 3 is a flowchart for determining the target noise removal amplitude according to an embodiment of the present application. FIG. 4 is a diagram of the frequency response curve of an FF filter at 16 noise removal levels according to an embodiment of the present application. FIG. 5 is a flowchart for determining the (k-1)th frame noise removal level according to an embodiment of the present application. FIG. 6 is a flowchart for determining target noise removal parameters according to an embodiment of the present application. FIG. 7 is a diagram of the structure of a headset according to an embodiment of the present application. FIG. 8 is a diagram of the structure of another headset according to an embodiment of the present application. FIG. 9 is a diagram of the structure of another headset according to an embodiment of the present application. FIG. 10 is a diagram of the structure of another headset according to an embodiment of the present application. FIG. 11 is a diagram of the structure of another headset according to an embodiment of the present application. FIG. 12 is a structural diagram of another headset according to an embodiment of the present application. FIG. 13 is a diagram of the structure of a noise removal device according to an embodiment of the present application. FIG. 14 is a diagram of the structure of another headset according to an embodiment of the present application. Specific details for implementing the invention
[0086] To more clearly explain the purpose, technical solution, and advantages of this application, the implementation of the present application is described in detail below with reference to the attached drawings.
[0087] Active noise-canceling headsets have gained popularity in recent years. Conventional noise-canceling headsets are typically in-ear or head-mounted types. This is because, in these two forms, the headset seals well with the ear canal, and acoustic leakage is stable when another person wears the headset. Technically, this allows for the implementation of active noise cancellation, which is more effective. Therefore, noise cancellation modes with a fixed coefficient are generally used. However, these two types of headsets also have some drawbacks. For example, the seal between the headset and the ear canal is so strong that it affects users' subjective comfort, which usually manifests as a foreign body sensation or a feeling of blockage while walking. They are also difficult to wear for extended periods.
[0088] Semi-open headsets are widely used by users due to their comfortable fit. However, because the headset does not seal well with the user's ear, ambient noise can be perceived more easily. Implementing active noise cancellation in semi-open headsets is more challenging. This is because wearing postures vary significantly when multiple people wear the headset, or even when the same person wears it at different times. Technically, the response between the headset and the ear canal, as well as the degree of acoustic leakage, differ considerably. Therefore, implementing adaptive noise cancellation and achieving optimal matching between the headset and the ear canal to address the issue of discrepancies in ear canal response is an urgent requirement for semi-open headsets. Furthermore, ear canal response is not absolutely uniform in in-ear or head-mounted forms, with variations existing both large and small. Currently, the industry is also exploring the feasibility of implementing adaptive noise cancellation in headsets.
[0089] As mentioned above, headsets come in multiple forms, such as in-ear, head-mount, semi-open, and open types. The audio performance of the speakers (i.e., loudspeakers) in a headset, particularly at low frequencies, is closely related to the specific form of the headset. In closed forms, such as in-ear or head-mount types, audio performance at high, mid, and low frequencies is generally guaranteed. In semi-open or open types, low-frequency response drops significantly due to severe acoustic leakage. This affects low-frequency sound quality and severely impacts the effectiveness of active noise cancellation (insufficient energy to generate inverse phase noise).
[0090] Considering the aforementioned problems, embodiments of the present application provide a noise cancellation method for implementing adaptive active noise cancellation (active noise cancellation, ANC) of a headset. Refer to FIG. 1. FIG. 1 is a diagram of a system architecture related to a noise cancellation method according to an embodiment of the present application. This system may be referred to as a headset noise cancellation system. The system includes a headset (101) and a user terminal (102). The headset (101) and the user terminal (102) are connected via a wired or wireless method to perform communication. For example, the headset (101) communicates with the user terminal (102) via Bluetooth or another wireless network.
[0091] Audio signals and control signals can be transmitted between the headset (101) and the user terminal (102). For example, the user terminal (102) transmits an audio signal, such as music or voice, to the headset (101) for playback. As another example, the user terminal (102) can transmit a control signal to the headset (101) to control whether the active noise cancellation function of the headset (101) is activated.
[0092] The user terminal (102) may be an electronic device such as a mobile phone or a computer (e.g., a laptop computer, a desktop computer, a portable tablet computer, or a vehicle-mounted tablet computer). The user terminal (102) may also be another electronic device, e.g., a smart speaker or a vehicle-mounted speaker. The type, structure, etc. of the user terminal (102) is not limited to the embodiments of the present application.
[0093] Optionally, the headset (101) provided in the embodiments of the present application may be wired or wireless. Additionally, regarding the wearing method, the headset (101) provided in the embodiments of the present application may be a neck mount type, an ear mount / ear clip type, a true wireless stereo (TWS) type, etc. Regarding the appearance, the headset (101) provided in the embodiments of the present application may be in the form of an in-ear type, a semi-open type, an open type, a head mount type, etc. The communication method, the wearing method, and the appearance of the headset are not limited to the embodiments of the present application. Hereinafter, the hardware structure of the headset provided in the embodiments of the present application will be described with reference to the method of wearing the headset on a person's ear.
[0094] As illustrated in FIG. 1, the headset (101) includes at least one speaker (i.e., a loudspeaker), a plurality of microphones, a microcontrol unit (MCU), an ANC chip, and memory. At least one speaker includes a target speaker, and the target speaker represents a speaker participating in noise cancellation, e.g., a loudspeaker (1). The target speaker may be a first speaker, and the first speaker must participate in noise cancellation. For example, the first speaker is a low-frequency and mid-frequency speaker, and the low-frequency and mid-frequency speaker must participate in noise cancellation. Optionally, at least one speaker further includes a second speaker, and the second speaker does not participate in noise cancellation. For example, the second speaker is a high-frequency speaker, and the high-frequency speaker does not need to participate in noise cancellation. Of course, for any speaker, regardless of whether the speaker is a high-frequency speaker or a mid-low frequency speaker, the speaker may or may not participate in noise cancellation. That is, in the embodiments of the present application, the sound generation frequency band of the first speaker participating in noise removal is not limited, and the sound generation frequency band of the second speaker not participating in noise removal is not limited. For example, the target speakers are the first speaker and the second speaker. A plurality of microphones includes at least one reference microphone and one error microphone. FIG. 1 illustrates one reference microphone as an example.
[0095] The speakers are configured to play a downlink signal (an audio signal, such as music or voice). At least one speaker is driven using an independent digital-to-analog converter (DAC) and power amplifier (PA). That is, one speaker corresponds to one DAC and one PA, and different speakers correspond to different DACs and PAs. Alternatively, at least one speaker may be driven using the same DAC and PA, or some of the speakers may be driven using the same DAC and PA while others are driven using different DACs and PAs. In the noise cancellation process, the target speakers are further configured to play out-phase noise, where the out-phase noise is used to reduce the noise signal in the user's ear canal to achieve an active noise cancellation effect.
[0096] A reference microphone is positioned outside the headset. After the headset is worn on a person's ears, the reference microphone is located outside the person's ears. The reference microphone is configured to collect noise signals from the external environment. In an embodiment of the present application, the noise signal collected by the reference microphone is referred to as a reference signal.
[0097] The error microphone is placed inside the headset. After the headset is worn on a person's ear, the error microphone is positioned inside the person's ear. The error microphone is configured to collect noise signals from the external auditory canal. In an embodiment of the present application, the noise signal collected by the error microphone is referred to as the error signal.
[0098] The micro-control unit is configured to process the reference signal collected by the reference microphone, the error signal collected by the error microphone, the downlink signal, etc., to determine the target noise cancellation parameter, and to record the target noise cancellation parameter on the ANC chip.
[0099] The ANC chip is configured to generate inverse phase noise by processing a reference signal collected by a reference microphone and an error signal collected by an error microphone based on target noise cancellation parameters, perform audio mixing of the generated inverse phase noise and a downlink signal to be played through a speaker, and output the mixed signal to the speaker to reduce the noise signal in the external auditory canal.
[0100] The memory is configured to store initial parameters, mapping relationships, etc., used when determining target noise removal parameters.
[0101] It should be noted that the microcontrol unit, ANC chip, and memory may be integrated on the same circuit board or placed on different circuit boards. This is not limited to the embodiments of this application. Furthermore, the microcontrol unit and the ANC chip are distinguished only in terms of logical functional description. In actual physical form, the microcontrol unit and the ANC chip may be integrated on a single chip or placed individually on multiple chips. For example, the microcontrol unit and the ANC chip are placed on two chips.
[0102] Optionally, the headset (101) may further include other elements configured to detect whether the headset (101) is inside the ear, for example, an optical proximity sensor. If the headset (101) is a wireless headset, the headset (101) may further include a wireless communication module, the wireless communication module may be a wireless local area network module or a Bluetooth module. The wireless communication module is used for the headset (101) to communicate with other devices.
[0103] It should be understood that the schematic structure of the embodiments of the present application does not constitute a limitation on the headset. In some other embodiments, the headset (101) may include more or fewer components than shown in the drawings, some components may be combined, some components may be divided, or other arrangements of components may be used. The components shown in the drawings may be implemented using hardware, software, or a combination of software and hardware.
[0104] The system architecture and service scenarios described in the embodiments of this application are intended to more clearly explain the technical solutions in the embodiments of this application and are not intended to limit the technical solutions provided in the embodiments of this application. Those skilled in the art will understand that, with the evolution of system architectures and the emergence of new service scenarios, the technical solutions provided in the embodiments of this application are applicable to similar technical problems.
[0105] FIG. 2 is a flowchart of a noise removal method according to an embodiment of the present application. This method is applied to a headset, which includes at least one first reference microphone, one error microphone, at least one first speaker, and one first FF filter. Refer to FIG. 2. This method comprises the following steps.
[0106] Step (201): A target noise removal parameter is determined based on a reference signal collected by at least one first reference microphone, an error signal collected by an error microphone, and an initial noise removal coefficient, wherein the target noise removal parameter includes the filter coefficient of the first FF filter.
[0107] According to the noise removal method provided in the embodiments of the present application, the target noise removal parameters may be determined on a frame-by-frame basis. That is, a group of target noise removal parameters is determined in each frame. Of course, the target noise removal parameters may, alternatively, be determined on other time units. For example, a group of target noise removal parameters is determined every two frames. Below, frames are used as the unit of description.
[0108] If the headset includes a first FF filter, the target noise cancellation parameter includes the k-th frame filter coefficient of the first FF filter, where k is an integer greater than 1. In some cases, the headset may also include an FB filter. In this case, the target noise cancellation parameter further includes the k-th frame filter coefficient of the FB filter. Additionally, if the headset further includes a downlink compensation filter, multiple groups of target noise cancellation parameters further include the k-th frame filter coefficient of the downlink compensation filter. Furthermore, if k is greater than 1, the target noise cancellation level may be additionally determined. Therefore, four parts are explained separately below.
[0109] (1) Determine the k-th frame filter coefficient of the first FF filter.
[0110] When k is 1, the initial filter coefficient of the first FF filter is determined as the k-th frame filter coefficient of the first FF filter, that is, the first frame filter coefficient of the first FF filter is determined as the initial filter coefficient of the first FF filter, or the k-th frame filter coefficient of the first FF filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the first FF filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the target noise removal level. That is, the k-th frame filter coefficient of the first FF filter is determined according to the adaptation method. This determination process is called the adaptation process and is also referred to as the iterative process.
[0111] The initial noise reduction coefficient includes the initial filter coefficient of the first FF filter, and the initial filter coefficient of the first FF filter may be predetermined and may or may not be zero. This is not limited to the embodiments of the present application. The initial noise reduction level may be a preset level, which is a level at which noise reduction can be performed normally using the corresponding noise reduction coefficient without causing stability issues. Of course, the initial noise reduction level may, alternatively, be a level determined based on an announcement sound, such as "noise reduction on" or "ding-dong," transmitted by the user terminal when noise reduction begins. The noise reduction coefficient corresponding to the level can better adapt to the current human ear and wearing position, and can reach a convergence state more quickly by performing adaptation iterations based on the noise reduction coefficient corresponding to the level. This is also not limited to the embodiments of the present application.
[0112] An implementation process for determining the k-th frame filter coefficient of a first FF filter based on a (k-1)th frame reference signal collected by at least one first reference microphone, a (k-1)th frame error signal collected by an error microphone, and a target noise removal level comprises: a step of determining the (k-1)th frame filter coefficient of a target SP based on a target noise removal level and a mapping relationship between the noise removal level and the filter coefficient of the SP—wherein the target SP is a path from the target speaker to the error microphone—and a step of determining the k-th frame filter coefficient of a first FF filter based on a (k-1)th frame reference signal collected by at least one first reference microphone, a (k-1)th frame error signal collected by an error microphone, and the (k-1)th frame filter coefficient of the target SP.
[0113] The mapping relationship between the noise removal level and the filter coefficients of the SP involves multiple noise removal levels. A mapping relationship exists between each noise removal level and the filter coefficients of the target SP, and the mapping relationships between different noise removal levels and the filter coefficients of the target SP may differ. Therefore, after the target noise removal level is determined, the filter coefficients corresponding to the target SP can be obtained from the mapping relationship between the noise removal level and the filter coefficients of the SP based on the target noise removal level, and the obtained filter coefficients are used as the (k-1)th frame filter coefficients of the target SP.
[0114] If the headset includes a first FF filter, the headset may additionally include an FB filter or not include an FB filter. Additionally, the headset may further include at least one second FF filter, and the filter coefficients of each second FF filter are fixed at the same noise removal level. In some cases, the method for determining the k-th frame filter coefficient of the first FF filter is different. That method is described separately below.
[0115] In the first case, the headset does not include an FB filter and at least one second FF filter. In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by an error microphone, and the (k-1)-th frame filter coefficient of the target SP. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0116] When the k-th frame frequency response information of the first FF filter is determined, the residual error can be determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone and the (k-1)-th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter can be determined based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the target SP, and the residual error.
[0117] In some embodiments, at least one first reference microphone includes one reference microphone. That is, the first FF filter corresponds to one reference microphone. In this case, the residual error is determined according to the subsequent formula (1) based on the (k-1)th frame reference signal collected by the reference microphone and the (k-1)th frame error signal collected by the error microphone.
[0118]
[0119] In the previous formula (1), represents the residual error, and represents the (k-1)th frame reference signal collected by the reference microphone, and represents the (k-1)th frame error signal collected by the error microphone.
[0120] In some other embodiments, at least one first reference microphone includes at least two reference microphones. That is, the first FF filter corresponds to at least two reference microphones. In this case, audio mixing is performed on the (k-1)th frame reference signal collected by at least two reference microphones to obtain the (k-1)th frame mixed reference signal. The residual error is determined based on the (k-1)th frame mixed reference signal and the (k-1)th frame error signal collected by the error microphone. In this way, the signal-to-noise ratio of the reference signal can be improved.
[0121] The method of determining the residual error based on the (k-1)th frame mixed reference signal and the (k-1)th frame error signal collected by the error microphone is similar to the method of determining the residual error according to the aforementioned formula (1). Specifically, the residual error is obtained by dividing the (k-1)th frame error signal collected by the error microphone by the (k-1)th frame mixed reference signal.
[0122] In some embodiments, frequency response information of the (k-1)th frame filter coefficient of the target SP can be determined, and based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, and residual error, the k-th frame frequency response information of the first FF filter can be determined according to the following formula (2).
[0123] (2)
[0124] In the previous formula (2), represents the k-th frame frequency response information of the first FF filter, and represents the (k-1)th frame frequency response information of the first FF filter, and represents a stage and is pre-set, represents the frequency response information of the (k-1)th frame filter coefficient of the target SP.
[0125] An implementation process for determining the k-th frame filter coefficient of a first FF filter based on the k-th frame frequency response information of a first FF filter includes the step of setting a loss function between the filter coefficient variable of the first FF filter and the k-th frame frequency response information of the first FF filter. The value of the filter coefficient variable is determined based on the loss function according to the gradient descent method, and the k-th frame filter coefficient of the first FF filter is determined based on the value of the filter coefficient variable. That is, a loss function is set between the filter coefficient variable of the first FF filter and the k-th frame frequency response information of the first FF filter. Since the optimal value of the variable is determined according to the gradient descent method, the k-th frame filter coefficient of the first FF filter is determined based on the optimal value of the variable.
[0126] In each frame, the filter coefficients of the first FF filter are determined according to the gradient descent method. Once the filter coefficients of the first FF filter are determined in each frame, one value of the loss function is determined. When the value of the loss function reaches a minimum threshold, it is determined that the filter coefficients of the first FF filter have reached the convergence stability condition. For example, in the case of the k-th frame filter coefficient of the first FF filter, if the value of the loss function between the filter coefficient variable and the k-th frame frequency response information of the first FF filter reaches the minimum threshold, it is determined that the k-th frame filter coefficient of the first FF filter has reached the convergence stability condition. If the value of the loss function does not reach the minimum threshold, it is determined that the k-th frame filter coefficient of the first FF filter has not reached the convergence stability condition. The minimum threshold is preset and may be adjusted according to other requirements depending on the case.
[0127] Optionally, the filter coefficients of each FF filter include at least one biquad filter coefficient and one gain. Variables corresponding to the biquad filter coefficient include the filter type, cutoff frequency, and quality factor. Of course, in actual applications, the filter coefficients of each FF filter may include more or fewer other parameters. This is not limited to the embodiments of this application.
[0128] The k-th frame filter coefficient of the first FF filter can be determined according to a relevant algorithm based on the value of the filter coefficient variable. This algorithm is not limited to the embodiments of the present application.
[0129] In some cases, background noise, or the noise floor problem, may occur in quiet environments. For example, semi-open headsets are more likely to have background noise problems than in-ear headsets in quiet environments. Additionally, powerful noise cancellation is not required in quiet environments, and some people may experience discomfort when powerful noise cancellation is performed in such environments. Furthermore, a higher noise cancellation intensity indicates a stronger sense of sound pressure for the user. Therefore, when determining the values of filter coefficient variables according to the gradient descent method, the subjective experience effect of adaptive noise cancellation can be enhanced by dynamically adjusting the target noise cancellation amplitude according to the ambient volume, so that the k-th frame filter coefficient of the first FF filter is determined based on the target noise cancellation amplitude. That is, the target noise cancellation amplitude is determined based on the ambient volume of the (k-1)-th frame and the ambient volume of the t frames prior to the (k-1)-th frame, where t is greater than or equal to 1 and less than k-1. The value of the filter coefficient variable is determined based on the loss function according to the target noise removal amplitude and the gradient descent method, and the k-th frame filter coefficient of the first FF filter is determined based on the value of the filter coefficient variable.
[0130] The target ambient sound level is determined based on the ambient sound level of the (k-1)th frame and the ambient sound levels of the t frames prior to the (k-1)th frame. If the target ambient sound level is less than or equal to the first sound level threshold, the first noise reduction amplitude is determined as the target noise reduction amplitude. If the target ambient sound level is greater than the first sound level threshold, it is determined whether the target ambient sound level has increased significantly or decreased significantly. If the target ambient sound level has increased significantly, the noise reduction amplitude of the (k-1)th frame is increased to obtain the target noise reduction amplitude. If the target ambient sound level has decreased significantly, the noise reduction amplitude of the (k-1)th frame is decreased to obtain the target noise reduction amplitude. If the target ambient sound level has neither increased significantly nor decreased significantly, the noise reduction amplitude of the (k-1)th frame is determined as the target noise reduction amplitude, i.e., the noise reduction amplitude remains unchanged.
[0131] For example, there are multiple methods for determining a target ambient sound level based on the ambient sound level of the (k-1)th frame and the ambient sound levels of the t frames prior to the (k-1)th frame, such as calculating an arithmetic mean or a weighted mean. This is not limited to the embodiments of the present application. The t frames may be any t frames prior to the (k-1)th frame, or t frames prior to the (k-1)th frame that are closest to the (k-1)th frame. This is not limited to the embodiments of the present application.
[0132] The first volume threshold is preset and indicates whether the current environment is quiet. That is, if the target environment volume is less than or equal to the first volume threshold, it indicates that the environment is quiet. If the target environment volume is greater than the first volume threshold, it indicates that the environment is not quiet. The first noise cancellation amplitude is preset for quiet environments and is used to perform mild noise cancellation so that background noise is not excessively amplified or causes subjective comfort issues. In actual application, the first volume threshold and the first noise cancellation amplitude can be adjusted according to various requirements.
[0133] For example, refer to Fig. 3. Whether the environment is quiet is determined based on the target environment volume, and if the environment is quiet, the first noise reduction amplitude is determined as the target noise reduction amplitude. In a non-quiet environment, if the target environment volume increases significantly, the (k-1)th frame noise reduction amplitude is increased to obtain the target noise reduction amplitude. If the target environment volume decreases significantly, the (k-1)th frame noise reduction amplitude is decreased to obtain the target noise reduction amplitude. If the target environment volume does not increase significantly and does not decrease significantly, the (k-1)th frame noise reduction amplitude is determined as the target noise reduction amplitude, that is, the noise reduction amplitude remains unchanged.
[0134] There are multiple methods for determining whether there has been a significant increase or a significant decrease in the target ambient volume. For example, if the target ambient volume determined in this instance is greater than the last determined target ambient volume, and the difference between the current and last determined target ambient volumes is greater than the second volume threshold, it is determined that the current target ambient volume has significantly increased. Similarly, if the target ambient volume determined in this instance is smaller than the last determined target ambient volume, and the difference between the current and last determined target ambient volumes is greater than the second volume threshold, it is determined that the current target ambient volume has significantly decreased.
[0135] A second volume threshold is also preset, for example, to 3dB. In actual application, the second volume threshold can be further adjusted according to various requirements.
[0136] In the second case, the headset includes one additional FB filter but does not include at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the FB filter.
[0137] Here, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the FB filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0138] When the k-th frame frequency response information of the first FF filter is determined, the residual error can be determined based on the (k-1)th frame reference signal collected by at least one first reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter can be determined based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, the residual error, and the (k-1)th frame filter coefficient of the FB filter.
[0139] For example, frequency response information of the (k-1)th frame filter coefficient of the target SP and frequency response information of the (k-1)th frame filter coefficient of the FB filter can be determined. Then, based on the (k-1)th frame frequency response information of the first FF filter, the frequency response information of the (k-1)th frame filter coefficient of the target SP, the residual error, and the frequency response information of the (k-1)th frame filter coefficient of the FB filter, the k-th frame frequency response information of the first FF filter is determined according to the following formula (3).
[0140] (3)
[0141] In the previous formula (3), represents the frequency response information of the (k-1)th frame filter coefficient of the FB filter. The meaning indicated by other letters is the same as the meaning of the preceding formula (2). Details are not described again in this specification.
[0142] The implementation process for determining the residual error and the implementation process for determining the k-th frame filter coefficient of the first FF filter based on the k-th frame frequency response information of the first FF filter are identical to the implementation process of the first case. For specific implementation processes, refer to the preceding description. Details are not described again in this specification.
[0143] In the third case, the headset does not include an FB filter, but includes at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame frequency response information of at least one second FF filter.
[0144] In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target SP, and the (k-1)th frame frequency response information of at least one second FF filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0145] When the k-th frame frequency response information of the first FF filter is determined, the residual error may be determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone and the (k-1)-th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter may be determined based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the target SP, the residual error, and the (k-1)-th frame response information of at least one second FF filter.
[0146] For example, frequency response information of the (k-1)th frame filter coefficient of the target SP can be determined. Then, based on the (k-1)th frame frequency response information of the first FF filter, the frequency response information of the (k-1)th frame filter coefficient of the target SP, the residual error, and the (k-1)th frame frequency response information of at least one second FF filter, the k-th frame frequency response information of the first FF filter can be determined according to the following formula (4).
[0147] (4)
[0148] In the aforementioned formula (4), is in at least one second FF filter j Represents the (k-1)th frame frequency response information of the th FF filter, and h represents the total quantity of at least one second FF filter. The meaning indicated by other characters is the same as the meaning of the preceding formula (2). Details are not described again in this specification.
[0149] The implementation process for determining the residual error and the implementation process for determining the k-th frame filter coefficient of the first FF filter based on the k-th frame frequency response information of the first FF filter are identical to the implementation process of the first case. For specific implementation processes, refer to the preceding description. Details are not described again in this specification. Additionally, the (k-1)-th frame frequency response information of at least one second FF filter may be determined according to a relevant algorithm based on each (k-1)-th frame filter coefficient. This algorithm is not limited to the embodiments of this application.
[0150] In the fourth case, the headset includes an FB filter and at least one second FF filter. In this case, the k-th frame filter coefficient of the first FF filter is determined based on the (k-1)-th frame reference signal collected by at least one first reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target SP, the (k-1)-th frame filter coefficient of the FB filter, and the (k-1)-th frame frequency response information of at least one second FF filter.
[0151] In this case, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target SP, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame frequency response information of at least one second FF filter. The k-th frame filter coefficient of the first FF filter is determined based on the k-th frame frequency response information of the first FF filter.
[0152] When the k-th frame frequency response information of the first FF filter is determined, the residual error may be determined based on the (k-1)th frame reference signal collected by at least one first reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the first FF filter may be determined based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, the residual error, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame frequency response information of at least one second FF filter.
[0153] For example, frequency response information of the (k-1)th frame filter coefficient of the target SP and frequency response information of the (k-1)th frame filter coefficient of the FB filter can be determined. Then, based on the (k-1)th frame frequency response information of the first FF filter, the frequency response information of the (k-1)th frame filter coefficient of the target SP, residual error, frequency response information of the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame frequency response information of at least one second FF filter, the k-th frame frequency response information of the first FF filter is determined according to the following formula (5).
[0154] (5)
[0155] In the previous formula (5), represents the frequency response information of the (k-1)th frame filter coefficient of the FB filter. The meaning indicated by other characters is the same as the meaning of the preceding formula. Details are not described again in this specification.
[0156] The implementation process for determining the residual error and the implementation process for determining the k-th frame filter coefficients of the first FF filter based on the k-th frame frequency response information of the first FF filter are identical to the implementation process of the first case. For specific implementation processes, refer to the preceding description. Details are not described again in this specification. Additionally, the frequency response information of the filter coefficients of the SP can be determined according to a relevant algorithm based on the filter coefficients of the SP, and the frequency response information of the FB filter coefficients can be determined according to a relevant algorithm based on the filter coefficients of the FB filter. This algorithm is not limited to the embodiments of this application.
[0157] In the process of determining the k-th frame frequency response information of the first FF filter described above, regardless of whether the headset includes an FB filter, the k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the target SP is determined based on the target noise removal level by querying the mapping relationship between the noise removal level and the filter coefficient of the SP. Specifically, the (k-1)-th frame filter coefficient of the target SP is an estimated value, and by determining the k-th frame frequency response information of the first FF filter based on the estimated value, dependency on the actual value of the target SP can be eliminated, and filter coefficient adaptation of the FF filter can be implemented even in the absence of a downlink signal.
[0158] (2) Determine the k-th frame filter coefficient of the FB filter.
[0159] When k is 1, the initial filter coefficient of the FB filter is determined as the k-th frame filter coefficient of the FB filter, that is, the first frame filter coefficient of the FB filter is determined as the initial filter coefficient of the FB filter, or the k-th frame filter coefficient of the FB filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the FB filter coefficient. When K is greater than 1, the k-th frame filter coefficient of the FB filter may be determined based on the target noise removal level. The initial noise removal coefficient includes the initial filter coefficient of the FB filter, and the initial filter coefficient may or may not be 0. This is not limited to the embodiments of the present application.
[0160] When k is greater than 1, the k-th frame filter coefficient of the FB filter can be determined in the following two ways.
[0161] In the first method, the k-th frame filter coefficient of the FB filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the FB filter coefficient.
[0162] The mapping relationship between the noise removal level and the FB filter coefficients includes multiple noise removal levels, and a mapping relationship exists between each noise removal level and the filter coefficient of the FB filter, and the mapping relationships between different noise removal levels and the filter coefficients of the FB filter may differ. Accordingly, based on the target noise removal level, filter coefficients corresponding to the FB filter can be obtained from the mapping relationship between the noise removal level and the FB filter coefficients, and the obtained filter coefficients are used as the k-th frame filter coefficients of the FB filter.
[0163] Since the mapping relationship between the noise removal level and the FB filter coefficients is stored in advance, determining the k-th frame filter coefficient of the FB filter in the first method is stable, simple to operate, and highly efficient.
[0164] In the second method, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the FB filter, and the target noise removal level.
[0165] Similar to what was explained earlier, when k is greater than 1, the k-th frame filter coefficient of the FB filter can be determined according to the adaptation method. The process of determining the k-th frame filter coefficient of the FB filter is an adaptation process, which can also be called an iterative process.
[0166] An implementation process for determining the k-th frame filter coefficient of an FB filter based on the (k-1)th frame error signal collected by an error microphone, the (k-1)th frame filter coefficient of an FB filter, and the target noise removal level comprises: the step of determining the (k-1)th frame filter coefficient of a target SP based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficient of the SP—wherein the target SP is the path from the target speaker to the error microphone—and the step of determining the k-th frame filter coefficient of an FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of an FB filter, and the (k-1)th frame filter coefficient of a target SP.
[0167] The k-th frame filter coefficient of the FB filter can be determined according to a relevant algorithm based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the (k-1)-th frame filter coefficient of the target SP. This algorithm is not limited to the embodiments of the present application.
[0168] The first frame filter coefficient of the FB filter can be determined based on the initial noise removal level by querying the mapping relationship between the noise removal level and the FB filter coefficient. Thus, when k is greater than or equal to 1, the k-th frame filter coefficient of the FB filter can be determined in two ways: (1) The k-th frame filter coefficient of the FB filter is determined by querying the mapping relationship between the noise removal level and the FB filter coefficient. (2) When k is 1, the k-th frame filter coefficient of the FB filter is determined by querying the mapping relationship between the noise removal level and the FB filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal level.
[0169] In the second method described above, the method of querying the mapping relationship between the noise removal level and the FB filter coefficients is combined with an adaptive method to improve the noise removal effect and ensure controllable stability with low complexity.
[0170] It should be noted that in the embodiments of the present application, the k-th frame filter coefficient of the FB filter can be determined in the two ways described above, and alternatively, the k-th frame filter coefficient of the FB filter can be determined in another way. For example, regardless of whether k is greater than 1 or equal to 1, the k-th frame filter coefficient of the FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal level. This is not limited to the embodiments of the present application.
[0171] (3) Determine the k-th frame filter coefficient of the downlink compensation filter.
[0172] When k is 1, the initial filter coefficient of the downlink compensation filter is determined as the k-th frame filter coefficient of the downlink compensation filter, or the k-th frame filter coefficient of the downlink compensation filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the downlink compensation filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the downlink compensation filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the downlink compensation filter coefficient.
[0173] The mapping relationship between noise reduction levels and downlink compensation filter coefficients involves multiple noise reduction levels, and a mapping relationship exists between each noise reduction level and the filter coefficients of the downlink compensation filter; furthermore, the mapping relationships between different noise reduction levels and the filter coefficients of the downlink compensation filter may differ. Therefore, after the target noise reduction level is determined, the corresponding downlink compensation filter coefficients can be obtained from the mapping relationship between the noise reduction level and the downlink compensation filter coefficients based on the target noise reduction level, and the obtained downlink compensation filter coefficients are used as the k-th frame filter coefficients of the downlink compensation filter.
[0174] (4) Determine the target noise reduction level.
[0175] The noise reduction level of the (k-1)th frame is determined, and the noise reduction levels of m frames prior to the (k-1)th frame are obtained, where m is greater than or equal to 1 and less than k-1. The target noise reduction level is determined based on the noise reduction level of the (k-1)th frame and the noise reduction levels of the m frames.
[0176] In the (k-1)th frame, a valid downlink signal may or may not exist, the environment may or may not be quiet, or, of course, an abnormal signal may exist. Depending on the case, the method for determining the noise cancellation level of the (k-1)th frame varies, and this is explained separately below.
[0177] In the first case, in the (k-1)th frame, there is no effective downlink signal and the environment is not quiet. In this case, the noise removal level of the (k-1)th frame is determined based on the reference filter coefficients of the first FF filter and the mapping relationship between the noise removal level and the frequency response information of the first FF filter. When k is 2, the reference filter coefficients are the initial filter coefficients of the first FF filter, or when k is greater than 2, the reference filter coefficients are the filter coefficients of the first FF filter that last satisfied the convergence stability condition before the kth frame, or the (k-1)th frame filter coefficients of the first FF filter.
[0178] For example, when an audio signal is played through the headset, such as when music is played or a call is made, the user terminal transmits control signaling for audio signal playback to the headset. Therefore, whether the headset is currently in a downlink active state can be determined based on whether the headset receives the control signaling. If the headset is not in a downlink active state, it is determined that there is no valid downlink signal in the (k-1)th frame. If the headset is in a downlink active state, sound may not be output continuously in the (k-1)th frame. For example, sound is not output during pauses in voice or transitions where music changes, and generally, these periods are not short. Therefore, if the headset is in a downlink active state, it can be further determined whether the (k-1)th frame is in a downlink intermittent period. If the (k-1)th frame is in a downlink intermittent period, it is determined that there is no valid downlink signal in the (k-1)th frame. If the (k-1)th frame is not in the downlink intermittent period, it is determined that a valid downlink signal exists in the (k-1)th frame.
[0179] In the case where there is no effective downlink signal in the (k-1)th frame but the environment is not quiet, the noise removal level of the (k-1)th frame may vary depending on different environmental noise. Therefore, the noise removal level of the (k-1)th frame must be determined based on the mapping relationship between the reference filter coefficient of the first FF filter, the noise removal level, and the frequency response information of the first FF filter.
[0180] In some embodiments, the reference frequency response information of the first FF filter is determined based on the reference filter coefficients of the first FF filter. The noise removal level corresponding to the reference frequency response information of the first FF filter is determined based on the mapping relationship between the noise removal level and the frequency response information of the first FF filter to obtain the (k-1)th frame noise removal level.
[0181] The reference frequency response information of the first FF filter can be determined according to a relevant algorithm based on the reference filter coefficients of the first FF filter. This algorithm is not limited to the embodiments of the present application.
[0182] If the noise removal level is different, the frequency response information of the FF filter may also be different. Therefore, a mapping relationship between the noise removal level and the frequency response information of the first FF filter can be stored in advance. In this way, after the reference frequency response information of the first FF filter is determined, a matching is performed between the reference frequency response information of the first FF filter and the frequency response information of the FF filter at different noise removal levels within the mapping relationship, and from the mapping relationship, the frequency response information that matches the reference frequency response information of the first FF filter is determined, and then the noise removal level corresponding to the matched frequency response information is used as the (k-1)th frame noise removal level.
[0183] The frequency response information of an FF filter can be expressed using a frequency response curve. Therefore, after the reference frequency response curve of the first FF filter is determined, matching can be performed between the reference frequency response curve of the first FF filter and the frequency response curve of the FF filter at different noise removal levels within the mapping relationship.
[0184] In actual application, matching may be performed between the complete reference frequency response curve of the first FF filter and the complete frequency response curve of the FF filter at different noise removal levels in the mapping relationship. Alternatively, matching may be performed between a curve that is within the reference frequency response curve of the first FF filter and belongs to the target frequency band, and a curve that is within the frequency response curve of the FF filter at different noise removal levels in the mapping relationship and belongs to the target frequency band. This is not limited to the embodiments of the present application.
[0185] It should be noted that the target frequency band is a frequency band with distinct characteristics in the frequency response curve, and that the target frequency band is preset. For example, the target frequency band is a frequency band from 100 Hertz to 200 Hertz (Hz). Of course, the value of the target frequency band may vary depending on the different acoustic conditions of the headset.
[0186] For example, the mapping relationship between the noise removal level and the frequency response information of the first FF filter includes the frequency response curve of the FF filter at 16 noise removal levels, and the frequency response curve of the FF filter at 16 noise removal levels is shown in FIG. 4. In FIG. 4, since the characteristics of the 100Hz–200Hz frequency band are clearly distinguished, the 100Hz–200Hz frequency band is used as the target frequency band. Then, matching is performed between the curve within the reference frequency response curve of the first FF filter that belongs to the 100Hz–200Hz range and the curve within the frequency response curve of the FF filter at 16 noise removal levels that belongs to the 100Hz–200Hz range.
[0187] The process of determining the filter coefficients of the first FF filter is an iterative process and can also be called an adaptive process; therefore, the aforementioned convergence stability condition indicates that the filter coefficients of the first FF filter converge to a state where they basically do not change. Additionally, since the filter coefficients of the first FF filter can be adaptively adjusted multiple times during the entire noise removal process, when the k-th frame filter coefficients of the first FF filter are determined, the filter coefficients of the first FF filter that last satisfied the convergence stability condition prior to the k-th frame can be used as reference filter coefficients, or the (k-1)-th frame filter coefficients of the first FF filter can be used as reference filter coefficients.
[0188] In the second case, an effective downlink signal exists in the (k-1)th frame. In this case, the noise cancellation level of the (k-1)th frame is determined based on the effective downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone.
[0189] In light of the above description, if the headset is in a downlink active state and is not in a downlink intermittent period, it is determined that an effective downlink signal exists in the (k-1)th frame. In this case, based on the (k-1)th frame effective downlink signal, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone, the effective downlink signal can be extracted from the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame noise removal level can be determined based on the extracted effective downlink signal.
[0190] Based on the (k-1)th frame effective downlink signal, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone, the effective downlink signal is extracted from the (k-1)th frame error signal collected by the error microphone according to a relevant algorithm, and the (k-1)th frame noise removal level can be determined based on the extracted effective downlink signal. This algorithm is not limited to the embodiments of the present application.
[0191] In the third case, there is no valid downlink signal in the (k-1)th frame and the environment is quiet, or there is an abnormal noise signal in the (k-1)th frame. In this case, the noise cancellation level of the (k-3)th frame is determined to be the noise cancellation level of the (k-1)th frame. That is, the noise cancellation level is not changed.
[0192] In the (k-1)th frame, if there is no valid downlink signal and the surrounding environment is quiet, the noise is essentially unchanged. In this case, the noise cancellation level may not change. If an abnormal noise signal is present in the (k-1)th frame, the noise cancellation level remains unchanged to perform robust control and prevent the divergence of the noise cancellation level.
[0193] Abnormal noise signals refer to signals that significantly affect the user's listening environment, such as howling, clipping, background noise, and wind noise. Howling is a phenomenon where the amplitude or energy of a single-frequency sound signal suddenly increases from a small value; it is generally caused by actions such as the user gripping the headset tightly or rapidly changing their posture while wearing it. The sound signal emitted during howling is called howling noise. Howling causes discomfort to the user, interferes with the playback of downlink signals, and severely affects audio playback effects. Clipping is a phenomenon where low-frequency signals overflow, generating cracking noise; this generated cracking noise is called clipping noise. Generally, clipping occurs when loud low-frequency noise bursts out in the environment. For example, loud low-frequency noise is generated when a vehicle crashes or an airplane lands. Background noise refers to ground noise, also known as the floor noise. Background noise is noise caused by performance limitations of the device's hardware (e.g., the circuitry or other components of a headset); for instance, it is noise such as rustling sounds in television audio that are separate from the program sound. In noisy environments, users cannot perceive or hear background noise. In quiet environments, users can perceive background noise. If background noise is too loud, it not only annoys the user but also obscures subtle details of the sound. Wind noise occurs when wind blows in the environment. Wind noise affects the user's normal use of the headset. Furthermore, because the direction of wind noise is random, its impact on the user's ears varies. In other words, the left and right ears experience different auditory sensations under the influence of wind noise.
[0194] The following is a brief summary of the three cases described above with reference to Fig. 5. Referring to Fig. 5, if an abnormal noise signal is present in the (k-1)th frame, the noise removal level does not change. If no abnormal noise signal is present in the (k-1)th frame, the downlink activation status is determined in the (k-1)th frame. If downlink activation is not performed in the (k-1)th frame, it is determined whether the environment is quiet in the (k-1)th frame. If the environment is quiet in the (k-1)th frame, the noise removal level does not change. If the environment is not quiet in the (k-1)th frame, the noise removal level of the (k-1)th frame is determined based on the reference filter coefficients of the FF filter. If downlink activation is performed in the (k-1)th frame, it is determined whether the (k-1)th frame is in a downlink intermittent period. If the (k-1)th frame is in a downlink intermittent period, the noise removal level of the (k-1)th frame is also determined in the same manner. If the (k-1)th frame is not in a downlink intermittent period, the (k-1)th frame noise removal level is determined based on the (k-1)th frame effective downlink signal, the (k-1)th frame reference signal collected by at least one first reference microphone, and the (k-1)th frame error signal collected by the error microphone.
[0195] After the noise reduction level of the (k-1)th frame is determined in the three cases above, the target noise reduction level can be determined by integrating the noise reduction level of the (k-1)th frame with the noise reduction levels of the m frames prior to the (k-1)th frame.
[0196] The noise reduction level in m frames may be the noise reduction level in any m frames prior to the (k-1)th frame, or the noise reduction level in the m frames prior to the (k-1)th frame and closest to the (k-1)th frame. This is not limited to the embodiments of the present application. Additionally, there are multiple implementations for determining a target noise reduction level based on the noise reduction level in the (k-1)th frame and the noise reduction levels in the m frames prior to the (k-1)th frame. For example, the noise reduction effect may be evaluated according to a relevant algorithm to determine the noise reduction probability corresponding to the (k-1)th frame noise reduction level and the noise reduction probability corresponding to the noise reduction level in m frames, and the noise reduction level with the highest noise reduction probability may be determined as the target noise reduction level. Alternatively, the target noise reduction level may be obtained by determining the arithmetic mean or weighted mean of the noise reduction level in the (k-1)th frame and the noise reduction levels in m frames. Alternatively, the noise reduction level that appears most frequently among the (k-1)th frame noise reduction level and the noise reduction levels in m frames is determined as the target noise reduction level, etc.
[0197] The various mapping relationships described above are predetermined. For example, when at least one speaker is operating and the other speaker is not operating, the various mapping relationships are determined based on a reference signal collected by at least one first reference microphone and an error signal collected by an error microphone in each of the plurality of leakage states. The plurality of leakage states are formed by the headset and the plurality of different external auditory canal environments, and the plurality of leakage states have a one-to-one correspondence with the plurality of noise cancellation levels.
[0198] In this case, the target noise removal parameter is determined. Using Fig. 6 as an example, the process of determining the target noise removal parameter is briefly summarized as follows. Refer to Fig. 6. Initial values, including the previously described initial noise removal level, initial filter coefficients, and various mapping relationships, can be set offline. Subsequently, in the (k-1)th frame, the (k-1)th frame noise removal level is determined according to various cases by determining whether there is a valid downlink signal, whether the environment is quiet, and whether there is an abnormal noise signal. The target noise removal level is determined based on the (k-1)th frame noise removal level and the previous noise removal levels of m frames. Then, the target noise removal amplitude is determined based on the (k-1)th frame environment loudness, and FB filter coefficient adaptation is performed based on the target noise removal level to determine the k-th frame filter coefficient of the FB filter. Finally, FF filter coefficient adaptation is performed based on the target noise removal level and the target noise removal amplitude to determine the k-th frame filter coefficient of the first FF filter.
[0199] Step (202): Based on the target noise removal parameter, noise removal is performed through at least one of the target speakers.
[0200] Target inverse phase noise is generated based on target noise removal parameters, and noise removal is performed through the target speaker among at least one speaker based on the target inverse phase noise.
[0201] If the target noise removal parameter includes the k-th frame filter coefficient of the first FF filter, the target inverse phase noise includes feedforward inverse phase noise. In this case, the k-th frame reference signal collected by at least one first reference microphone can be processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise.
[0202] In light of the above description, at least one first reference microphone may include one reference microphone or at least two reference microphones. If at least one first reference microphone includes one reference microphone, the k-th frame reference signal collected by the reference microphone is directly processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise. If at least one first reference microphone includes at least two reference microphones, audio mixing is performed on the k-th frame reference signal collected by at least two reference microphones to obtain a k-th frame mixed reference signal, and then the k-th frame mixed reference signal is processed based on the k-th frame filter coefficient of the first FF filter to obtain feedforward inverse phase noise.
[0203] Optionally, the headset may further include at least one second FF filter. In this case, the k-th frame filter coefficient of at least one second FF filter may be determined. Noise removal is performed through a target speaker based on target noise removal parameters and the k-th frame filter coefficient of at least one second FF filter. That is, the k-th frame reference signal collected by at least one first reference microphone is processed based on the k-th frame filter coefficient of the first FF filter to obtain a first feedforward inverse phase noise. The k-th frame reference signal collected by at least one first reference microphone is processed based on the k-th frame filter coefficient of at least one second FF filter to obtain at least one second feedforward inverse phase noise.
[0204] If k is 1, the initial filter coefficient of at least one second FF filter is determined as the k-th frame filter coefficient of at least one second FF filter, that is, the first frame filter coefficient of at least one second FF filter is the initial filter coefficient of the corresponding second FF filter, or the k-th frame filter coefficient of at least one second FF filter is determined based on the initial noise removal level and the mapping relationship between the noise removal level and the second FF filter coefficient. If k is greater than 1, the k-th frame filter coefficient of at least one second FF filter is determined based on the target noise removal level and the mapping relationship between the noise removal level and the second FF filter coefficient. The initial noise removal coefficient includes the initial filter coefficient of at least one second FF filter, and the initial filter coefficient may be 0 or not 0. This is not limited to the embodiments of the present application.
[0205] The mapping relationship between the noise removal level and the second FF filter coefficient includes multiple noise removal levels, and a mapping relationship exists between each noise removal level and the filter coefficient of at least one second FF filter, and the mapping relationships between different noise removal levels and the filter coefficient of at least one second FF filter may differ. Accordingly, after the target noise removal level is determined, filter coefficients corresponding to at least one second FF filter can be obtained from the mapping relationship between the noise removal level and the second FF filter coefficients based on the target noise removal level, and the obtained filter coefficients are used as the k-th frame filter coefficients of at least one second FF filter. The same applies to the initial noise removal level.
[0206] The foregoing description is provided using an example in which the first FF filter and at least one second FF filter both correspond to at least one first reference microphone. In actual applications, the first FF filter and at least one second FF filter may correspond to different reference microphones. For example, a headset further includes a plurality of second reference microphones, the first FF filter corresponds to at least one first reference microphone, and each second FF filter corresponds to at least one of the plurality of second reference microphones. In this case, the k-th frame reference signal collected by at least one first reference microphone can be processed based on the k-th frame filter coefficient of the first FF filter to obtain the first feedforward inverse phase noise. The k-th frame reference signal collected by at least one second reference microphone corresponding to each second FF filter can be processed based on the k-th frame filter coefficient of each second FF filter to obtain at least one second feedforward inverse phase noise.
[0207] If the headset additionally includes an FB filter, the target inverse phase noise further includes feedback inverse phase noise. That is, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal. Specifically, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal based on the k-th frame filter coefficients of the downlink compensation filter. Then, after inversion is performed on the k-th frame downlink signal obtained through downlink compensation, audio mixing is performed on the inverted k-th frame downlink signal and the k-th frame error signal collected by the error microphone to obtain the k-th frame noise signal collected by the error microphone. The k-th frame noise signal collected by the error microphone is processed based on the k-th frame filter coefficients of the FB filter to obtain feedback inverse phase noise.
[0208] Downlink compensation can be used to prevent sound quality degradation of the downlink signal by removing all downlink signals from the error signal collected by the error microphone and performing noise removal only on the residual noise signal through an FB filter. Additionally, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal to remove downlink signals from all speakers in the error microphone, thereby preventing sound quality degradation of the full-band downlink signal.
[0209] In light of the above description, when the target noise removal parameter is determined on a frame-by-frame basis, since a single frame may include one sample point or multiple sample points, when the target inverse phase noise is generated, a target inverse phase noise group may be generated at each sample point or a target inverse phase noise group may be generated in a single frame.
[0210] In an embodiment of the present application, when the target noise removal parameter is determined, frequency division is not performed on the downlink signal; that is, the target noise removal parameter is determined based on the full-band downlink signal. In this manner, after the target inverse phase noise is generated based on the target noise removal parameter, the frequency band of the target inverse phase noise includes the sound generation frequency band of at least one speaker; that is, the frequency band of the target inverse phase noise is the entire frequency band.
[0211] After the target inverse phase noise is generated, the target inverse phase noise is mixed with the k-th frame downlink signal to be played through the target speaker, and then the mixed signal is played through the target speaker to achieve noise removal.
[0212] At least one speaker may include a single speaker or a plurality of speakers. If at least one speaker includes a single speaker, the speaker may be a full-band speaker. If at least one speaker includes a plurality of speakers, some of the plurality of speakers may be high-band speakers and others may be low-band speakers. Alternatively, some of the plurality of first speakers may be full-band speakers and others may be non-full-band speakers. In other words, the sound generation frequency bands of the plurality of speakers may differ from one another. Alternatively, all of the plurality of speakers may be full-band speakers. Alternatively, all of the plurality of speakers may be non-full-band speakers. If all of the plurality of speakers are full-band speakers, the k-th frame downlink signal played through the plurality of speakers is the k-th frame downlink signal transmitted from the user terminal. If not all of the plurality of speakers are full-band speakers, frequency division must be performed on the k-th frame downlink signal transmitted by the user terminal based on the sound generation frequency band of each speaker to obtain the k-th frame downlink signal to be played through each speaker.
[0213] In an embodiment of the present application, noise cancellation is performed through a target speaker among at least one speaker. If the at least one speaker includes one speaker, that speaker is the target speaker. If the at least one speaker includes a plurality of speakers and the plurality of speakers includes a first speaker and a second speaker to which digital frequency division is performed, the target speaker is the first speaker, and the second speaker does not participate in noise cancellation. However, the second speaker may participate in downlink compensation (i.e., downlink compensation is performed on a downlink signal transmitted by a user terminal, wherein the downlink signal is a full-band audio signal including an audio signal in the sound generation frequency band of the second speaker). In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker or a mid-band speaker or a low-band speaker. Optionally, the second speaker may not participate in downlink compensation. In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker.
[0214] Optionally, at least one speaker may include a first speaker and a second speaker on which analog frequency division is performed as an alternative. In this case, the target speakers are the first speaker and the second speaker, that is, both the first speaker and the second speaker participate in noise cancellation.
[0215] If the target speakers are the first and second speakers, the first and second speakers may be an analog frequency division combination of the two speakers. That is, the first and second speakers are driven using the same DAC and PA, and the combination of the first and second speakers can be considered as a single speaker.
[0216] It should be noted that the sound generation frequency band of at least one second speaker is higher than the sound generation frequency band of at least one first speaker. Of course, the sound generation frequency band of at least one second speaker may alternatively be lower than the sound generation frequency band of at least one first speaker. This is not limited to the embodiments of the present application.
[0217] Furthermore, the process of determining the target noise removal parameter according to the aforementioned adaptation method requires a specific amount of time. If a single frame contains multiple sample points and the duration of the frame is long, the duration for determining the target noise removal parameter may be shorter than the duration of the frame. Therefore, based on the relevant data of the (k-1)th frame, calculations can be performed during a portion of the duration of the k-th frame to obtain the target noise removal parameter for the k-th frame, and active noise removal can be performed during another portion of the duration of the k-th frame based on the target noise removal parameter. However, if a single frame contains only one sample point, or if a single frame contains multiple sample points and the duration of the frame is short, the duration for determining the target noise removal parameter may be equal to the duration of the frame. In this case, to obtain the target noise removal parameter, calculations may need to be performed over the entire duration of the k-th frame based on the relevant data of the (k-1)th frame. In this case, the target noise reduction parameter can be determined as the target noise reduction parameter of the (k+1)th frame, and active noise reduction can be performed during the period of the (k+1)th frame based on the target noise reduction parameter of the (k+1)th frame. The preceding content is explained using the latter case as an example.
[0218] In conclusion, in an embodiment of the present application, the target noise removal parameter is determined based on a reference signal collected by at least one first reference microphone, an error signal collected by an error microphone, and an initial noise removal coefficient. This eliminates dependency on the downlink signal, so that the target noise removal parameter can be determined and adaptive noise removal can be performed even in the absence of a downlink signal.
[0219] The following describes several possible headset architectures in the embodiments of the present application.
[0220] FIG. 7 is a diagram of the structure of a headset according to an embodiment of the present application. Refer to FIG. 7. The headset includes f first reference microphones, one error microphone, one first FF filter, one FF adaptation engine, one FB filter, one FB adaptation engine, h second FF filters, n speakers, downlink compensation filters, downlink compensation adaptation engines (not shown in the drawing), and a digital frequency divider. At the same noise rejection level, the filter coefficients of each second FF filter are fixed, and f and n are both integers greater than or equal to 1, and f and n may be equal or not equal.
[0221] F first reference microphones are configured to collect noise signals from the external environment, i.e., reference signals. Error microphones are configured to collect noise signals from the external auditory canal, i.e., error signals. The FF adaptation engine is configured to determine the k-th frame filter coefficients of the first FF filter and to refresh the determined k-th frame filter coefficients into the first FF filter. The FB adaptation engine is configured to determine the k-th frame filter coefficients of the FB filter and to refresh the determined k-th frame filter coefficients into the FB filter. The downlink compensation adaptation engine is configured to determine the k-th frame filter coefficients of the downlink compensation filter and to refresh the determined k-th frame filter coefficients into the downlink compensation filter. A digital frequency divider is configured to perform frequency division on the k-th frame downlink signal transmitted by the user terminal to obtain the k-th frame downlink signal corresponding to each speaker.
[0222] When noise removal is performed, the first FF filter is configured to process the k-th frame reference signal collected by f first reference microphones based on the k-th frame filter coefficients to obtain the first feedforward inverse phase noise. Each second FF filter is configured to process the k-th frame reference signal collected by f first reference microphones based on the k-th frame filter coefficients of the filter to obtain the second feedforward inverse phase noise. The downlink compensation filter is configured to perform downlink compensation on the k-th frame downlink signal transmitted by the user terminal based on the k-th frame filter coefficients of the downlink compensation filter. Then, after inversion is performed on the k-th frame downlink signal obtained through downlink compensation, audio mixing is performed on the inverted k-th frame downlink signal and the k-th frame error signal collected by the error microphone to obtain the k-th frame noise signal collected by the error microphone. The FB filter is configured to process the k-th frame noise signal collected by the error microphone based on the k-th frame filter coefficients of the FB filter to obtain feedback inverse phase noise. Then, after performing audio mixing on the first feedforward inverse phase noise, the second feedforward inverse phase noise, the feedback inverse phase noise, and the k-th frame downlink signal of the target speaker (i.e., speaker 1) among n speakers, the acquired signal is played back through the target speaker to implement noise removal.
[0223] FIG. 8 is a diagram of the structure of another headset according to an embodiment of the present application. Refer to FIG. 8. The headset includes one first reference microphone, one error microphone, one first FF filter, one FF adaptation engine, one FB filter, one FB adaptation engine, one second FF filter, one speaker, a downlink compensation filter, and a downlink compensation adaptation engine (not shown in the drawing). The speaker is a target speaker.
[0224] FIG. 9 is a diagram of the structure of another headset according to an embodiment of the present application. Refer to FIG. 9. The headset includes one first reference microphone, one error microphone, one first FF filter, one FF adaptation engine, one FB filter, one FB adaptation engine, one speaker, a downlink compensation filter, and a downlink compensation adaptation engine (not shown in the drawing). The speaker is a target speaker.
[0225] FIG. 10 is a diagram of the structure of another headset according to an embodiment of the present application. Refer to FIG. 10. The headset includes one first reference microphone, one error microphone, one first FF filter, one FF adaptation engine, one FB filter, one FB adaptation engine, two speakers, a downlink compensation filter, and a downlink compensation adaptation engine (not shown in the drawing). The two first speakers are high-frequency and low-frequency speakers. The two first speakers are one speaker in terms of physical substance. Therefore, the headset may not include a digital frequency divider. That is, the two speakers are driven by performing analog frequency division using the same DAC and PA. When noise cancellation is performed, both speakers can be used as target speakers participating in the noise cancellation.
[0226] FIG. 11 is a diagram of the structure of another headset according to an embodiment of the present application. Refer to FIG. 11. The headset includes one first reference microphone, one error microphone, one first FF filter, one FF adaptation engine, one FB filter, one FB adaptation engine, two speakers, a downlink compensation filter, a downlink compensation adaptation engine (not shown in the drawing), and a digital frequency divider. The difference from FIG. 10 is that the two speakers are two speakers with separate loudspeakers, are driven by different DACs and PAs, and must perform digital frequency division at the two speakers. To support high sound quality, the two speakers may use high and low loudspeakers. Speaker 1 is a low-frequency unit and is the main driving loudspeaker for ANC. Speaker 2 is a high-frequency unit that provides only a sound quality channel and does not contribute to noise cancellation.
[0227] When noise cancellation is performed, one of the two speakers (i.e., Speaker 1 used as the target speaker) participates in the noise cancellation. The other speaker acting as the second speaker (i.e., Speaker 2) does not participate in the noise cancellation, but the second speaker participates in downlink compensation (i.e., downlink compensation is performed on a downlink signal transmitted by the user terminal, where the downlink signal is a full-band audio signal containing an audio signal in the sound generation frequency band of the second speaker). The first speaker may be a mid-band speaker, a low-band speaker, or a full-band speaker. The second speaker may be a high-band speaker. Of course, the second speaker may also be a mid-band speaker or a low-band speaker.
[0228] Optionally, current mainstream ANC chips may not be able to obtain signals from high-frequency speakers. Accordingly, refer to FIG. 12. The second speaker may not participate in downlink compensation (i.e., after digital frequency division is performed on the downlink signal transmitted by the user terminal, the downlink signal corresponding to the first speaker is obtained, downlink compensation is performed on the downlink signal corresponding to the first speaker, and not on the downlink signal corresponding to the second speaker). In this case, the first speaker may be a low-frequency and mid-frequency speaker or a full-frequency speaker, and the second speaker may be a high-frequency speaker. To reduce the damage of ANC to downlink sound quality, the frequency division point of the high-frequency speaker may be greater than 6 kHz, that is, audio signals greater than 6 kHz are not compensated. Of course, in the embodiments of the present application, the frequency division point of 6 kHz is not limited, and other high-frequency division points may exist.
[0229] It should be noted that the above-mentioned FF adaptation engine, FB adaptation engine, and downlink compensation adaptation engine can be placed in a microcontroller unit. The FF filter, FB filter, and downlink compensation filter can be placed in an ANC chip. The microcontroller unit and the ANC chip can be integrated into a single chip or placed on multiple chips.
[0230] FIG. 13 is a diagram of the structure of a noise removal device according to an embodiment of the present application. The noise removal device may be implemented as part or all of a headset by software, hardware, or a combination thereof. The headset may be the headset illustrated in FIG. 1. Refer to FIG. 13. The device includes a noise removal parameter determination module (1301) and a noise removal module (1302).
[0231] A noise removal parameter determination module (1301) is configured to determine a target noise removal parameter based on a reference signal collected by at least one first reference microphone, an error signal collected by an error microphone, and an initial noise removal coefficient, wherein the target noise removal parameter includes a filter coefficient of the first FF filter.
[0232] The noise removal module (1302) is configured to perform noise removal through at least one of the target speakers based on the target noise removal parameters.
[0233] Optionally, the initial noise removal coefficient includes the initial filter coefficient of the first FF filter, and the target noise removal parameter includes the k-th frame filter coefficient of the first FF filter, where k is an integer greater than or equal to 1.
[0234] The noise removal parameter determination module (1301) is:
[0235] A first FF filter coefficient determination submodule configured to determine the initial filter coefficient of the first FF filter as the k-th frame filter coefficient of the first FF filter when k is 1, or to determine the k-th frame filter coefficient of the first FF filter based on an initial noise removal level and a mapping relationship between the noise removal level and the first FF filter coefficient, or
[0236] When k is greater than 1, it includes a second FF filter coefficient determination submodule configured to determine the k-th frame filter coefficient of the first FF filter based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by an error microphone, and the target noise removal level.
[0237] Optionally, the second FF filter coefficient determination submodule specifically:
[0238] Determine the (k-1)th frame filter coefficients of the target auxiliary path (SP) based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficients of the SP—where the target SP is the path from the target speaker to the error microphone—,
[0239] It is configured to determine the k-th frame filter coefficient of the first FF filter based on the (k-1)th frame reference signal collected by at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of the target SP.
[0240] Optionally, the second FF filter coefficient determination submodule specifically:
[0241] Determining a residual error based on a (k-1)th frame reference signal collected by at least one first reference microphone and a (k-1)th frame error signal collected by an error microphone,
[0242] The k-th frame frequency response information of the first FF filter is determined based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficients of the target SP, and the residual error, and
[0243] It is configured to determine the k-th frame filter coefficients of the first FF filter based on the k-th frame frequency response information of the first FF filter.
[0244] Optionally, the headset includes one additional feedback FB filter.
[0245] The second FF filter coefficient determination submodule specifically:
[0246] It is further configured to determine the k-th frame frequency response information of the first FF filter based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the FB filter, the (k-1)-th frame filter coefficient of the target SP, and the residual error.
[0247] Optionally, the second FF filter coefficient determination submodule specifically:
[0248] A loss function is established between the filter coefficient variables of the first FF filter and the k-th frame frequency response information of the first FF filter, and
[0249] The value of the filter coefficient variable is determined based on the loss function according to the gradient descent method, and
[0250] It is configured to determine the k-th frame filter coefficient of the first FF filter based on the value of the filter coefficient variable.
[0251] Optionally, the second FF filter coefficient determination submodule specifically:
[0252] Determine the target noise reduction amplitude based on the ambient sound level, and
[0253] It is further configured to determine the values of filter coefficient variables based on the target noise removal amplitude and the loss function according to the gradient descent method.
[0254] Optionally, the headset further includes a feedback FB filter, the initial noise rejection factor includes the initial filter factor of the FB filter, and the target noise rejection parameter further includes the k-th frame filter factor of the FB filter, where k is an integer greater than or equal to 1.
[0255] The noise removal parameter determination module (1301) is:
[0256] A first FB filter coefficient determination submodule configured to determine the initial filter coefficient of the FB filter as the k-th frame filter coefficient of the FB filter when k is 1, or to determine the k-th frame filter coefficient of the FB filter based on an initial noise rejection level and a mapping relationship between the noise rejection level and the FB filter coefficient, or
[0257] When k is greater than 1, it further includes a second FB filter coefficient determination submodule configured to determine the k-th frame filter coefficient of the FB filter based on the target noise removal level and the mapping relationship between the noise removal level and the FB filter coefficient, or to determine the k-th frame filter coefficient of the FB filter based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal level.
[0258] Optionally, the second FB filter coefficient determination submodule specifically:
[0259] Determine the (k-1)th frame filter coefficients of the target auxiliary path (SP) based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficients of the SP—where the target SP is the path from the target speaker to the error microphone—,
[0260] It is configured to determine the k-th frame filter coefficient of the FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame filter coefficient of the target SP.
[0261] Optionally, the filter coefficients of the first FF filter include at least one biquad filter coefficient and one gain.
[0262] Optionally, the headset further includes at least one second FF filter.
[0263] The noise removal module (1302) specifically:
[0264] Determine the k-th frame filter coefficient of at least one second FF filter, and
[0265] It is configured to perform noise removal through a target speaker based on a target noise removal parameter and the k-th frame filter coefficient of at least one second FF filter.
[0266] Optionally, the initial noise removal coefficient includes the initial filter coefficient of at least one second FF filter.
[0267] The noise removal module (1302) specifically:
[0268] If k is 1, the initial filter coefficient of at least one second FF filter is determined as the k-th frame filter coefficient of at least one second FF filter, or the k-th frame filter coefficient of at least one second FF filter is determined based on an initial noise removal level and a mapping relationship between the noise removal level and the second FF filter coefficient, and
[0269] When k is greater than 1, it is further configured to determine the k-th frame filter coefficient of at least one second FF filter based on the target noise removal level and the mapping relationship between the noise removal level and the second FF filter coefficient.
[0270] Optionally, the headset further includes a downlink compensation filter, and the initial noise cancellation coefficient includes the initial filter coefficient of the downlink compensation filter. The noise cancellation parameter determination module (1301) is:
[0271] A first compensation coefficient determination submodule configured to determine the initial filter coefficient of the downlink compensation filter as the k-th frame filter coefficient of the downlink compensation filter when k is 1, or to determine the k-th frame filter coefficient of the downlink compensation filter based on an initial noise removal level and a mapping relationship between the noise removal level and the downlink compensation filter coefficient, or
[0272] It further includes a second compensation coefficient determination submodule configured to determine the k-th frame filter coefficient of the downlink compensation filter based on the target noise removal level and the mapping relationship between the noise removal level and the downlink compensation filter coefficient.
[0273] In conclusion, in an embodiment of the present application, the target noise removal parameter is determined based on a reference signal collected by at least one first reference microphone, an error signal collected by an error microphone, and an initial noise removal coefficient. This eliminates dependency on the downlink signal, so that the target noise removal parameter can be determined and adaptive noise removal can be performed even in the absence of a downlink signal.
[0274] It should be noted that during noise removal performed by the noise removal device provided in the embodiments, the classification of functional modules is used only as an example for illustrative purposes. In actual application, functions may be assigned to different functional modules and implemented according to requirements. That is, the internal structure of the device is divided into different functional modules to implement all or part of the functions described above. Furthermore, the embodiments of the noise removal device and the noise removal method provided in the embodiments relate to the same concept. For the specific implementation process of the noise removal device, refer to the method embodiments. Details are not described again in this specification.
[0275] Refer to FIG. 14. FIG. 14 is a diagram of the structure of another headset according to an embodiment of the present application. The headset includes one or more processors (1401), a communication bus (1402), a memory (1403), and one or more communication interfaces (1404).
[0276] The processor (1401) is a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits configured to implement the solution of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the PLD is a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a general array logic (GAL), or any combination thereof.
[0277] The communication bus (1402) is configured to transmit information between the aforementioned components. Optionally, the communication bus (1402) may be classified into an address bus, a data bus, a control bus, etc. In the drawing, the bus is represented using only a single thick line for convenience of representation, but this does not mean that there is only one bus or only one type of bus.
[0278] Optionally, the memory (1403) may be read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory (CD-ROM), compact disc, laser disc, digital multi-purpose disc, Blu-ray disc, etc.), magnetic disc storage medium, other magnetic storage device, or any other medium that can be used to transmit or store program code expected in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory (1403) may exist independently, be connected to the processor (1401) via a communication bus (1402), or the memory (1403) may be integrated with the processor (1401).
[0279] The communication interface (1404) is configured to communicate with another device or communication network using any transceiver type device. The communication interface (1404) includes a wired communication interface or optionally includes a wireless communication interface. The wired communication interface is, for example, an Ethernet interface. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area network (WLAN) interface, a cellular network communication interface, a combination thereof, etc.
[0280] In some embodiments, memory (1403) is configured to store program code (1405) for executing the solution of the present application. A processor (1401) can execute the program code (1405) stored in memory (1403). The program code includes one or more software modules, and the headset can implement the noise removal method provided in the embodiment of FIG. 2 through the program code (1405) of the processor (1401) and memory (1403).
[0281] All or part of the foregoing embodiments may be implemented using software, hardware, firmware, or any combination thereof. Where software is used to implement the foregoing embodiments, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product comprises one or more computer instructions. When computer instructions are loaded and executed on a computer, a procedure or function according to an embodiment of the present application is created in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. Computer instructions may be stored on a computer-readable storage medium or transferred from a computer-readable storage medium to another computer-readable storage medium. For example, computer instructions may be transferred from a website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, or Digital Subscriber Line (DSL)) or wireless (e.g., infrared, radio, or microwave). A computer-readable storage medium may be any available medium accessible to a computer or a data storage device, such as a server or data center, that incorporates one or more available media. Available media may be magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital multi-purpose disk (DVD)), semiconductor media (e.g., solid-state disk (SSD)), or similar media. It should be noted that the computer-readable storage media mentioned in the embodiments of this application may be non-volatile storage media, i.e., non-transient storage media.
[0282] Embodiments of the present application further provide a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0283] Embodiments of the present application further provide a computer program product. The computer program product stores computer instructions, and when the computer instructions are executed by a processor, the steps of the method described above are implemented.
[0284] In this specification, "at least one" should be understood to mean one or more, and "plural" to mean two or more. In the description of the embodiments of this application, " / " indicates "or" unless otherwise specified. For example, A / B may mean A or B. In this specification, "and / or" describes the association between related objects and indicates that three relationships may exist. For example, A and / or B may mean the following three cases: A alone, A and B both, or B alone. To clearly describe the technical solutions in the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items that provide essentially the same function or purpose. Those skilled in the art will understand that terms such as "first" and "second" do not limit quantity or order of execution, and that terms such as "first" and "second" do not imply a clear difference.
[0285] In the embodiments of this application, information (including, but not limited to, user equipment information, user's personal information, etc.), data (including, but not limited to, data used for analysis, stored data, display data, etc.), and signals are used with the consent of the user or full consent of all parties, and it should be noted that the capture, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0286] The foregoing description is merely an example of the present application and is not intended to limit the application. Any modification, equivalent substitution, or improvement that does not depart from the spirit and principles of the present application shall fall within the scope of protection of the present application.
Claims
Claim 1 A noise removal method applied to a headset, wherein the headset comprises at least one first reference microphone, one error microphone, at least one speaker, and one first feedforward (FF) filter, and the method comprises the steps of: determining a target noise removal parameter based on a reference signal collected by the at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient—the target noise removal parameter includes a filter coefficient of the first FF filter—and performing noise removal through a target speaker among the at least one speaker based on the target noise removal parameter, wherein the initial noise removal coefficient includes an initial filter coefficient of the first FF filter, and the target noise removal parameter includes a k-th frame filter coefficient of the first FF filter, where k is an integer greater than or equal to 1, and the step of determining a target noise removal parameter based on a reference signal collected by the at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient comprises, when k is 1, the initial filter coefficient of the first FF filter and the k-th frame filter of the first FF filter A noise removal method comprising the steps of determining the k-th frame filter coefficient of the first FF filter based on a coefficient, or based on an initial noise removal level and a mapping relationship between the noise removal level and the first FF filter coefficient, or, if k is greater than 1, determining the k-th frame filter coefficient of the first FF filter based on a (k-1)-th frame reference signal collected by at least one first reference microphone, a (k-1)-th frame error signal collected by the error microphone, and a target noise removal level. Claim 2 delete Claim 3 A noise removal method according to claim 1, wherein the step of determining the k-th frame filter coefficient of the first FF filter based on the (k-1)th frame reference signal collected by the at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, and the target noise removal level comprises the step of determining the (k-1)th frame filter coefficient of the target auxiliary path (SP) based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficient of the auxiliary path (SP)—wherein the target SP is a path from the target speaker to the error microphone—and the step of determining the k-th frame filter coefficient of the first FF filter based on the (k-1)th frame reference signal collected by the at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of the target SP. Claim 4 In claim 3, the step of determining the k-th frame filter coefficient of the first FF filter based on the (k-1)th frame reference signal collected by the at least one first reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of the target SP comprises: the step of determining a residual error based on the (k-1)th frame reference signal collected by the at least one first reference microphone and the (k-1)th frame error signal collected by the error microphone; the step of determining the k-th frame frequency response information of the first FF filter based on the (k-1)th frame frequency response information of the first FF filter, the (k-1)th frame filter coefficient of the target SP, and the residual error; and the step of determining the k-th frame filter coefficient of the first FF filter based on the k-th frame frequency response information of the first FF filter. Claim 5 In claim 4, the headset further comprises one feedback (FB) filter, and the step of determining the k-th frame frequency response information of the first FF filter based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the target SP, and the residual error comprises the step of determining the k-th frame frequency response information of the first FF filter based on the (k-1)-th frame frequency response information of the first FF filter, the (k-1)-th frame filter coefficient of the FB filter, the (k-1)-th frame filter coefficient of the target SP, and the residual error, a noise removal method. Claim 6 In claim 4, the step of determining the k-th frame filter coefficient of the first FF filter based on the k-th frame frequency response information of the first FF filter comprises: a step of setting a loss function between the filter coefficient variable of the first FF filter and the k-th frame frequency response information of the first FF filter; a step of determining the value of the filter coefficient variable based on the loss function according to a gradient descent method; and a step of determining the k-th frame filter coefficient of the first FF filter based on the value of the filter coefficient variable. Claim 7 A noise removal method according to claim 6, wherein the step of determining the value of the filter coefficient variable based on the loss function according to the gradient descent method comprises the step of determining the target noise removal amplitude based on the environmental sound level, and the step of determining the value of the filter coefficient variable based on the target noise removal amplitude and the loss function according to the gradient descent method. Claim 8 In claim 1, the headset further comprises one feedback (FB) filter, the initial noise removal coefficient comprises the initial filter coefficient of the FB filter, and the target noise removal parameter further comprises the k-th frame filter coefficient of the FB filter, wherein k is an integer greater than or equal to 1, and the step of determining the target noise removal parameter based on a reference signal collected by at least one first reference microphone, an error signal collected by the error microphone, and the initial noise removal coefficient comprises, when k is 1, determining the initial filter coefficient of the FB filter as the k-th frame filter coefficient of the FB filter, or determining the k-th frame filter coefficient of the FB filter based on an initial noise removal level and a mapping relationship between the noise removal level and the FB filter coefficient, or when k is greater than 1, determining the k-th frame filter coefficient of the FB filter based on a target noise removal level and a mapping relationship between the noise removal level and the FB filter coefficient, or the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the FB filter, and the target noise removal A noise removal method comprising the step of determining the k-th frame filter coefficient of the FB filter based on the level. Claim 9 In claim 8, the step of determining the k-th frame filter coefficient of the FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the FB filter, and the target noise removal level comprises: the step of determining the (k-1)th frame filter coefficient of the target auxiliary path (SP) based on the target noise removal level and the mapping relationship between the noise removal level and the filter coefficient of the auxiliary path (SP)—wherein the target SP is a path from the target speaker to the error microphone—and the step of determining the k-th frame filter coefficient of the FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the FB filter, and the (k-1)th frame filter coefficient of the target SP. Claim 10 A noise removal method according to claim 1, wherein the filter coefficients of the first FF filter include at least one biquad filter coefficient and one gain. Claim 11 A noise removal method according to claim 1, wherein the headset further comprises at least one second FF filter, and the step of performing noise removal through a target speaker among the at least one speaker based on the target noise removal parameter comprises the step of determining the k-th frame filter coefficient of the at least one second FF filter, and the step of performing noise removal through the target speaker based on the target noise removal parameter and the k-th frame filter coefficient of the at least one second FF filter. Claim 12 A noise removal method according to claim 11, wherein the initial noise removal coefficient includes the initial filter coefficient of the at least one second FF filter, and the step of determining the k-th frame filter coefficient of the at least one second FF filter comprises, if k is 1, determining the initial filter coefficient of the at least one second FF filter as the k-th frame filter coefficient of the at least one second FF filter, or determining the k-th frame filter coefficient of the at least one second FF filter based on an initial noise removal level and a mapping relationship between the noise removal level and the second FF filter coefficient, or if k is greater than 1, determining the k-th frame filter coefficient of the at least one second FF filter based on a target noise removal level and a mapping relationship between the noise removal level and the second FF filter coefficient. Claim 13 A noise removal method according to claim 1, wherein the headset further comprises a downlink compensation filter, the initial noise removal coefficient comprises an initial filter coefficient of the downlink compensation filter, and the step of determining a target noise removal parameter based on a reference signal collected by at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient further comprises, when k is 1, determining the initial filter coefficient of the downlink compensation filter as the k-th frame filter coefficient of the downlink compensation filter, or determining the k-th frame filter coefficient of the downlink compensation filter based on an initial noise removal level and a mapping relationship between the noise removal level and the downlink compensation filter coefficient, or when k is greater than 1, determining the k-th frame filter coefficient of the downlink compensation filter based on a target noise removal level and a mapping relationship between the noise removal level and the downlink compensation filter coefficient. Claim 14 As a headset, the headset comprises at least one first reference microphone, one error microphone, at least one speaker, one first feedforward (FF) filter, and one noise removal processor, wherein the noise removal processor is configured to implement steps of a method comprising: determining a target noise removal parameter based on a reference signal collected by the at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient—the target noise removal parameter includes a filter coefficient of the first FF filter—and performing noise removal through a target speaker among the at least one speaker based on the target noise removal parameter, wherein the initial noise removal coefficient includes an initial filter coefficient of the first FF filter, and the target noise removal parameter includes a k-th frame filter coefficient of the first FF filter, where k is an integer greater than or equal to 1, and the step of determining a target noise removal parameter based on a reference signal collected by the at least one first reference microphone, an error signal collected by the error microphone, and an initial noise removal coefficient comprises, when k is 1, the initial filter coefficient of the first FF filter A headset comprising the steps of determining the k-th frame filter coefficient of the first FF filter, or determining the k-th frame filter coefficient of the first FF filter based on an initial noise removal level and a mapping relationship between the noise removal level and the first FF filter coefficient, or, if k is greater than 1, determining the k-th frame filter coefficient of the first FF filter based on a (k-1)-th frame reference signal collected by at least one first reference microphone, a (k-1)-th frame error signal collected by the error microphone, and a target noise removal level. Claim 15 In paragraph 14, the above-mentioned at least one speaker comprises a first speaker and a second speaker on which digital frequency division is performed, the target speaker is the first speaker, and the second speaker does not participate in noise removal, a headset. Claim 16 A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to claim 1 are implemented. Claim 17 A computer program stored on a computer-readable storage medium, wherein the computer program includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method according to claim 1 are implemented. Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete Claim 23 delete
Citation Information
Patent Citations
Filter design method and device and in-ear active noise reduction earphone
CN113132848A
Active noise cancellation featuring secondary path estimation
US20160300563A1
Hybrid active noise cancellation filter adaptation
US20210375254A1