Noise cancellation method, headset, device, storage medium, and computer program product

The noise cancellation method in headsets addresses inconsistent noise leakage by using adaptive noise cancellation parameters and full-band anti-phase noises, enhancing sound clarity and effectiveness across different user fits and environments.

JP2025539869APending Publication Date: 2025-12-09HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025530756
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-28
Filing Date
2023-06-28
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Active noise cancellation in headsets is challenged by variable and irregular environmental noise, and the varying fit between headsets and human ears, leading to inconsistent noise leakage and cancellation effectiveness.

Method used

A noise cancellation method utilizing multiple reference and error microphones, feedforward and feedback filters, and adaptive noise cancellation parameters to generate full-band anti-phase noises for each speaker, dynamically adjusting to the user's ear fit and noise environment.

Benefits of technology

Enhances noise cancellation across various environments and user fits by optimizing noise cancellation parameters on a frame-by-frame basis, utilizing full-band anti-phase noises to improve sound clarity and reduce noise leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539869000001_ABST
    Figure 2025539869000001_ABST
Patent Text Reader

Abstract

This application discloses a noise cancellation method, a headset, an apparatus, a storage medium, and a computer program product, which relate to the field of audio processing technology. The headset includes at least one reference microphone, one error microphone, and a plurality of first speakers. The method includes determining a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; generating a plurality of groups of target anti-phase noises corresponding one-to-one to the plurality of first speakers based on the plurality of groups of target noise cancellation parameters, wherein the frequency band of each target anti-phase noise in the plurality of groups of target anti-phase noises covers the sound-generating frequency band of the plurality of first speakers; and performing noise cancellation through the plurality of first speakers by using the plurality of groups of target anti-phase noises. This solution can improve the noise cancellation effect of the headset by using full-band anti-phase noise in the plurality of noise cancellation channels.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to Chinese Patent Application No. 202211506453.2, entitled "NOISE CANCELLATION METHOD, HEADSET, APPARATUS, STORAGE MEDIUM, AND COMPUTER PROGRAM PRODUCT," filed on November 28, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of audio processing technology, and in particular to a noise cancellation method, a headset, a device, a storage medium, and a computer program product. [Background technology]

[0003] When a user wears a headset to listen to audio signals such as music or voice, the existence of environmental noise will affect the clarity of the audio signals heard by the user, and when the environmental noise is severe, the user cannot even clearly hear the audio signals in the headset. Therefore, it is necessary to implement active noise cancellation in the headset to eliminate as much as possible the environmental noise heard by the headset wearer.

[0004] Active noise cancellation in headsets faces many challenges. Environmental noise is variable and irregular. In addition, the degree to which environmental noise leaks into the ear canal is related to the degree of fit between the headset and the human ear. However, the size and shape of the ear canal vary from person to person, and different people wearing the same headset will have different levels of fit between the headset and the human ear, resulting in different levels of noise leakage. Even when the same user wears the same headset multiple times, the fit between the headset and the human ear may also vary. Therefore, current research focuses on how to improve the effectiveness of active noise cancellation in headsets to minimize the impact of environmental noise on the headset wearer. Summary of the Invention

[0005] The present application provides a noise cancellation method, a headset, a device, a storage medium, and a computer program product for improving the effect of active noise cancellation of a headset. The technical solutions are as follows:

[0006] According to a first aspect, there is provided a noise cancellation method applied to a headset. The headset includes at least one reference microphone, one error microphone, and a plurality of first speakers. The method includes the steps of: determining a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; generating a plurality of groups of target anti-phase noises corresponding one-to-one to the plurality of first speakers based on the plurality of groups of target noise cancellation parameters, wherein a frequency band of each target anti-phase noise in the plurality of groups of target anti-phase noises covers a sound-generating frequency band of the plurality of first speakers; and performing noise cancellation via the plurality of first speakers by using the plurality of groups of target anti-phase noises.

[0007] The multiple groups of target anti-phase noises correspond one-to-one to the multiple first speakers, so that the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound-generating frequency band of the multiple first speakers. In other words, each target anti-phase noise is a full-band anti-phase noise. Therefore, regardless of whether the first speaker is a high-band speaker, a low-band speaker, or a full-band speaker, when the multiple groups of target anti-phase noises are used to perform noise cancellation, the noise cancellation capability of each first speaker can be fully utilized. In other words, in a headset architecture including multiple noise cancellation channels and multiple speakers, this solution can improve the noise cancellation effect of the headset by using the full-band anti-phase noise of the multiple noise cancellation channels.

[0008] According to the noise cancellation method provided in the present application, the plurality of groups of target noise cancellation parameters can be determined on a frame-by-frame basis. In other words, the plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers are determined for each frame. Of course, the target noise cancellation parameters may alternatively be determined on another time basis. For example, the plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers are determined every two frames. In the following, a frame is used as a unit for description.

[0009] The headset has multiple feedforward amplifiers that correspond one-to-one with multiple primary speakers. (F In this case, the plurality of groups of target noise cancellation parameters include a k-th frame filter coefficient of a plurality of FF filters, where k is an integer equal to or greater than 1. In some cases, the headset includes a plurality of feedback filters that correspond one-to-one to the plurality of first speakers. (FB) further includes a filter. In other words, the multiple FB filters correspond one-to-one to the multiple FF filters. In this case, the multiple groups of target noise cancellation parameters further include the k-th frame filter coefficients of the multiple FB filters. In addition, if the headset further includes a downlink compensation filter, the multiple groups of target noise cancellation parameters further include the k-th frame filter coefficients of the downlink compensation filter. In addition, if k is greater than 1, a target noise cancellation level can also be determined. Therefore, the following four parts will be described separately.

[0010] (1) The k-th frame filter coefficients of a plurality of FF filters are determined.

[0011] When k is equal to 1, the initial filter coefficients of the multiple FF filters are determined as the k-th frame filter coefficients of the multiple FF filters, i.e., the first frame filter coefficients of the multiple FF filters are the initial filter coefficients of the corresponding FF filters, or the k-th frame filter coefficients of the multiple FF filters are determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the FF filter coefficients. When k is greater than 1, the k-th frame filter coefficients of the multiple FF filters are determined based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the target noise canceling level. In other words, the k-th frame filter coefficients of the multiple FF filters are determined according to an adaptive method. The determination process is an adaptive process and may also be called an iterative process.

[0012] It should be noted that the initial filter coefficients of the multiple FF filters may be the same or different, and the initial filter coefficients may be zero or non-zero. This is not a limitation in the embodiments of the present application. The initial noise cancellation level may be a preset level at which noise cancellation can be successfully performed by using the corresponding noise cancellation coefficients without introducing stability issues. Of course, the initial noise cancellation level may alternatively be a level determined based on a prompt tone, such as "Noise cancellation on" or "Dingdong," transmitted by the user terminal when noise cancellation starts. The noise cancellation coefficients corresponding to this level may better adapt to the current human ear and wearing posture, and performing adaptive iterations based on the noise cancellation coefficients corresponding to this level may more quickly reach a convergence state. This is also not a limitation in the embodiments of the present application.

[0013] The implementation process of determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the target noise canceling level includes determining the target noise canceling level, the noise canceling level, and the secondary path (S P), where the plurality of SPs are paths from a plurality of first speakers to an error microphone; and determining k-th frame filter coefficients of the plurality of FF filters based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the (k-1)-th frame filter coefficients of the plurality of SPs.

[0014] The k-th frame filter coefficients of the multiple FF filters can be determined using a multi-channel linkage method. In addition, when the headset includes multiple FF filters, the headset may further include multiple FB filters that correspond one-to-one to the multiple first speakers, or may not include multiple FB filters. In different cases, the methods for determining the k-th frame filter coefficients of the multiple FF filters are different. The methods will be described separately below.

[0015] Since the process of determining the k-th frame filter coefficient of an FF filter based on the (k-1)th frame reference signal collected by at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of a plurality of SPs is the same, one of the processes will be used as an example below for description. In other words, one of the plurality of FF filters is used as a target FF filter, and the k-th frame filter coefficient of the target FF filter is determined as follows: For the process of determining the k-th frame filter coefficient of another FF filter of the plurality of FF filters, please refer to the process of determining the k-th frame filter coefficient of the target FF filter.

[0016] In the first case, the headset does not include multiple FB filters. If the target FF filter is the first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of the target SP, where the target reference microphone is the reference microphone corresponding to the target FF filter, and the target SP is the path from the first speaker to the error microphone corresponding to the target FF filter. If the target FF filter is a non-first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the multiple SPs, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0017] If the target FF filter is the first FF filter, a residual is determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the (k-1)th frame filter coefficient of the target SP, and the residual. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0018] One of the multiple FF filters corresponds to one reference microphone. In other words, the target reference microphone includes one reference microphone. In this case, the residual is determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone.

[0019] One of the multiple FF filters corresponds to at least two reference microphones. In other words, the target reference microphone includes at least two reference microphones. In this case, audio mixing is performed on the (k-1)th frame reference signals collected by the at least two reference microphones included in the target reference microphone to obtain a (k-1)th frame mixed reference signal. The residual is determined based on the (k-1)th frame mixed reference signal and the (k-1)th frame error signal collected by the error microphone. In this way, the signal-to-noise ratio of the reference signal can be improved.

[0020] If the target FF filter is a non-first FF filter, a residual is determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple SPs, and the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter. The kth frame filter coefficient of the target FF filter is determined based on the kth frame frequency response information of the target FF filter.

[0021] When determining the kth frame frequency response information of the target FF filter, the kth frame frequency response information of the target FF filter can be determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficient of the target SP, the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter, and the (k-1)th frame filter coefficient of the SP corresponding to each FF filter before the target FF filter.

[0022] An implementation process for determining the kth frame filter coefficients of the target FF filter based on the kth frame frequency response information of the target FF filter includes establishing a loss function between filter coefficient variables of the target FF filter and the kth frame frequency response information of the target FF filter. Based on the loss function, values ​​of the filter coefficient variables are determined using a gradient descent method, and the kth frame filter coefficients of the target FF filter are determined based on the values ​​of the filter coefficient variables. In other words, a loss function is established between the filter coefficient variables of the target FF filter and the kth frame frequency response information of the target FF filter. Optimal values ​​of the variables are determined using a gradient descent method, and the kth frame filter coefficients of the target FF filter are determined based on the optimal values ​​of the variables.

[0023] The filter coefficients of the target FF filter for each frame are determined according to a gradient descent method. Once the filter coefficients of the target FF filter for each frame are determined, a value of a loss function is determined. When the value of the loss function reaches a minimum threshold, it is determined that the filter coefficients of the target FF filter have reached a convergence and stability condition. For example, for the kth frame filter coefficient of the target FF filter, when the value of the loss function between the filter coefficient variable and the kth frame frequency response information of the target FF filter reaches a minimum threshold, it is determined that the kth frame filter coefficient of the target FF filter has reached a convergence and stability condition. When the value of the loss function does not reach the minimum threshold, it is determined that the kth frame filter coefficient of the target FF filter has not reached a convergence and stability condition. The minimum threshold may be preset and adjusted based on different requirements in different cases.

[0024] Optionally, the filter coefficients of each FF filter include at least one biquadratic filter coefficient and one gain. Variables corresponding to the biquadratic filter coefficient include a filter type, a cutoff frequency, and a quality factor. Of course, in practical applications, the filter coefficients of each FF filter may further include more or fewer other parameters. This is not a limitation in the present application.

[0025] In some cases, background noise, i.e., noise floor issues, may occur in quiet environments. For example, semi-open headsets are more likely to have background noise issues in quiet environments than in-ear headsets. Additionally, strong noise cancellation is unnecessary in quiet environments, and some people find strong noise cancellation uncomfortable. Furthermore, a stronger noise cancellation intensity indicates a stronger sense of negative pressure. Therefore, when the value of the filter coefficient variable is determined using the gradient descent method, the target noise cancellation amplitude can be dynamically adjusted based on the ambient sound volume. The k-th frame filter coefficient of the target FF filter is determined based on this target noise cancellation amplitude, thereby improving the subjective experience of adaptive noise cancellation. In other words, the target noise cancellation amplitude is determined based on the ambient sound volume of the (k-1)th frame and the ambient sound volume in t frames prior to the (k-1)th frame, where t is greater than or equal to 1 and less than k-1. The value of the filter coefficient variable is determined based on the target noise cancellation amplitude and a loss function using the gradient descent method, and the k-th frame filter coefficient of the target FF filter is determined based on the value of the filter coefficient variable.

[0026] In a second case, the headset further includes a plurality of FB filters. When the target FF filter is the first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, and the (k-1)th frame filter coefficients of the plurality of FB filters. When the target FF filter is a non-first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, the (k-1)th frame filter coefficients of the plurality of FB filters, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0027] When the target FF filter is the first FF filter, a residual may be determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple FB filters, and the (k-1)th frame filter coefficients of the multiple SPs. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0028] If the target FF filter is a non-first FF filter, the residual may be determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple SPs, the (k-1)th frame filter coefficients of the multiple FB filters, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0029] In the above-described process for determining the k-th frame frequency response information of the target FF filter, regardless of whether the headset includes a target FB filter, the k-th frame frequency response information of the target FF filter is determined based on the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the target SP is determined based on the target noise canceling level by querying the mapping relationship between the noise canceling level and the filter coefficient of the SP. Specifically, the (k-1)-th frame filter coefficient of the target SP is an estimated value, and the k-th frame frequency response information of the target FF filter is determined based on this estimated value, thereby eliminating dependence on the actual value of the target SP and enabling adaptation of the filter coefficient of the FF filter even when there is no downlink signal.

[0030] (2) The k-th frame filter coefficients of a plurality of FB filters are determined.

[0031] When k is equal to 1, the initial filter coefficients of the plurality of FB filters are determined as the k-th frame filter coefficients of the plurality of FB filters, that is, the first frame filter coefficients of the plurality of FB filters are the initial filter coefficients of the corresponding FB filters, or the k-th frame filter coefficients of the plurality of FB filters are determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficients. When k is greater than 1, the k-th frame filter coefficients of the plurality of FB filters can be determined based on the target noise canceling level.

[0032] It should be noted that the initial filter coefficients of the multiple FB filters may be the same or different, and the initial filter coefficients may be 0 or not 0. This is not limited in the embodiments of the present application.

[0033] Since the process of determining the k-th frame filter coefficient of the FB filter based on the target noise canceling level is the same, the following description will use one of the processes as an example. In other words, one of the multiple FB filters is used as the target FB filter, and the k-th frame filter coefficient of the target FB filter is determined in the following two ways: For the process of determining the k-th frame filter coefficient of another FB filter of the multiple FB filters, please refer to the process of determining the k-th frame filter coefficient of the target FB filter. In other words, when k is greater than 1, the k-th frame filter coefficient of the target FB filter can be determined in the following two ways:

[0034] In the first method, the k-th frame filter coefficient of the target FB filter is determined based on the target noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficient.

[0035] Since the mapping relationship between the noise canceling level and the FB filter coefficient is stored in advance, determining the k-th frame filter coefficient of the target FB filter in the first method is stable, simple in operation, and highly efficient.

[0036] In the second method, when the target FB filter is a first type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on a target noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficient. When the target FB filter is a second type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on a (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level.

[0037] The first frame filter coefficient of the target FB filter can be determined based on the initial noise canceling level by querying the mapping relationship between the noise canceling level and the FB filter coefficient. Therefore, when k is 1 or greater, this is equivalent to determining the kth frame filter coefficient of the target FB filter in three ways. Specifically, (1) the kth frame filter coefficient of the target FB filter can be determined by querying the mapping relationship between the noise canceling level and the FB filter coefficient. (2) If the target FB filter is a first type FB filter, the kth frame filter coefficient of the target FB filter can be determined by querying the mapping relationship between the noise canceling level and the FB filter coefficient. If the target FB filter is a second type FB filter, the kth frame filter coefficient of the target FB filter can be determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level. (3) When the target FB filter is a first type FB filter, or when the target FB filter is a second type FB filter and k is equal to 1, the k-th frame filter coefficient of the target FB filter is determined by consulting the mapping relationship between the noise canceling level and the FB filter coefficient. When the target FB filter is a second type FB filter and k is greater than 1, the k-th frame filter coefficient of the target FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level.

[0038] An implementation process for determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level includes: a step of determining the (k-1)th frame filter coefficient of the target SP based on the target noise canceling level and a mapping relationship between the noise canceling level and the filter coefficient of the SP, where the target SP is a path from the first speaker corresponding to the target FB filter to the error microphone; and a step of determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the (k-1)th frame filter coefficient of the target SP.

[0039] The sound-producing frequency band of the first speaker corresponding to the first type of FB filter is higher than the sound-producing frequency band of the first speaker corresponding to the second type of FB filter. In other words, the first speaker corresponding to the first type of FB filter is a high-bandwidth speaker, and the first speaker corresponding to the second type of FB filter is a low-bandwidth speaker. Of course, the first type of FB filter and the second type of FB filter may be distinguished in a manner other than based on the sound-producing frequency band. This is also not a limitation of the present application.

[0040] In the second and third methods described above, the method of querying the mapping relationship between the noise cancellation level and the FB filter coefficient is combined with the adaptive method, which can improve the noise cancellation effect without high complexity and with controllable stability.

[0041] It should be noted that the k-th frame filter coefficient of the target FB filter may be determined by the three methods described above, and alternatively, the k-th frame filter coefficient of the target FB filter may be determined by another method. For example, regardless of whether the target FB filter is a first type FB filter or a second type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level. This is not limited to the embodiments of the present application.

[0042] (3) Determine the k-th frame filter coefficient of the downlink compensation filter.

[0043] When k is equal to 1, the initial filter coefficient of the downlink compensation filter is determined as the k-th frame filter coefficient of the downlink compensation filter, or the k-th frame filter coefficient of the downlink compensation filter is determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the downlink compensation filter coefficient.When k is greater than 1, the k-th frame filter coefficient of the downlink compensation filter is determined based on the target noise canceling level and the mapping relationship between the noise canceling level and the downlink compensation filter coefficient.

[0044] The mapping relationship between the noise canceling level and the downlink compensation filter coefficient includes a plurality of noise canceling levels, and there is a mapping relationship between each noise canceling level and the filter coefficient of the downlink compensation filter, and the mapping relationship between different noise canceling levels and the filter coefficient of the downlink compensation filter may be different. Therefore, after the target noise canceling level is determined, the corresponding downlink compensation filter coefficient can be obtained from the mapping relationship between the noise canceling level and the downlink compensation filter coefficient based on the target noise canceling level, and the obtained downlink compensation filter coefficient is used as the k-th frame filter coefficient of the downlink compensation filter.

[0045] (4) Determine the target noise canceling level.

[0046] A (k-1)th frame noise canceling level is determined, and noise canceling levels in m frames before the (k-1)th frame are obtained, where m is greater than or equal to 1 and less than k-1. A target noise canceling level is determined based on the (k-1)th frame noise canceling level and the noise canceling levels in m frames.

[0047] In the (k-1)th frame, there may be a valid downlink signal, there may not be a valid downlink signal, the environment may be quiet, the environment may not be quiet, or of course, there may be an abnormal signal. In different cases, the manner of determining the (k-1)th frame noise canceling level is different, which will be described separately below.

[0048] In the first case, in the (k-1)th frame, there is no valid downlink signal and the environment is not quiet. In this case, the (k-1)th frame noise canceling level is determined based on the reference filter coefficients of multiple FF filters and the mapping relationship between the noise canceling level and the frequency response information of the FF filters. When k is equal to 2, the reference filter coefficient is the initial filter coefficient of the corresponding FF filter; or when k is greater than 2, the reference filter coefficient is the filter coefficient of the corresponding FF filter that last satisfies the convergence and stability condition before the kth frame, or is the (k-1)th frame filter coefficient of the corresponding FF filter.

[0049] Reference frequency response information of the plurality of FF filters is determined based on reference filter coefficients of the plurality of FF filters. Noise canceling levels matching the reference frequency response information of the plurality of FF filters are determined based on a mapping relationship between the noise canceling levels and the frequency response information of the FF filters to obtain a plurality of reference noise canceling levels. The (k-1)th frame noise canceling level is determined based on the plurality of reference noise canceling levels.

[0050] There are several ways to determine the noise canceling level of the (k-1)th frame based on the multiple reference noise canceling levels. For example, the noise canceling level of the (k-1)th frame may be determined based on the average value of the multiple reference noise canceling levels. Alternatively, the noise canceling level of the (k-1)th frame may be determined based on the reference noise canceling level having the maximum amount among the multiple reference noise canceling levels.

[0051] When the (k-1)th frame noise canceling level is determined based on an average value of a plurality of reference noise canceling levels, the average value of the plurality of reference noise canceling levels may be determined as the (k-1)th frame noise canceling level as is, or the average value of the plurality of reference noise canceling levels may be adjusted to obtain the (k-1)th frame noise canceling level. Similarly, when the (k-1)th frame noise canceling level is determined based on the reference noise canceling level having the maximum amount among the plurality of reference noise canceling levels, the reference noise canceling level having the maximum amount among the plurality of reference noise canceling levels may be determined as the (k-1)th frame noise canceling level as is, or alternatively, the reference noise canceling level having the maximum amount among the plurality of reference noise canceling levels may be adjusted to obtain the (k-1)th frame noise canceling level.

[0052] In the second case, a valid downlink signal exists in the (k-1)th frame, and the (k-1)th frame noise canceling level is determined based on the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by the error microphone.

[0053] In consideration of the above description, when the headset is in a downlink enabled state and is not in a downlink intermittent period, it is determined that a valid downlink signal exists in the (k-1)th frame. In this case, based on the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by the error microphone, the valid downlink signal can be extracted from the (k-1)th frame error signal collected by the error microphone, and a (k-1)th frame noise canceling level can be determined based on the extracted valid downlink signal.

[0054] In the third case, there is no valid downlink signal in the (k-1)th frame, and the environment is quiet or an abnormal noise signal exists in the (k-1)th frame. In this case, the (k-3)th frame noise canceling level is determined as the (k-1)th frame noise canceling level. In other words, the noise canceling level remains unchanged.

[0055] In the (k-1)th frame, when there is no valid downlink signal and the environment is quiet, the noise basically remains unchanged. In this case, the noise canceling level remains unchanged. If there is an abnormal noise signal in the (k-1)th frame, the noise canceling level remains unchanged to perform robustness control and avoid divergence of the noise canceling level.

[0056] After the (k-1)th frame noise canceling level is determined in the three cases mentioned above, the target noise canceling level can be determined by integrating the (k-1)th frame noise canceling level with the noise canceling levels in the m frames prior to the (k-1)th frame.

[0057] The noise canceling level in the m frames may be the noise canceling level in any m frames before the (k-1)th frame, or the noise canceling level in the m frames before and closest to the (k-1)th frame. In addition, there are several implementations for determining the target noise canceling level based on the (k-1)th frame noise canceling level and the noise canceling levels in the m frames before the (k-1)th frame. For example, the noise cancellation effect is evaluated according to a related algorithm to determine the noise cancellation probability corresponding to the (k-1)th frame noise canceling level and the noise cancellation probability corresponding to the noise canceling levels in the m frames, and the noise canceling level with the maximum noise cancellation probability is determined as the target noise canceling level. Alternatively, the arithmetic average or weighted average of the (k-1)th frame noise canceling level and the noise canceling levels in the m frames is determined to obtain the target noise canceling level. Alternatively, the noise canceling level that appears most frequently among the (k-1)th frame noise canceling level and the noise canceling levels in the m frames may be determined as the target noise canceling level.

[0058] In consideration of the above description, the multiple groups of target noise cancellation parameters may be referred to as noise cancellation parameters of multiple noise cancellation channels. In this way, the multiple groups of generated target anti-phase noise may also be referred to as anti-phase noise of multiple noise cancellation channels. Because the process of generating anti-phase noise for the noise cancellation channels is the same, the following uses one of the noise cancellation channels as an example for explanation.

[0059] One of the multiple noise cancellation channels is used as a target noise cancellation channel, and the target noise cancellation channel includes a target FF filter and a target first speaker, and a reference microphone corresponding to the target FF filter is called a target reference microphone. In this case, the target anti-phase noise includes feedforward anti-phase noise. In other words, the kth frame reference signal collected by the target reference microphone is processed based on the kth frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise.

[0060] In consideration of the foregoing description, the target reference microphone may include one reference microphone or at least two reference microphones. When the target reference microphone includes one reference microphone, the k-th frame reference signal collected by the target reference microphone may be directly processed based on the k-th frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise. When the target reference microphone includes at least two reference microphones, audio mixing is performed on the k-th frame reference signals collected by the at least two reference microphones to obtain a mixed reference signal of the k-th frame, and then the mixed reference signal of the k-th frame is processed based on the k-th frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise.

[0061] If the headset further includes an FB filter, the target noise cancellation channel further includes a target FB filter. In this case, the target anti-phase noise further includes feedback anti-phase noise. In other words, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal based on the k-th frame filter coefficient of the downlink compensation filter. Then, negation is performed on the k-th frame downlink signal obtained through downlink compensation, and audio mixing is performed on the negated k-th frame downlink signal and the k-th frame error signal collected by the error microphone to obtain the k-th frame noise signal collected by the error microphone. The k-th frame noise signal collected by the error microphone is processed based on the k-th frame filter coefficient of the target FB filter to obtain the feedback anti-phase noise.

[0062] Downlink compensation can be used to remove all downlink signals in the error signal collected by the error microphone, so that noise cancellation is only performed on the residual noise signal through the FB filter to avoid damaging the sound quality of the downlink signal. In addition, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal, so that the downlink signals of all speakers in the error microphone can be removed to avoid damaging the sound quality of the full-band downlink signal.

[0063] Considering the above description, when multiple groups of target noise cancellation parameters are determined for each frame, a frame may include one sample point or multiple sample points, so when the target anti-phase noise is generated, a group of target anti-phase noise may be generated at each sample point, or a group of target anti-phase noise may be generated in one frame.

[0064] When the multiple groups of target noise cancellation parameters are determined, no frequency division is performed on the downlink signal, i.e., the multiple groups of target noise cancellation parameters are determined based on the full-band downlink signal. In this way, after the multiple groups of target anti-phase noises corresponding one-to-one to the multiple first speakers are generated based on the multiple groups of target noise cancellation parameters, the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound generation frequency band of the multiple first speakers, i.e., the frequency band of each target anti-phase noise is the full frequency band.

[0065] After multiple groups of target anti-phase noise are generated, the multiple groups of target anti-phase noise are respectively mixed with the kth frame downlink signal to be played through multiple first speakers, and then the mixed signal is played through the corresponding first speakers to achieve noise cancellation.

[0066] Some of the multiple first speakers may be high-band speakers and others may be low-band speakers. Alternatively, some of the multiple first speakers may be full-band speakers and others may not be full-band speakers. In other words, the sound generation frequency bands of the multiple first speakers may be different. Alternatively, all of the multiple first speakers are full-band speakers. Alternatively, all of the multiple first speakers are not full-band speakers. If all of the multiple first speakers are full-band speakers, the k-th frame downlink signals to be played through the multiple first speakers are all k-th frame downlink signals transmitted by the user terminal. If all of the multiple first speakers are not full-band speakers, it is necessary to perform frequency division on the k-th frame downlink signals transmitted by the user terminal based on the sound generation frequency bands of each first speaker to obtain the k-th frame downlink signals to be played through each first speaker.

[0067] Two of the first speakers of the plurality of first speakers may include two first speakers formed by one dual-diaphragm (also called dual-dynamic) loudspeaker. Alternatively, the plurality of first speakers may include a plurality of split speakers / loudspeakers.

[0068] Optionally, the headset may further include at least one second speaker, where the at least one second speaker is not involved in noise cancellation. In this case, the second speaker may be involved in downlink compensation (i.e., downlink compensation is performed on a downlink signal transmitted by the user terminal, and the downlink signal is a full-band audio signal including an audio signal in the sound-generating frequency band of the second speaker). In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker, a mid-band speaker, or a low-band speaker. Optionally, the second speaker may not be involved in downlink compensation. In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker is a high-band speaker.

[0069] The aforementioned process of determining the multiple groups of target noise cancellation parameters according to the adaptive method requires a certain time. When a frame includes multiple sample points and the duration of the frame is long, the duration of determining the multiple groups of target noise cancellation parameters is shorter than the duration of the frame. Therefore, during a part of the period of the kth frame, calculations may be performed based on the relevant data of the (k-1)th frame to obtain the multiple groups of target noise cancellation parameters for the kth frame, and during another part of the period of the kth frame, active noise cancellation may be performed based on the multiple groups of target noise cancellation parameters for the kth frame. However, when a frame includes one sample point or when a frame includes multiple sample points and the duration of the frame is short, the duration of determining the multiple groups of target noise cancellation parameters may be equal to the duration of the frame. In this case, it may be necessary to perform calculations for the entire period of the kth frame based on the relevant data of the (k-1)th frame to obtain the multiple groups of target noise cancellation parameters. In this case, the multiple groups of target noise cancellation parameters can be determined as the multiple groups of target noise cancellation parameters in the (k+1)th frame, and then active noise cancellation is performed during the (k+1)th frame based on the multiple groups of target noise cancellation parameters in the (k+1)th frame. The above content will be described using the former case as an example.

[0070] According to a second aspect, there is provided a headset, the headset including at least one reference microphone, one error microphone, a plurality of first speakers, and a noise cancellation processor, the noise cancellation processor configured to perform the steps of the method according to the first aspect.

[0071] Optionally, the plurality of first speakers includes two first speakers formed by one dual diaphragm loudspeaker, or the plurality of first speakers includes a plurality of speakers with separate loudspeakers.

[0072] Optionally, the headset further comprises at least one second speaker, wherein the at least one second speaker does not participate in noise cancellation.

[0073] According to a third aspect, there is provided a noise cancellation device having a function of performing the operations of the noise cancellation method of the first aspect. The noise cancellation device includes one or more modules, and the one or more modules are configured to perform the noise cancellation method provided in the first aspect.

[0074] According to a fourth aspect, there is provided a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the noise cancellation method described in the first aspect.

[0075] According to a fifth aspect, there is provided a computer program product comprising instructions which, when executed on a computer, enable the computer to perform the noise cancellation method described in the first aspect.

[0076] The technical effects obtained in the second to fifth aspects are the same as those obtained by the corresponding technical means in the first aspect, and the details will not be described again in this specification. [Brief explanation of the drawings]

[0077] [Figure 1] 1 is a diagram of a system architecture associated with a noise cancellation method according to an embodiment of the present application; [Figure 2] 1 is a flowchart of a noise cancellation method according to an embodiment of the present application. [Figure 3] 1 is a flowchart for determining a target noise cancellation amplitude according to an embodiment of the present application. [Figure 4] FIG. 10 illustrates frequency response curves of the FF filter at 16 noise canceling levels according to an embodiment of the present application. [Figure 5] 10 is a flowchart for determining a (k-1)th frame noise canceling level according to an embodiment of the present application. [Figure 6] 1 is a flowchart for determining multiple groups of target noise cancellation parameters according to an embodiment of the present application. [Figure 7] 1 is a diagram of a headset structure according to an embodiment of the present application. [Figure 8] FIG. 10 is a diagram of another headset configuration according to an embodiment of the present application. [Figure 9] FIG. 10 is a diagram of another headset configuration according to an embodiment of the present application. [Figure 10] FIG. 10 is a diagram of another headset configuration according to an embodiment of the present application. [Figure 11] FIG. 10 is a diagram of another headset configuration according to an embodiment of the present application. [Figure 12] 1 is a diagram of a structure of a noise cancellation device according to an embodiment of the present application; [Figure 13] FIG. 10 is a diagram of another headset configuration according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0078] To make the objectives, technical solutions and advantages of the present application clearer, the following further describes in detail the implementation of the present application with reference to the accompanying drawings.

[0079] Active noise-canceling headsets have become popular in recent years. Conventional noise-canceling headsets are generally in-ear or head-mounted. This is because these two types of headsets provide a good seal between the headset and the ear canal, ensuring consistent acoustic leakage even when worn by different people. This allows active noise cancellation to be technically realized with better results. Therefore, noise-canceling modes with fixed coefficients are generally used. However, these two types of headsets also have some drawbacks. For example, the seal between the headset and the ear canal is too good, which affects people's subjective comfort, typically characterized by a foreign body sensation and a feeling of blockage during walking and other conditions. Wearing them for long periods of time is difficult.

[0080] Semi-open headsets are widely accepted by users due to their comfortable fit. However, in semi-open headsets, the headset and the human ear are not sufficiently sealed, making it more likely that people will perceive environmental noise. Achieving active noise cancellation in semi-open headsets is more difficult because different people wear headsets, and even the same person wears a headset at different times, have significantly different wearing postures. Technically, the response function and degree of acoustic leakage between the headset and the ear canal vary significantly. Therefore, how to achieve adaptive noise cancellation and optimal matching between the headset and the ear canal to address the issue of ear canal response differences are urgent requirements for semi-open headsets. In addition, whether in-ear or head-mounted, ear canal responses are not absolutely identical and still vary greatly. Currently, the industry is also exploring the feasibility of adaptive noise cancellation in headsets.

[0081] As mentioned above, headsets come in multiple forms, such as in-ear, head-mounted, semi-open, and open. The audio performance of the speakers (i.e., loudspeakers) in a headset, especially in low frequencies, is closely related to the specific form of the headset. Closed headsets, such as in-ear or head-mounted headsets, generally ensure high, mid, and low frequency audio performance. Semi-open or open headsets suffer from significant acoustic leakage and a significant reduction in low-frequency response. This affects the performance of low-frequency sound quality and severely impacts the active noise cancellation effect (which is insufficient to generate anti-phase noise with sufficient energy).

[0082] In consideration of the aforementioned problems, an embodiment of the present application provides an adaptive active noise cancellation for a headset. (A The present invention provides a noise cancellation method for achieving noise cancellation (NC). Please refer to FIG. 1. FIG. 1 is a diagram of a system architecture related to the noise cancellation method according to an embodiment of the present application. The system may be referred to as a headset noise cancellation system. The system includes a headset 101 and a user terminal 102. The headset 101 and the user terminal 102 are connected and communicate with each other in a wired or wireless manner. For example, the headset 101 communicates with the user terminal 102 via Bluetooth (registered trademark) or another wireless network.

[0083] Audio and control signals can be transmitted between the headset 101 and the user terminal 102. For example, the user terminal 102 transmits audio signals, such as music or voice, to the headset 101 for playback. In another example, the user terminal 102 transmits control signals to the headset 101 to control, for example, whether an active noise cancellation function of the headset 101 is enabled.

[0084] The user terminal 102 may be an electronic device such as a mobile phone or a computer (e.g., a notebook computer, a desktop computer, a handheld tablet computer, or an in-vehicle tablet computer). Alternatively, the user terminal 102 may be another electronic device, such as a smart speaker or an in-vehicle speaker. The type, structure, etc. of the user terminal 102 are not limited in the embodiments of the present application.

[0085] Optionally, the headset 101 provided in the embodiment of the present application may be wired or wireless. In addition, from the viewpoint of wearing style, the headset 101 provided in the embodiment of the present application may be neck-worn, ear-worn / ear-clip, true wireless stereo, (T In terms of appearance, the headset 101 provided in the embodiment of the present application may be an in-ear type, a semi-open type, an open type, a head-mounted type, etc. The communication method, wearing method, and appearance of the headset are not limited to those in the embodiment of the present application. The hardware structure of the headset provided in the embodiment of the present application will be described below with reference to the wearing method of the headset in the human ear.

[0086] As shown in FIG. 1, the headset 101 includes a plurality of speakers (i.e., loudspeakers), a plurality of microphones, a microcontroller unit, (MThe plurality of speakers includes a first loudspeaker (CU), an ANC chip, and a memory. The plurality of speakers includes a plurality of first speakers, for example, loudspeaker 1 and loudspeaker 2. The plurality of first speakers need to participate in noise cancellation. For example, the first speakers are low-band and mid-band speakers, and the low-band and mid-band speakers need to participate in noise cancellation. Optionally, the plurality of speakers further includes at least one second speaker, and the at least one second speaker does not participate in noise cancellation. For example, the second speaker is a high-band speaker, and the high-band speaker does not need to participate in noise cancellation. Of course, for any speaker, regardless of whether the speaker is a high-band speaker or a low-band and mid-band speaker, the speaker may or may not participate in noise cancellation. In other words, in the embodiment of the present application, the sound-generating frequency band of the first speaker involved in noise cancellation is not limited, and the sound-generating frequency band of the second speaker not involved in noise cancellation is not limited. The multiple microphones include at least one reference microphone and one error microphone. Figure 1 will be described using one reference microphone as an example.

[0087] The speakers are configured to play downlink signals (e.g., audio signals such as music or voice). Each speaker includes an independent digital-to-analog converter. (D AC) and power amplifiers (P A). In other words, one speaker corresponds to one DAC and one PA, and a different speaker corresponds to a different DAC and PA. In the noise cancellation process, the first speaker is further configured to play anti-phase noise, which is used to reduce the noise signal in the user's ear canal to achieve an active noise cancellation effect.

[0088] The reference microphone is disposed outside the headset. After the headset is worn on a human ear, the reference microphone is located outside the human ear. The reference microphone is configured to collect a noise signal of an external environment. In the embodiment of the present application, the noise signal collected by the reference microphone is referred to as a reference signal.

[0089] The error microphone is disposed inside the headset. After the headset is worn on a human ear, the error microphone is located inside the human ear. The error microphone is configured to collect a noise signal in the ear canal. In the embodiment of the present application, the noise signal collected by the error microphone is referred to as an error signal.

[0090] The microcontrol unit is configured to process the reference signal collected by the reference microphone, the error signal collected by the error microphone, the downlink signal, etc. to determine a group of target noise cancellation parameters corresponding to each of the plurality of first speakers, and write the group of target noise cancellation parameters corresponding to each first speaker to the ANC chip.

[0091] The ANC chip is configured to process the reference signal collected by the reference microphone and the error signal collected by the error microphone based on a group of target noise cancellation parameters corresponding to each first speaker to generate anti-phase noise, perform audio mixing on the generated anti-phase noise and the downlink signal to be played through the first speaker, and output the mixed signal to the corresponding first speaker to reduce the noise signal in the ear canal.

[0092] The memory is configured to store initial parameters, mapping relationships, etc. used when determining target noise cancellation parameters corresponding to each first speaker.

[0093] It should be noted that the microcontroller unit, the ANC chip, and the memory may be integrated on the same circuit board or may be located on different circuit boards. This is not a limitation in the embodiments of the present application. In addition, the microcontroller unit and the ANC chip are distinguished only from the perspective of describing logical functions. In actual physical form, the microcontroller unit and the ANC chip may be integrated on one chip or may be located separately on multiple chips. For example, the microcontroller unit and the ANC chip may be located on two chips.

[0094] Optionally, headset 101 may further include another element, such as an optical proximity sensor, configured to detect whether headset 101 is in the ear. If headset 101 is a wireless headset, headset 101 may further include a wireless communication module, which may be a wireless local area network module or a Bluetooth module. The wireless communication module is used by headset 101 to communicate with another device.

[0095] It can be understood that the general structure in the embodiments of the present application does not constitute a limitation on the headset. In some other embodiments, the headset 101 may include more or fewer components than those shown in the figures, or some components may be combined, or some components may be divided, or there may be a different component arrangement. The components shown in the figures may be implemented using hardware, software, or a combination of software and hardware.

[0096] The system architecture and service scenarios described in the embodiments of the present application are intended to more clearly explain the technical solutions in the embodiments of the present application, and do not constitute limitations on the technical solutions provided in the embodiments of the present application. Those skilled in the art can know the following: With the evolution of system architecture and the emergence of new service scenarios, the technical solutions provided in the embodiments of the present application can also be applied to similar technical problems.

[0097] 2 is a flowchart of a noise cancellation method according to an embodiment of the present application. The method is applied to a headset, which includes at least one reference microphone, one error microphone, and multiple first speakers. See FIG. 2. The method includes the following steps:

[0098] Step 201: Determine a plurality of groups of target noise cancellation parameters that correspond one-to-one to a plurality of first speakers.

[0099] According to the noise cancellation method provided in this embodiment of the present application, the plurality of groups of target noise cancellation parameters can be determined on a frame-by-frame basis. In other words, the plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers are determined in each frame. Of course, the target noise cancellation parameters may alternatively be determined on another time basis. For example, the plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers are determined every two frames. In the following, a frame is used as a unit for description.

[0100] In some embodiments, the headset further includes a plurality of FF filters in one-to-one correspondence with the plurality of first speakers. In this case, the plurality of groups of target noise cancellation parameters include k-th frame filter coefficients of the plurality of FF filters, where k is an integer greater than or equal to 1. In some cases, the headset further includes a plurality of FB filters in one-to-one correspondence with the plurality of first speakers. In other words, the plurality of FB filters are in one-to-one correspondence with the plurality of FF filters. In this case, the plurality of groups of target noise cancellation parameters further include k-th frame filter coefficients of the plurality of FB filters. In addition, if the headset further includes a downlink compensation filter, the plurality of groups of target noise cancellation parameters further include k-th frame filter coefficients of the downlink compensation filter. In addition, if k is greater than 1, a target noise cancellation level can also be determined. Therefore, the four parts will be described separately below.

[0101] It should be noted that the multiple groups of target noise cancellation parameters may also be referred to as noise cancellation parameters of multiple noise cancellation channels, and one noise cancellation channel includes one FF filter and one first speaker. When the headset further includes multiple FB filters that correspond one-to-one to the multiple first speakers, one noise cancellation channel further includes one FB filter.

[0102] (1) The k-th frame filter coefficients of a plurality of FF filters are determined.

[0103] When k is equal to 1, the initial filter coefficients of the multiple FF filters are determined as the k-th frame filter coefficients of the multiple FF filters, i.e., the first frame filter coefficients of the multiple FF filters are the initial filter coefficients of the corresponding FF filters, or the k-th frame filter coefficients of the multiple FF filters are determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the FF filter coefficients. When k is greater than 1, the k-th frame filter coefficients of the multiple FF filters are determined based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the target noise canceling level. In other words, the k-th frame filter coefficients of the multiple FF filters are determined according to an adaptive method. The determination process is an adaptive process and may also be called an iterative process.

[0104] It should be noted that the initial filter coefficients of the multiple FF filters may be the same or different, and the initial filter coefficients may be zero or non-zero. This is not a limitation in the embodiments of the present application. The initial noise cancellation level may be a preset level at which noise cancellation can be successfully performed by using the corresponding noise cancellation coefficients without introducing stability issues. Of course, the initial noise cancellation level may alternatively be a level determined based on a prompt tone, such as "Noise cancellation on" or "Dingdong," transmitted by the user terminal when noise cancellation starts. The noise cancellation coefficients corresponding to this level may better adapt to the current human ear and wearing posture, and performing adaptive iterations based on the noise cancellation coefficients corresponding to this level may more quickly reach a convergence state. This is also not a limitation in the embodiments of the present application.

[0105] An implementation process for determining the k-th frame filter coefficients of a plurality of FF filters based on the (k-1)th frame reference signal collected by at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and a target noise canceling level includes the steps of determining the (k-1)th frame filter coefficients of a plurality of SPs based on the target noise canceling level and a mapping relationship between the noise canceling level and the filter coefficients of the SPs, where the plurality of SPs are paths from a plurality of first speakers to the error microphone; and determining the k-th frame filter coefficients of a plurality of FF filters based on the (k-1)th frame reference signal collected by at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficients of the plurality of SPs.

[0106] The multiple SPs may also be referred to as multiple noise cancellation channel SPs. The mapping relationship between the noise cancellation levels and the filter coefficients of the SPs includes multiple noise cancellation levels. There is a mapping relationship between each noise cancellation level and the filter coefficients of the multiple SPs, and the mapping relationship between different noise cancellation levels and the filter coefficients of the multiple SPs may be different. Therefore, after the target noise cancellation level is determined, filter coefficients corresponding to the multiple SPs can be obtained from the mapping relationship between the noise cancellation levels and the filter coefficients of the SPs based on the target noise cancellation level, and the obtained filter coefficients are used as the (k-1)th frame filter coefficients of the multiple SPs. The same applies to the initial noise cancellation level.

[0107] The k-th frame filter coefficients of the multiple FF filters can be determined using a multi-channel linkage method. In addition, when the headset includes multiple FF filters, the headset may further include multiple FB filters that correspond one-to-one to the multiple first speakers, or may not include multiple FB filters. In different cases, the methods for determining the k-th frame filter coefficients of the multiple FF filters are different. The methods will be described separately below.

[0108] Since the process of determining the k-th frame filter coefficient of an FF filter based on the (k-1)th frame reference signal collected by at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of a plurality of SPs is the same, one of the processes will be used as an example below for description. In other words, one of the plurality of FF filters is used as a target FF filter, and the k-th frame filter coefficient of the target FF filter is determined as follows: For the process of determining the k-th frame filter coefficient of another FF filter of the plurality of FF filters, please refer to the process of determining the k-th frame filter coefficient of the target FF filter.

[0109] In the first case, the headset does not include multiple FB filters. If the target FF filter is the first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficient of the target SP, where the target reference microphone is the reference microphone corresponding to the target FF filter, and the target SP is the path from the first speaker to the error microphone corresponding to the target FF filter. If the target FF filter is a non-first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the multiple SPs, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0110] If the target FF filter is the first FF filter, a residual is determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the (k-1)th frame filter coefficient of the target SP, and the residual. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0111] In some embodiments, one of the multiple FF filters corresponds to one reference microphone. In other words, the target reference microphone includes one reference microphone. In this case, the residual is determined according to the following formula (1) based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone:

number

[0112] In the above formula (1), Res k-1 indicates the residual, and Ref k-1 denotes the (k-1)th frame reference signal collected by the target reference microphone, and Err k-1 denotes the (k-1)th frame error signal collected by the error microphone.

[0113] In some other embodiments, one of the multiple FF filters corresponds to at least two reference microphones. In other words, the target reference microphone includes at least two reference microphones. In this case, audio mixing is performed on the (k-1)th frame reference signals collected by the at least two reference microphones included in the target reference microphone to obtain a (k-1)th frame mixed reference signal. The residual is determined based on the (k-1)th frame mixed reference signal and the (k-1)th frame error signal collected by the error microphone. In this way, the signal-to-noise ratio of the reference signal can be improved.

[0114] The method of determining the residual based on the mixed reference signal of the (k-1)th frame and the (k-1)th frame error signal collected by the error microphone is the same as the above method of determining the residual according to the above formula (1). Specifically, the (k-1)th frame error signal collected by the error microphone is divided by the mixed reference signal of the (k-1)th frame to obtain the residual.

[0115] In some embodiments, frequency response information of the (k-1)th frame filter coefficient of the target SP may be determined, and then the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the frequency response information of the (k-1)th frame filter coefficient of the target SP, and the residual, according to the following equation (2):

number

[0116] In the above formula (2), FF k denotes the k-th frame frequency response information of the target FF filter, and FF k-1 denotes the (k-1)th frame frequency response information of the target FF filter, μ denotes the step, which is preset, and SP k-1 denotes the frequency response information of the (k-1)th frame filter coefficient of the target SP.

[0117] If the target FF filter is a non-first FF filter, a residual is determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple SPs, and the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter. The kth frame filter coefficient of the target FF filter is determined based on the kth frame frequency response information of the target FF filter.

[0118] The method of determining the residual is the same as that described above. For the detailed implementation process, please refer to the above content. The details will not be described again in this specification.

[0119] When determining the kth frame frequency response information of the target FF filter, the kth frame frequency response information of the target FF filter can be determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficient of the target SP, the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter, and the (k-1)th frame filter coefficient of the SP corresponding to each FF filter before the target FF filter.

[0120] For example, the frequency response information of the (k-1)th frame filter coefficient of the target SP and the frequency response information of the (k-1)th frame filter coefficient of the SP corresponding to each FF filter before the target FF filter may be determined. Then, according to the following equation (3), the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the frequency response information of the (k-1)th frame filter coefficient of the target SP, the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter, and the frequency response information of the (k-1)th frame filter coefficient of the SP corresponding to each FF filter before the target FF filter.

number

[0121] In the above formula (3), FF i,k indicates the k-th frame frequency response information of the target FF filter, i.e., the target FF filter is the i-th FF filter among the multiple FF filters, and FF i,k-1 indicates the (k-1)th frame frequency response information of the target FF filter, and Res i,k-1 indicates the residual, and SP i,k-1 denotes the frequency response information of the (k-1)th frame filter coefficient of the target SP, and FF j,k denotes the k-th frame frequency response information of the j-th FF filter before the target FF filter, and FF j,k-1 denotes the (k-1)th frame frequency response information of the jth FF filter before the target FF filter, and SP j,k-1 denotes the frequency response information of the (k-1)th frame filter coefficient of the SP corresponding to the jth FF filter before the target FF filter.

[0122] An implementation process for determining the kth frame filter coefficients of the target FF filter based on the kth frame frequency response information of the target FF filter includes establishing a loss function between filter coefficient variables of the target FF filter and the kth frame frequency response information of the target FF filter. Based on the loss function, values ​​of the filter coefficient variables are determined using a gradient descent method, and the kth frame filter coefficients of the target FF filter are determined based on the values ​​of the filter coefficient variables. In other words, a loss function is established between the filter coefficient variables of the target FF filter and the kth frame frequency response information of the target FF filter. Optimal values ​​of the variables are determined using a gradient descent method, and the kth frame filter coefficients of the target FF filter are determined based on the optimal values ​​of the variables.

[0123] The filter coefficients of the target FF filter for each frame are determined according to a gradient descent method. Once the filter coefficients of the target FF filter for each frame are determined, a value of a loss function is determined. When the value of the loss function reaches a minimum threshold, it is determined that the filter coefficients of the target FF filter have reached a convergence and stability condition. For example, for the kth frame filter coefficient of the target FF filter, when the value of the loss function between the filter coefficient variable and the kth frame frequency response information of the target FF filter reaches a minimum threshold, it is determined that the kth frame filter coefficient of the target FF filter has reached a convergence and stability condition. When the value of the loss function does not reach the minimum threshold, it is determined that the kth frame filter coefficient of the target FF filter has not reached a convergence and stability condition. The minimum threshold may be preset and adjusted based on different requirements in different cases.

[0124] Optionally, the filter coefficients of each FF filter include at least one biquadratic filter coefficient and one gain. Variables corresponding to the biquadratic filter coefficient include a filter type, a cutoff frequency, and a quality factor. Of course, in practical applications, the filter coefficients of each FF filter may further include more or fewer other parameters. This is not limited in the embodiments of the present application.

[0125] The k-th frame filter coefficient of the first FF filter may be determined according to a related algorithm based on the values ​​of the filter coefficient variables, where the algorithm is not limited in the embodiments of the present application.

[0126] In some cases, background noise, i.e., noise floor issues, may occur in quiet environments. For example, semi-open headsets are more likely to have background noise issues in quiet environments than in-ear headsets. Additionally, strong noise cancellation is unnecessary in quiet environments, and some people find strong noise cancellation uncomfortable. Furthermore, a stronger noise cancellation intensity indicates a stronger sense of negative pressure. Therefore, when the value of the filter coefficient variable is determined using the gradient descent method, the target noise cancellation amplitude can be dynamically adjusted based on the ambient sound volume. The k-th frame filter coefficient of the target FF filter is determined based on this target noise cancellation amplitude, thereby improving the subjective experience of adaptive noise cancellation. In other words, the target noise cancellation amplitude is determined based on the ambient sound volume of the (k-1)th frame and the ambient sound volume in t frames prior to the (k-1)th frame, where t is greater than or equal to 1 and less than k-1. The value of the filter coefficient variable is determined based on the target noise cancellation amplitude and a loss function using the gradient descent method, and the k-th frame filter coefficient of the target FF filter is determined based on the value of the filter coefficient variable.

[0127] A target ambient volume is determined based on the ambient volume of the (k-1)th frame and the ambient volume in t frames prior to the (k-1)th frame. If the target ambient volume is equal to or less than a first volume threshold, the first noise cancellation amplitude is determined as the target noise cancellation amplitude. If the target ambient volume is greater than the first volume threshold, it is determined whether the target ambient volume has significantly increased or decreased, and if the target ambient volume has significantly increased, the noise cancellation amplitude of the (k-1)th frame is increased to obtain the target noise cancellation amplitude. If the ambient volume has significantly decreased, the noise cancellation amplitude of the (k-1)th frame is decreased to obtain the target noise cancellation amplitude. If the target ambient volume has neither significantly increased nor significantly decreased, the noise cancellation amplitude of the (k-1)th frame is determined as the target noise cancellation amplitude, i.e., the noise cancellation amplitude remains unchanged.

[0128] There are several methods for determining the target ambient volume based on the ambient volume of the (k-1)th frame and the ambient volume in t frames before the (k-1)th frame, for example, methods for obtaining an arithmetic average value or a weighted average value. This is not a limitation in the embodiment of the present application. The t frames may be any t frames before the (k-1)th frame, or may be the t frames before the (k-1)th frame and closest to the (k-1)th frame. This is not a limitation in the embodiment of the present application.

[0129] It should be noted that the first volume threshold is preset and indicates whether the environment is currently quiet. In other words, if the target environment volume is equal to or less than the first volume threshold, it indicates that the environment is quiet. If the target environment volume is greater than the first volume threshold, it indicates that the environment is not quiet. The first noise cancellation amplitude is preset for a quiet environment and is used to perform weak noise cancellation, thereby avoiding excessive amplification of background noise or introducing subjective comfort issues. In practical applications, the first volume threshold and the first noise cancellation amplitude can be adjusted based on different requirements.

[0130] For example, see FIG. 3. Whether the environment is quiet is determined based on the target environmental volume. If the environment is quiet, the first noise cancellation amplitude is determined as the target noise cancellation amplitude. In a non-quiet environment, if the target environmental volume increases significantly, the noise cancellation amplitude of the (k-1)th frame is increased to obtain the target noise cancellation amplitude. If the target environmental volume decreases significantly, the noise cancellation amplitude of the (k-1)th frame is decreased to obtain the target noise cancellation amplitude. If the target environmental volume neither increases nor decreases significantly, the noise cancellation amplitude of the (k-1)th frame is determined as the target noise cancellation amplitude, i.e., the noise cancellation amplitude remains unchanged.

[0131] There are several ways to determine whether the target environment volume has significantly increased or decreased. For example, if the currently determined target environment volume is higher than the previously determined target environment volume and the difference between the currently determined target environment volume and the previously determined target environment volume is greater than a second volume threshold, it is determined that the currently determined target environment volume has significantly increased. Similarly, if the currently determined target environment volume is lower than the previously determined target environment volume and the difference between the currently determined target environment volume and the previously determined target environment volume is greater than a second volume threshold, it is determined that the currently determined target environment volume has significantly decreased.

[0132] The second volume threshold is also preset, for example, 3 dB. In practical applications, the second volume threshold can be further adjusted based on different requirements.

[0133] In a second case, the headset further includes a plurality of FB filters. When the target FF filter is the first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, and the (k-1)th frame filter coefficients of the plurality of FB filters. When the target FF filter is a non-first FF filter, the k-th frame filter coefficient of the target FF filter is determined based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, the (k-1)th frame filter coefficients of the plurality of FB filters, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0134] When the target FF filter is the first FF filter, a residual may be determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple FB filters, and the (k-1)th frame filter coefficients of the multiple SPs. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0135] In one example, frequency response information of the (k-1)th frame filter coefficients of the plurality of FB filters and frequency response information of the (k-1)th frame filter coefficients of the plurality of SPs may be determined. Then, according to the following Equation (4), the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the frequency response information of the (k-1)th frame filter coefficients of the plurality of FB filters, and the frequency response information of the (k-1)th frame filter coefficients of the plurality of SPs.

number

[0136] In the above formula (4), FF 1,k denotes the k-th frame frequency response information of the target FF filter, and FF 1,k-1 indicates the (k-1)th frame frequency response information of the target FF filter, and Res 1,k-1 indicates the residual, and SP 1,k-1 indicates the frequency response information of the (k-1)th frame filter coefficient of the target SP, and FB j,k-1 indicates the frequency response information of the (k-1)-th frame filter coefficient of the j-th FB filter among the multiple FB filters, and SP j,k-1denotes frequency response information of the (k-1)th frame filter coefficient of the SP corresponding to the jth FB filter, and n denotes the total number of FB filters, that is, the total number of noise cancellation channels.

[0137] If the target FF filter is a non-first FF filter, the residual may be determined based on the (k-1)th frame reference signal collected by the target reference microphone and the (k-1)th frame error signal collected by the error microphone, and the k-th frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the (k-1)th frame filter coefficients of the multiple SPs, the (k-1)th frame filter coefficients of the multiple FB filters, and the k-th frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter. The k-th frame filter coefficient of the target FF filter is determined based on the k-th frame frequency response information of the target FF filter.

[0138] In one example, frequency response information of the (k-1)th frame filter coefficients of the plurality of SPs and frequency response information of the (k-1)th frame filter coefficients of the plurality of FB filters may be determined. Then, according to the following equation (5), the kth frame frequency response information of the target FF filter is determined based on the (k-1)th frame frequency response information of the target FF filter, the residual, the frequency response information of the (k-1)th frame filter coefficients of the plurality of SPs, the frequency response information of the (k-1)th frame filter coefficients of the plurality of FB filters, and the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

number

[0139] In the above formula (5), FB j,k-1indicates the frequency response information of the (k-1)th frame filter coefficient of the jth FB filter among the multiple FB filters, n indicates the total number of the multiple FB filters, i.e., the total number of the multiple noise cancellation channels, and the meanings represented by the other letters are the same as those in the above equation (3).

[0140] The implementation process of determining the k-th frame filter coefficients of the target FF filter based on the k-th frame frequency response information of the target FF filter is the same as the first case. For details, please refer to the above description. The details will not be described again in this specification. In addition, the frequency response information of the filter coefficients of the SP can be determined according to a related algorithm based on the filter coefficients of the SP, and the frequency response information of the FB filter coefficients can also be determined according to a related algorithm based on the filter coefficients of the FB filter. The algorithm is not limited in the embodiments of the present application.

[0141] In the above-described process for determining the k-th frame frequency response information of the target FF filter, regardless of whether the headset includes a target FB filter, the k-th frame frequency response information of the target FF filter is determined based on the (k-1)-th frame filter coefficient of the target SP, and the (k-1)-th frame filter coefficient of the target SP is determined based on the target noise canceling level by querying the mapping relationship between the noise canceling level and the filter coefficient of the SP. Specifically, the (k-1)-th frame filter coefficient of the target SP is an estimated value, and the k-th frame frequency response information of the target FF filter is determined based on this estimated value, thereby eliminating dependence on the actual value of the target SP and enabling adaptation of the filter coefficient of the FF filter even when there is no downlink signal.

[0142] (2) The k-th frame filter coefficients of a plurality of FB filters are determined.

[0143] When k is equal to 1, the initial filter coefficients of the plurality of FB filters are determined as the k-th frame filter coefficients of the plurality of FB filters, that is, the first frame filter coefficients of the plurality of FB filters are the initial filter coefficients of the corresponding FB filters, or the k-th frame filter coefficients of the plurality of FB filters are determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficients. When k is greater than 1, the k-th frame filter coefficients of the plurality of FB filters can be determined based on the target noise canceling level.

[0144] It should be noted that the initial filter coefficients of the multiple FB filters may be the same or different, and the initial filter coefficients may be 0 or not 0. This is not limited in the embodiments of the present application.

[0145] Since the process of determining the k-th frame filter coefficient of the FB filter based on the target noise canceling level is the same, the following description will use one of the processes as an example. In other words, one of the multiple FB filters is used as the target FB filter, and the k-th frame filter coefficient of the target FB filter is determined in the following two ways: For the process of determining the k-th frame filter coefficient of another FB filter of the multiple FB filters, please refer to the process of determining the k-th frame filter coefficient of the target FB filter. In other words, when k is greater than 1, the k-th frame filter coefficient of the target FB filter can be determined in the following two ways:

[0146] In the first method, the k-th frame filter coefficient of the target FB filter is determined based on the target noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficient.

[0147] The mapping relationship between the noise canceling level and the FB filter coefficient includes a plurality of noise canceling levels, and there is a mapping relationship between each noise canceling level and the filter coefficients of the plurality of FB filters, and the mapping relationships between different noise canceling levels and the filter coefficients of the plurality of FB filters may be different. Therefore, a filter coefficient corresponding to the target FB filter can be obtained from the mapping relationship between the noise canceling level and the FB filter coefficient based on the target noise canceling level, and the obtained filter coefficient is used as the k-th frame filter coefficient of the target FB filter.

[0148] Since the mapping relationship between the noise canceling level and the FB filter coefficient is stored in advance, determining the k-th frame filter coefficient of the target FB filter in the first method is stable, simple in operation, and highly efficient.

[0149] In the second method, when the target FB filter is a first type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on a target noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficient. When the target FB filter is a second type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on a (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level.

[0150] Similarly to the above description, when the target FB filter is a second type FB filter, the k-th frame filter coefficient of the target FB filter can be determined according to an adaptive method. The process of determining the k-th frame filter coefficient of the target FB filter is an adaptive process, which may also be called an iterative process.

[0151] An implementation process for determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level includes: a step of determining the (k-1)th frame filter coefficient of the target SP based on the target noise canceling level and a mapping relationship between the noise canceling level and the filter coefficient of the SP, where the target SP is a path from the first speaker corresponding to the target FB filter to the error microphone; and a step of determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the (k-1)th frame filter coefficient of the target SP.

[0152] According to the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the (k-1)th frame filter coefficient of the target SP, the kth frame filter coefficient of the target FB filter can be determined according to a related algorithm. The algorithm is not limited in the embodiments of the present application.

[0153] The first frame filter coefficient of the target FB filter may be an initial filter coefficient, or may be determined based on the initial noise canceling level by querying the mapping relationship between the noise canceling level and the FB filter coefficient. Therefore, when k is 1 or greater, this is equivalent to determining the kth frame filter coefficient of the target FB filter in three ways. Specifically, (1) the kth frame filter coefficient of the target FB filter is determined by querying the mapping relationship between the noise canceling level and the FB filter coefficient. (2) If the target FB filter is a first type FB filter, the kth frame filter coefficient of the target FB filter is determined by querying the mapping relationship between the noise canceling level and the FB filter coefficient. If the target FB filter is a second type FB filter, the kth frame filter coefficient of the target FB filter is determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level. (3) When the target FB filter is a first type FB filter, or when the target FB filter is a second type FB filter and k is equal to 1, the k-th frame filter coefficient of the target FB filter is determined by consulting the mapping relationship between the noise canceling level and the FB filter coefficient. When the target FB filter is a second type FB filter and k is greater than 1, the k-th frame filter coefficient of the target FB filter is determined based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level.

[0154] The sound-producing frequency band of the first speaker corresponding to the first type of FB filter is higher than the sound-producing frequency band of the first speaker corresponding to the second type of FB filter. In other words, the first speaker corresponding to the first type of FB filter is a high-bandwidth speaker, and the first speaker corresponding to the second type of FB filter is a low-bandwidth speaker. Of course, the first type of FB filter and the second type of FB filter may be distinguished in a manner other than based on the sound-producing frequency band. This is also not a limitation of the embodiments of the present application.

[0155] In the second and third methods described above, the method of querying the mapping relationship between the noise cancellation level and the FB filter coefficient is combined with the adaptive method, which can improve the noise cancellation effect without high complexity and with controllable stability.

[0156] In the embodiment of the present application, the k-th frame filter coefficient of the target FB filter may be determined by the three methods described above, and it should be noted that the k-th frame filter coefficient of the target FB filter may alternatively be determined by another method. For example, regardless of whether the target FB filter is a first type FB filter or a second type FB filter, the k-th frame filter coefficient of the target FB filter is determined based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level. This is not limited to the embodiment of the present application.

[0157] (3) Determine the k-th frame filter coefficient of the downlink compensation filter.

[0158] When k is equal to 1, the initial downlink compensation filter coefficient is determined as the k-th frame filter coefficient of the downlink compensation filter, or the k-th frame filter coefficient of the downlink compensation filter is determined based on the initial noise canceling level and the mapping relationship between the noise canceling level and the downlink compensation filter coefficient. When k is greater than 1, the k-th frame filter coefficient of the downlink compensation filter is determined based on the target noise canceling level and the mapping relationship between the noise canceling level and the downlink compensation filter coefficient.

[0159] The mapping relationship between the noise canceling level and the downlink compensation filter coefficient includes a plurality of noise canceling levels, and there is a mapping relationship between each noise canceling level and the filter coefficient of the downlink compensation filter, and the mapping relationship between different noise canceling levels and the filter coefficient of the downlink compensation filter may be different. Therefore, after the target noise canceling level is determined, the corresponding downlink compensation filter coefficient can be obtained from the mapping relationship between the noise canceling level and the downlink compensation filter coefficient based on the target noise canceling level, and the obtained downlink compensation filter coefficient is used as the k-th frame filter coefficient of the downlink compensation filter.

[0160] (4) Determine the target noise canceling level.

[0161] A (k-1)th frame noise canceling level is determined, and noise canceling levels in m frames before the (k-1)th frame are obtained, where m is greater than or equal to 1 and less than k-1. A target noise canceling level is determined based on the (k-1)th frame noise canceling level and the noise canceling levels in m frames.

[0162] In the (k-1)th frame, there may be a valid downlink signal, there may not be a valid downlink signal, the environment may be quiet, the environment may not be quiet, or of course, there may be an abnormal signal. In different cases, the manner of determining the (k-1)th frame noise canceling level is different, which will be described separately below.

[0163] In the first case, in the (k-1)th frame, there is no valid downlink signal and the environment is not quiet. In this case, the (k-1)th frame noise canceling level is determined based on the reference filter coefficients of multiple FF filters and the mapping relationship between the noise canceling level and the frequency response information of the FF filters. When k is equal to 2, the reference filter coefficient is the initial filter coefficient of the corresponding FF filter; or when k is greater than 2, the reference filter coefficient is the filter coefficient of the corresponding FF filter that last satisfies the convergence and stability condition before the kth frame, or is the (k-1)th frame filter coefficient of the corresponding FF filter.

[0164] When an audio signal is played through a headset, for example, when music is played or a call is made, the user terminal sends control signaling to the headset to play the audio signal. Therefore, whether the headset is currently in a downlink enabled state can be determined based on whether the headset receives the control signaling. When the headset is not in a downlink enabled state, it is determined that a valid downlink signal does not exist in the (k-1)th frame. When the headset is in a downlink enabled state, sound is not necessarily output continuously in the (k-1)th frame. For example, sound is not output during speech pauses and transition periods between music tracks, and these periods are usually not short. Therefore, when the headset is in a downlink enabled state, it can further be determined whether the (k-1)th frame is within a downlink discontinuity period. If the (k-1)th frame is within a downlink discontinuity period, it is determined that a valid downlink signal does not exist in the (k-1)th frame. If the (k-1)th frame is not within a downlink discontinuity period, it is determined that a valid downlink signal exists in the (k-1)th frame.

[0165] If there is no valid downlink signal in the (k-1)th frame but the environment is not quiet, the (k-1)th frame noise canceling level may change according to different environmental noises. Therefore, the (k-1)th frame noise canceling level needs to be determined based on the reference filter coefficients of multiple FF filters and the mapping relationship between the noise canceling level and the frequency response information of the FF filters.

[0166] In some embodiments, reference frequency response information of the plurality of FF filters is determined based on reference filter coefficients of the plurality of FF filters.Noise canceling levels that match the reference frequency response information of the plurality of FF filters are determined based on a mapping relationship between the noise canceling levels and the frequency response information of the FF filters to obtain a plurality of reference noise canceling levels.The (k-1)th frame noise canceling level is determined based on the plurality of reference noise canceling levels.

[0167] The reference frequency response information of the plurality of FF filters may be determined based on the reference filter coefficients of the plurality of FF filters according to a related algorithm, which is not limited in the embodiments of the present application.

[0168] When the noise canceling levels are different, the frequency response information of the FF filters may also be different. Therefore, a mapping relationship between the noise canceling levels and the frequency response information of the FF filters can be pre-stored. In this way, after the reference frequency response information of the multiple FF filters is determined, for any one of the multiple FF filters, a matching is performed between the reference frequency response information of the FF filter and the frequency response information of the FF filter at different noise canceling levels in the mapping relationship. From this mapping relationship, frequency response information matching the reference frequency response information of the FF filter is determined, and then the noise canceling level corresponding to the matched frequency response information is used as the reference noise canceling level. The other FF filters of the multiple FF filters are processed in the same manner, thereby obtaining multiple reference noise canceling levels.

[0169] The frequency response information of the FF filter can be represented by using a frequency response curve. Therefore, after the reference frequency response curves of the multiple FF filters are determined, for any one of the multiple FF filters, matching is performed between the reference frequency response curve of the FF filter and the frequency response curves of the FF filter at different noise canceling levels in a mapping relationship.

[0170] In practical application, matching can be performed between the complete reference frequency response curve of the FF filter and the complete frequency response curve of the FF filter at different noise canceling levels in the mapping relationship. Alternatively, matching can be performed between a curve that is within the reference frequency response curve of the FF filter and within the target frequency band and a curve that is within the frequency response curve of the FF filter at different noise canceling levels in the mapping relationship and within the target frequency band. This is not limited to the embodiments of the present application.

[0171] It should be noted that the target frequency band is a frequency band that has a clearly prominent feature in the frequency response curve, and the target frequency band is preset. For example, the target frequency band is a frequency band from 100 Hertz (Hz) to 200 Hertz (Hz). Of course, the value of the target frequency band may be different depending on the acoustic conditions of the headset.

[0172] For example, the mapping relationship between the noise canceling level and the frequency response information of the FF filter includes frequency response curves of the FF filter at 16 noise canceling levels, which are shown in FIG. 4. Because the features in the 100 Hz to 200 Hz frequency band in FIG. 4 are clearly prominent, the 100 Hz to 200 Hz frequency band is used as the target frequency band. Then, a matching is performed between the reference frequency response curve of the FF filter, which falls within the range of 100 Hz to 200 Hz, and the frequency response curve of the FF filter at the 16 noise canceling levels, which falls within the range of 100 Hz to 200 Hz.

[0173] There are several ways to determine the noise canceling level of the (k-1)th frame based on the multiple reference noise canceling levels. For example, the noise canceling level of the (k-1)th frame may be determined based on the average value of the multiple reference noise canceling levels. Alternatively, the noise canceling level of the (k-1)th frame may be determined based on the reference noise canceling level having the maximum amount among the multiple reference noise canceling levels.

[0174] When the (k-1)th frame noise canceling level is determined based on the average value of multiple reference noise canceling levels, the average value of the multiple reference noise canceling levels may be determined as the (k-1)th frame noise canceling level, or the average value of the multiple reference noise canceling levels may be adjusted to obtain the (k-1)th frame noise canceling level. Similarly, when the (k-1)th frame noise canceling level is determined based on the reference noise canceling level having the largest amount among the multiple reference noise canceling levels, the reference noise canceling level having the largest amount among the multiple reference noise canceling levels may be determined as the (k-1)th frame noise canceling level, or alternatively, the reference noise canceling level having the largest amount among the multiple reference noise canceling levels may be adjusted to obtain the (k-1)th frame noise canceling level. This is not limited to the embodiments of the present application.

[0175] Since the process of determining the filter coefficients of the FF filter is an iterative process, also known as an adaptive process, the aforementioned convergence and stability condition indicates that the filter coefficients of the FF filter converge and remain essentially unchanged. In addition, since the filter coefficients of the FF filter can be adaptively adjusted multiple times throughout the noise cancellation process, when the k-th frame filter coefficients of the FF filter are determined, the filter coefficients of the FF filter that last satisfy the convergence and stability condition before the k-th frame may be used as the reference filter coefficients, or the (k-1)-th frame filter coefficients of the FF filter may be used as the reference filter coefficients.

[0176] In the second case, a valid downlink signal exists in the (k-1)th frame, and the (k-1)th frame noise canceling level is determined based on the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by the error microphone.

[0177] In consideration of the above description, when the headset is in a downlink enabled state and is not in a downlink intermittent period, it is determined that a valid downlink signal exists in the (k-1)th frame. In this case, based on the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by the error microphone, the valid downlink signal can be extracted from the (k-1)th frame error signal collected by the error microphone, and a (k-1)th frame noise canceling level can be determined based on the extracted valid downlink signal.

[0178] According to the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by the error microphone, the valid downlink signal can be extracted from the (k-1)th frame error signal collected by the error microphone according to a related algorithm, and the (k-1)th frame noise canceling level can be determined based on the extracted valid downlink signal. The algorithm is not limited in the embodiments of the present application.

[0179] In the third case, there is no valid downlink signal in the (k-1)th frame, and the environment is quiet or an abnormal noise signal exists in the (k-1)th frame. In this case, the (k-3)th frame noise canceling level is determined as the (k-1)th frame noise canceling level. In other words, the noise canceling level remains unchanged.

[0180] In the (k-1)th frame, when there is no valid downlink signal and the environment is quiet, the noise basically remains unchanged. In this case, the noise canceling level remains unchanged. If there is an abnormal noise signal in the (k-1)th frame, the noise canceling level remains unchanged to perform robustness control and avoid divergence of the noise canceling level.

[0181] Abnormal noise signals indicate signals that seriously affect the user's listening experience, such as feedback, clipping, background noise, and wind noise. Feedback is a phenomenon in which the amplitude or energy of a single-frequency acoustic signal suddenly increases from a small value, and is usually caused by a user gripping a headset tightly or rapidly changing the headset's wearing position. The acoustic signal emitted during feedback is called feedback noise. Feedback causes discomfort to the user, interferes with the playback of downlink signals, and seriously affects the audio playback effect. Clipping is a phenomenon in which a low-band signal overflows, causing a crackling noise. The generated crackling noise is called clipping noise. Clipping generally occurs when a loud low-band noise bursts in the environment. For example, a loud low-band noise occurs when a vehicle crashes or an airplane lands. Background noise is ground noise, and is sometimes referred to as the noise floor. Background noise is noise caused by performance limitations of the device's hardware (e.g., circuits or other components in the headset), such as an unpleasant sound other than the program sound in a television broadcast. In a noisy environment, a user cannot perceive or hear the background noise. When the environment is quiet, a user can perceive the background noise. When the background noise is too strong, it not only makes people uncomfortable but also buries the subtle details in the sound. Wind noise occurs when there is wind in the environment. Wind noise affects the user's normal use of the headset. In addition, since the direction of the wind noise is randomized, the impact of wind noise on the user's ears is different. In other words, the left ear and right ear hear differently under the influence of wind noise.

[0182] The three cases are briefly summarized below with reference to FIG. 5. Referring to FIG. 5, if an abnormal noise signal exists in the (k-1)th frame, the noise canceling level remains unchanged. If no abnormal noise signal exists in the (k-1)th frame, it is determined whether downlink enablement is performed in the (k-1)th frame. If downlink enablement is not performed in the (k-1)th frame, it is determined whether the environment is quiet in the (k-1)th frame. If the environment is quiet in the (k-1)th frame, the noise canceling level remains unchanged. If the environment is not quiet in the (k-1)th frame, a corresponding reference noise canceling level is determined based on the reference filter coefficients of the FF filter. After polling of all channels is completed, the noise canceling level for the (k-1)th frame is determined based on multiple reference noise canceling levels. If downlink enablement is performed in the (k-1)th frame, it is determined whether the (k-1)th frame is in a downlink discontinuity period. When the (k-1)th frame is in a downlink intermittent period, multiple reference noise canceling levels are determined in the same manner, and then the (k-1)th frame noise canceling level is determined based on the multiple reference noise canceling levels.When the (k-1)th frame is not in a downlink intermittent period, the (k-1)th frame noise canceling level is determined based on the valid downlink signal of the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by an error microphone.

[0183] After the (k-1)th frame noise canceling level is determined in the three cases mentioned above, the target noise canceling level can be determined by integrating the (k-1)th frame noise canceling level with the noise canceling levels in the m frames prior to the (k-1)th frame.

[0184] The noise canceling level in the m frames may be the noise canceling level in any m frames before the (k-1)th frame, or the noise canceling level in the m frames before and closest to the (k-1)th frame. This is not limited to the embodiments of the present application. In addition, there are multiple implementation forms for determining the target noise canceling level based on the (k-1)th frame noise canceling level and the noise canceling levels in the m frames before the (k-1)th frame. For example, the noise cancellation effect is evaluated according to a related algorithm to determine the noise cancellation probability corresponding to the (k-1)th frame noise canceling level and the noise cancellation probability corresponding to the noise canceling levels in the m frames, and the noise canceling level with the maximum noise cancellation probability is determined as the target noise canceling level. Alternatively, the arithmetic average or weighted average of the (k-1)th frame noise canceling level and the noise canceling levels in the m frames is determined to obtain the target noise canceling level. Alternatively, the noise canceling level that appears most frequently among the (k-1)th frame noise canceling level and the noise canceling levels in the m frames may be determined as the target noise canceling level.

[0185] The various mapping relationships are determined in advance. For example, when one of the multiple first speakers is operating and the other first speakers are not operating, the various mapping relationships are determined based on a reference signal collected by at least one reference microphone and an error signal collected by an error microphone in each of the multiple leakage states. The multiple leakage states are formed by the headset and multiple different ear canal environments, and the multiple leakage states correspond one-to-one to the multiple noise canceling levels.

[0186] In this case, multiple groups of target noise cancellation parameters are determined. The process of determining multiple groups of target noise cancellation parameters is briefly summarized below using FIG. 6 as an example. See FIG. 6. Initial values, including the initial noise cancellation level, initial filter coefficients, and various mapping relationships described above, can be set offline. Then, in the (k-1)th frame, it is determined whether a valid downlink signal exists, whether the environment is quiet, and whether an abnormal noise signal exists, and the (k-1)th frame noise cancellation level is determined based on the different cases. The target noise cancellation level is determined based on the (k-1)th frame noise cancellation level and the previous noise cancellation levels in m frames. Next, a target noise cancellation amplitude is determined based on the (k-1)th frame environmental volume, and FB filter coefficient adaptation is performed based on the target noise cancellation level to determine the kth frame filter coefficients of multiple FB filters. Finally, FF filter coefficient adaptation is performed based on the target noise cancellation level and the target noise cancellation amplitude to determine the kth frame filter coefficients of multiple FF filters.

[0187] Step 202: Based on the multiple groups of target noise cancellation parameters, multiple groups of target anti-phase noises are generated that correspond one-to-one to the multiple first speakers, and the frequency bands of each target anti-phase noise in the multiple groups of target anti-phase noises cover the sound generation frequency bands of the multiple first speakers.

[0188] In consideration of the above description, the multiple groups of target noise cancellation parameters may be referred to as noise cancellation parameters of multiple noise cancellation channels. In this way, the multiple groups of generated target anti-phase noise may also be referred to as anti-phase noise of multiple noise cancellation channels. Because the process of generating anti-phase noise for the noise cancellation channels is the same, the following uses one of the noise cancellation channels as an example for explanation.

[0189] One of the multiple noise cancellation channels is used as a target noise cancellation channel, and the target noise cancellation channel includes a target FF filter and a target first speaker, and a reference microphone corresponding to the target FF filter is called a target reference microphone. In this case, the target anti-phase noise includes feedforward anti-phase noise. In other words, the kth frame reference signal collected by the target reference microphone is processed based on the kth frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise.

[0190] In consideration of the foregoing description, the target reference microphone may include one reference microphone or at least two reference microphones. When the target reference microphone includes one reference microphone, the k-th frame reference signal collected by the target reference microphone may be directly processed based on the k-th frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise. When the target reference microphone includes at least two reference microphones, audio mixing is performed on the k-th frame reference signals collected by the at least two reference microphones to obtain a mixed reference signal of the k-th frame, and then the mixed reference signal of the k-th frame is processed based on the k-th frame filter coefficient of the target FF filter to obtain the feedforward anti-phase noise.

[0191] If the headset further includes an FB filter, the target noise cancellation channel further includes a target FB filter. In this case, the target anti-phase noise further includes feedback anti-phase noise. In other words, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal based on the k-th frame filter coefficient of the downlink compensation filter. Then, negation processing is performed on the k-th frame downlink signal obtained through downlink compensation, and audio mixing is performed on the negated k-th frame downlink signal and the k-th frame error signal collected by the error microphone to obtain the k-th frame noise signal collected by the error microphone. The k-th frame noise signal collected by the error microphone is processed based on the k-th frame filter coefficient of the target FB filter to obtain the feedback anti-phase noise.

[0192] Downlink compensation can be used to remove all downlink signals in the error signal collected by the error microphone, so that noise cancellation is only performed on the residual noise signal through the FB filter to avoid damaging the sound quality of the downlink signal. In addition, downlink compensation is performed on the k-th frame downlink signal transmitted by the user terminal, so that the downlink signals of all speakers in the error microphone can be removed to avoid damaging the sound quality of the full-band downlink signal.

[0193] Considering the above description, when multiple groups of target noise cancellation parameters are determined for each frame, a frame may include one sample point or multiple sample points, so when the target anti-phase noise is generated, a group of target anti-phase noise may be generated at each sample point, or a group of target anti-phase noise may be generated in one frame.

[0194] In an embodiment of the present application, when the multiple groups of target noise cancellation parameters are determined, no frequency division is performed on the downlink signal, i.e., the multiple groups of target noise cancellation parameters are determined based on the full-band downlink signal. In this way, after the multiple groups of target anti-phase noises that correspond one-to-one to the multiple first speakers are generated based on the multiple groups of target noise cancellation parameters, the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound-generating frequency band of the multiple first speakers, i.e., the frequency band of each target anti-phase noise is the full frequency band.

[0195] Step 203: Perform noise cancellation via multiple first speakers by using multiple groups of target anti-phase noises.

[0196] After multiple groups of target anti-phase noise are generated, the multiple groups of target anti-phase noise are respectively mixed with the kth frame downlink signal to be played through multiple first speakers, and then the mixed signal is played through the corresponding first speakers to achieve noise cancellation.

[0197] Some of the multiple first speakers may be high-band speakers and others may be low-band speakers. Alternatively, some of the multiple first speakers may be full-band speakers and others may not be full-band speakers. In other words, the sound generation frequency bands of the multiple first speakers may be different. Alternatively, all of the multiple first speakers are full-band speakers. Alternatively, all of the multiple first speakers are not full-band speakers. If all of the multiple first speakers are full-band speakers, the k-th frame downlink signals to be played through the multiple first speakers are all k-th frame downlink signals transmitted by the user terminal. If all of the multiple first speakers are not full-band speakers, it is necessary to perform frequency division on the k-th frame downlink signals transmitted by the user terminal based on the sound generation frequency bands of each first speaker to obtain the k-th frame downlink signals to be played through each first speaker.

[0198] Two of the first speakers of the plurality of first speakers may include two first speakers formed by one dual-diaphragm (also called dual-dynamic) loudspeaker. Alternatively, the plurality of first speakers may include a plurality of split speakers / loudspeakers.

[0199] Optionally, the headset may further include at least one second speaker, where the at least one second speaker is not involved in noise cancellation. In this case, the second speaker may be involved in downlink compensation (i.e., downlink compensation is performed on a downlink signal transmitted by the user terminal, and the downlink signal is a full-band audio signal including an audio signal in the sound-generating frequency band of the second speaker). In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker may be a high-band speaker, a mid-band speaker, or a low-band speaker. Optionally, the second speaker may not be involved in downlink compensation. In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker is a high-band speaker.

[0200] It should be noted that the sound-producing frequency band of the at least one second speaker is higher than the sound-producing frequency band of the at least one first speaker. Of course, the sound-producing frequency band of the at least one second speaker may alternatively be lower than the sound-producing frequency band of the at least one first speaker. This is not a limitation in the embodiments of the present application.

[0201] In addition, the aforementioned process of determining the multiple groups of target noise cancellation parameters according to the adaptive method requires a certain time. When a frame includes multiple sample points and the duration of the frame is long, the duration of determining the multiple groups of target noise cancellation parameters is shorter than the duration of the frame. Therefore, during a part of the period of the kth frame, calculations may be performed based on the relevant data of the (k-1)th frame to obtain the multiple groups of target noise cancellation parameters for the kth frame, and during another part of the period of the kth frame, active noise cancellation may be performed based on the multiple groups of target noise cancellation parameters for the kth frame. However, when a frame includes one sample point or when a frame includes multiple sample points and the duration of the frame is short, the duration of determining the multiple groups of target noise cancellation parameters may be equal to the duration of the frame. In this case, it may be necessary to perform calculations for the entire period of the kth frame based on the relevant data of the (k-1)th frame to obtain the multiple groups of target noise cancellation parameters. In this case, the multiple groups of target noise cancellation parameters can be determined as the multiple groups of target noise cancellation parameters in the (k+1)th frame, and then active noise cancellation is performed during the (k+1)th frame based on the multiple groups of target noise cancellation parameters in the (k+1)th frame. The above content will be described using the former case as an example.

[0202] In conclusion, in the embodiment of the present application, the multiple groups of target anti-phase noises correspond one-to-one to the multiple first speakers, and the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound-generating frequency band of the multiple first speakers. In other words, each target anti-phase noise is a full-band anti-phase noise. Therefore, regardless of whether the first speaker is a high-band speaker, a low-band speaker, or a full-band speaker, when the multiple groups of target anti-phase noises are used to perform noise cancellation, the noise cancellation capability of each first speaker can be fully utilized. In other words, in a headset architecture including multiple noise cancellation channels and multiple speakers, this solution can improve the noise cancellation effect of the headset by using the full-band anti-phase noise of the multiple noise cancellation channels.

[0203] In the following, by way of example, some possible headset architectures are described in embodiments of the present application.

[0204] 7 is a diagram of a headset structure according to an embodiment of the present application. Please refer to FIG. 7. The headset includes f reference microphones, one error microphone, n FF filters, n FF adaptation engines that correspond one-to-one to the n FF filters, n FB filters, n FB adaptation engines that correspond one-to-one to the n FB filters, n first speakers (i.e., speaker 1 to speaker n), a downlink compensation filter, a downlink compensation adaptation engine (not shown), a digital divider, and n equalizers. (E Q) Calibrator. Both f and n are integers equal to or greater than 1, and f and n may or may not be equal.

[0205] The f reference microphones are configured to collect noise signals from the external environment, i.e., reference signals. The error microphone is configured to collect noise signals in the ear canal, i.e., error signals. The n FF adaptation engines are configured to determine k-th frame filter coefficients of FF filters corresponding to each of the n FF adaptation engines and refresh the corresponding FF filters with the determined k-th frame filter coefficients. The n FB adaptation engines are configured to determine k-th frame filter coefficients of FB filters corresponding to each of the n FB adaptation engines and refresh the corresponding FB filters with the determined k-th frame filter coefficients. The downlink compensation adaptation engine is configured to determine k-th frame filter coefficients of downlink compensation filters and refresh the downlink compensation filters with the determined k-th frame filter coefficients.

[0206] The digital divider is configured to perform frequency division on the k-th frame downlink signal transmitted by the user terminal based on the sound generation frequency bands of the n first speakers to obtain the k-th frame downlink signal corresponding to each first speaker. The n EQ calibrators are configured to calibrate mass production parameters of the corresponding first speakers, so that tolerances of the mass production parameters of the n first speakers are consistently matched.

[0207] When noise cancellation is performed, the n FF filters are configured to process the kth frame reference signal collected by the reference microphone corresponding to the n FF filters based on their respective kth frame filter coefficients to obtain feedforward anti-phase noise. The downlink compensation filter is configured to perform downlink compensation on the kth frame downlink signal transmitted by the user terminal based on the kth frame filter coefficients of the downlink compensation filter. Then, after negation processing is performed on the kth frame downlink signal obtained through downlink compensation, audio mixing is performed on the negated kth frame downlink signal and the kth frame error signal collected by the error microphone to obtain the kth frame noise signal collected by the error microphone. The n FB filters are configured to process the kth frame noise signal collected by the error microphone based on their respective kth frame filter coefficients to obtain feedback anti-phase noise. Then, the feedforward antiphase noise, the feedback antiphase noise, and the kth frame downlink signal of the first speaker of each noise cancellation channel are mixed, and then the mixed signal is played back through the corresponding first speaker to achieve noise cancellation.

[0208] FIG. 8 is a diagram of another headset structure according to an embodiment of the present application. See FIG. 8. The headset includes one reference microphone, one error microphone, two FF filters, two FF adaptation engines that correspond one-to-one to the two FF filters, two FB filters, two FB adaptation engines that correspond one-to-one to the two FB filters, two first speakers (i.e., Speaker 1 and Speaker 2), a downlink compensation filter, a downlink compensation adaptation engine (not shown), and two EQ calibrators. Both of the two FF filters correspond to the reference microphone, and the two first speakers are formed by dual-diaphragm (also called dual dynamic) loudspeakers. In this case, the headset may not include a digital divider. In addition, the two first speakers may be considered as a combination of two dynamic loudspeakers, but they share a magnetic circuit and are physically considered as one speaker. Since the two dynamic loudspeakers both have good full-band sound production capabilities, they can be thought of as two full-band ANC noise cancellation modules stacked together, thus making full use of the noise cancellation capabilities of the speakers.

[0209] FIG. 9 is a diagram of another headset structure according to an embodiment of the present application. Please refer to FIG. 9. The headset includes one reference microphone, one error microphone, two FF filters, two FF adaptation engines that correspond one-to-one to the two FF filters, two FB filters, two FB adaptation engines that correspond one-to-one to the two FB filters, two first speakers (i.e., Speaker 1 and Speaker 2), a downlink compensation filter, a downlink compensation adaptation engine (not shown), a digital divider, and two EQ calibrators. The difference from FIG. 8 is that there are two physical entities of the first speakers, and the two divided first speakers may be different or the same. Although the two first speakers are not completely the same, from an independent perspective, each first speaker can be designed as a full-band noise cancellation unit. In this way, the maximum noise cancellation capability of each speaker is fully utilized.

[0210] FIG. 10 is a diagram of another headset structure according to an embodiment of the present application. Please refer to FIG. 10. The headset includes two reference microphones, one error microphone, two FF filters, two FF adaptation engines that correspond one-to-one to the two FF filters, two FB filters, two FB adaptation engines that correspond one-to-one to the two FB filters, two first speakers (i.e., Speaker 1 and Speaker 2), one second speaker (i.e., Speaker 3), a downlink compensation filter, a downlink compensation adaptation engine (not shown), a digital divider, and three EQ calibrators. In FIG. 10, three speakers are used to meet high-quality sound requirements. The frequency responses of the three speakers focus separately on low, mid, and high frequency bands. Two of the three speakers functioning as first speakers (i.e., Speaker 1 and Speaker 2) are involved in noise cancellation, and the remaining speaker functioning as second speaker (i.e., Speaker 3) is not involved in noise cancellation, but is involved in downlink compensation (i.e., downlink compensation is performed on the downlink signal transmitted by the user terminal, and the downlink signal is a full-band audio signal including an audio signal in the sound-generating frequency band of the second speaker). The first speaker may be a low-band and mid-band speaker, or a full-band speaker, and the second speaker may be a high-band speaker, a mid-band speaker, or a low-band speaker.

[0211] Optionally, current mainstream ANC chips may not acquire signals from the high-band speaker. Therefore, see FIG. 11. The second speaker may not be involved in downlink compensation (i.e., after digital frequency division is performed on the downlink signal transmitted by the user terminal, downlink signals corresponding to the two first speakers are acquired, and downlink compensation is performed on the downlink signals corresponding to the two first speakers, but not on the downlink signal corresponding to the second speaker). In this case, the first speaker may be a low-band and mid-band speaker or a full-band speaker, and the second speaker is a high-band speaker. To reduce the damage caused by ANC to downlink sound quality, the frequency division point of the high-band speaker may be above 6 kHz, i.e., audio signals above 6 kHz are not compensated. Of course, the frequency division point of 6 kHz is not limited in the embodiments of the present application, and another higher frequency division point may be used.

[0212] It should be noted that the FF adaptation engine, FB adaptation engine, and downlink compensation adaptation engine described above may be located on the microcontroller unit. The FF filter, FB filter, and downlink compensation filter may be located on the ANC chip. The microcontroller unit and the ANC chip are sometimes collectively referred to as a noise cancellation processor. The microcontroller unit and the ANC chip may be integrated on one chip or located on two chips.

[0213] 12 is a diagram of the structure of a noise cancellation device according to an embodiment of the present application. The noise cancellation device may be implemented as part or all of a headset by software, hardware, or a combination thereof. The headset may be the headset shown in FIG. 1. See FIG. 12. The device includes a noise cancellation parameter determination module 1201, an anti-phase noise generation module 1202, and a noise cancellation module 1203.

[0214] The noise cancellation parameter determination module 1201 is configured to determine a plurality of groups of target noise cancellation parameters that correspond one-to-one to a plurality of first speakers.

[0215] The anti-phase noise generation module 1202 is configured to generate, based on the multiple groups of target noise cancellation parameters, multiple groups of target anti-phase noises that correspond one-to-one to the multiple first speakers, and the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound generation frequency band of the multiple first speakers.

[0216] The noise cancellation module 1203 is configured to perform noise cancellation via the multiple first speakers by using multiple groups of target anti-phase noises.

[0217] Optionally, the headset further includes a plurality of FF filters corresponding one-to-one to the plurality of first speakers, and the plurality of groups of target noise cancellation parameters include k-th frame filter coefficients of the plurality of FF filters, where k is an integer greater than or equal to 1.

[0218] The noise cancellation parameter determination module 1201 includes: a first FF filter coefficient determination submodule configured to determine initial filter coefficients of the plurality of FF filters as k-th frame filter coefficients of the plurality of FF filters when k is equal to 1, or to determine k-th frame filter coefficients of the plurality of FF filters based on an initial noise canceling level and a mapping relationship between the noise canceling level and the FF filter coefficients; a second FF filter coefficient determination submodule configured to determine the k-th frame filter coefficients of a plurality of FF filters based on a (k-1)-th frame reference signal collected by at least one reference microphone, a (k-1)-th frame error signal collected by an error microphone, and a target noise canceling level, when k is greater than 1;

[0219] Optionally, the second FF filter coefficient determination sub-module is specifically configured to: determining (k-1)-th frame filter coefficients of a plurality of secondary paths (SPs) based on a target noise canceling level and a mapping relationship between the noise canceling level and filter coefficients of the SPs, where the plurality of SPs are paths from a plurality of first speakers to an error microphone; and Determining the k-th frame filter coefficients of a plurality of FF filters based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the (k-1)-th frame filter coefficients of a plurality of SPs.

[0220] Optionally, the second FF filter coefficient determination sub-module is further configured to specifically: Determine a k-th frame filter coefficient of a target FF filter by using one of the plurality of FF filters as a target FF filter based on the following operations until a k-th frame filter coefficient of each FF filter is determined: If the target FF filter is the first FF filter, determining the k-th frame filter coefficient of the target FF filter based on the (k-1)-th frame reference signal collected by the target reference microphone, the (k-1)-th frame error signal collected by the error microphone, and the (k-1)-th frame filter coefficient of the target SP, where the target reference microphone is the reference microphone corresponding to the target FF filter, and the target SP is the path from the first speaker corresponding to the target FF filter to the error microphone; or If the target FF filter is a non-first FF filter, determine the kth frame filter coefficient of the target FF filter based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the multiple SPs, and the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0221] Optionally, the headset further includes a plurality of FB filters in one-to-one correspondence with the plurality of first speakers.

[0222] The second FF filter coefficient determination sub-module is specifically further configured to: Determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)-th frame reference signal collected by at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficients of the plurality of SPs, and the (k-1)-th frame filter coefficients of the plurality of FB filters.

[0223] Optionally, the second FF filter coefficient determination sub-module is further configured to specifically: Determine a k-th frame filter coefficient of a target FF filter by using one of the plurality of FF filters as a target FF filter based on the following operations until a k-th frame filter coefficient of each FF filter is determined: When the target FF filter is a first FF filter, determining the k-th frame filter coefficient of the target FF filter based on the (k-1)-th frame reference signal collected by the target reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficients of the plurality of SPs, and the (k-1)-th frame filter coefficients of the plurality of FB filters, where the target reference microphone is the reference microphone corresponding to the target FF filter; or When the target FF filter is a non-first FF filter, determine the kth frame filter coefficient of the target FF filter based on the (k-1)th frame reference signal collected by the target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the multiple SPs, the (k-1)th frame filter coefficients of the multiple FB filters, and the kth frame frequency response information and the (k-1)th frame frequency response information of each FF filter before the target FF filter.

[0224] Optionally, the headset further includes a plurality of feedback (FB) filters corresponding one-to-one to the plurality of first speakers, and the plurality of groups of target noise cancellation parameters further include a kth frame filter coefficient of the plurality of FB filters, where k is an integer greater than or equal to 1.

[0225] The noise cancellation parameter determination module includes: a first FB filter coefficient determination submodule configured to determine initial filter coefficients of the plurality of FB filters as k-th frame filter coefficients of the plurality of FB filters when k is equal to 1, or to determine the k-th frame filter coefficients of the plurality of FB filters based on an initial noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficients; a second FB filter coefficient determination submodule configured to determine a k-th frame filter coefficient of the plurality of FB filters based on a target noise canceling level, when k is greater than 1;

[0226] Optionally, the second FB filter coefficient determination sub-module is specifically configured to: Determine a k-th frame filter coefficient of a target FB filter by using one of the plurality of FB filters as a target FB filter based on the following operations until a k-th frame filter coefficient of each FB filter is determined: Determining a k-th frame filter coefficient of the target FB filter based on a target noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficient; or If the target FB filter is a first type FB filter, determine the k-th frame filter coefficient of the target FB filter based on a target noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficient; or if the target FB filter is a second type FB filter, determine the k-th frame filter coefficient of the target FB filter based on a (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level.

[0227] Optionally, the second FB filter coefficient determination sub-module is further specifically configured to: Determine a (k-1)-th frame filter coefficient of a secondary path (SP) based on a target noise canceling level and a mapping relationship between the noise canceling level and a filter coefficient of the secondary path (SP), where the target SP is a path from the first speaker to the error microphone corresponding to the target FB filter; and determining the k-th frame filter coefficient of the target FB filter based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the (k-1)-th frame filter coefficient of the target SP;

[0228] Optionally, a sound-producing frequency band of the first speaker corresponding to the first type of FB filter is higher than a sound-producing frequency band of the first speaker corresponding to the second type of FB filter.

[0229] Optionally, the apparatus includes: a first noise canceling level determination module configured to determine a (k-1)th frame noise canceling level; a noise canceling level acquisition module configured to acquire noise canceling levels in m frames before the (k-1)th frame, where m is equal to or greater than 1 and less than k-1; and a second noise canceling level determination module configured to determine a target noise canceling level based on the (k-1)th frame noise canceling level and the noise canceling levels in the m frames;

[0230] Optionally, the first noise canceling level determination module is specifically configured to: When there is no valid downlink signal in the (k-1)th frame and the environment is not quiet, determine a (k-1)th frame noise canceling level according to reference filter coefficients of a plurality of FF filters and a mapping relationship between the noise canceling level and frequency response information of the FF filters, wherein: When k is equal to 2, the reference filter coefficients are the initial filter coefficients of the corresponding FF filter, or when k is greater than 2, the reference filter coefficients are the filter coefficients of the corresponding FF filter that last satisfied the convergence stability condition before the kth frame, or are the (k-1)th frame filter coefficients of the corresponding FF filter.

[0231] Optionally, the first noise canceling level determination module is specifically configured to: determining reference frequency response information of the plurality of FF filters based on reference filter coefficients of the plurality of FF filters; determining noise canceling levels that match the reference frequency response information of the plurality of FF filters based on a mapping relationship between the noise canceling levels and the frequency response information of the FF filters to obtain a plurality of reference noise canceling levels; Determining a (k-1)th frame noise canceling level based on a plurality of reference noise canceling levels.

[0232] Optionally, the first noise canceling level determination module is further specifically configured to: determining a (k-1)th frame noise canceling level based on an average value of a plurality of reference noise canceling levels; or Determining a (k-1)th frame noise canceling level based on a reference noise canceling level having the maximum amount among a plurality of reference noise canceling levels.

[0233] Optionally, the first noise canceling level determination module is specifically configured to: If a valid downlink signal exists in the (k-1)th frame, determine a (k-1)th frame noise canceling level based on the valid downlink signal in the (k-1)th frame, the (k-1)th frame reference signal collected by at least one reference microphone, and the (k-1)th frame error signal collected by an error microphone.

[0234] Optionally, the filter coefficients of each FF filter include at least one biquad filter coefficient and one gain.

[0235] In conclusion, in the embodiment of the present application, the multiple groups of target anti-phase noises correspond one-to-one to the multiple first speakers, and the frequency band of each target anti-phase noise in the multiple groups of target anti-phase noises covers the sound-generating frequency band of the multiple first speakers. In other words, the target anti-phase noise is full-band anti-phase noise. Therefore, regardless of whether the first speaker is a high-band speaker, a low-band speaker, or a full-band speaker, when the target anti-phase noise is used to perform noise cancellation, the noise cancellation capability of each first speaker can be fully utilized. In other words, in a headset architecture including multiple noise cancellation channels and multiple speakers, this solution can improve the noise cancellation effect of the headset by using full-band anti-phase noise of the multiple noise cancellation channels.

[0236] It should be noted that the noise cancellation performed by the noise cancellation device provided in the embodiments is divided into functional modules only as an example for explanation. In actual applications, functions may be allocated to different functional modules and implemented according to requirements. In other words, the internal structure of the device is divided into different functional modules to implement all or part of the above-mentioned functions. In addition, the noise cancellation device provided in the embodiments and the noise cancellation method embodiments belong to the same concept. For specific implementation processes of the noise cancellation device, please refer to the method embodiments. Details will not be described again in this specification.

[0237] See Figure 13. Figure 13 is a diagram of another headset structure according to an embodiment of the present application. The headset includes one or more processors 1301, a communication bus 1302, a memory 1303, and one or more communication interfaces 1304.

[0238] The processor 1301 is a general-purpose central processing unit. (C PU), Network Processor (N P), a microprocessor, or one or more integrated circuits configured to implement the solutions of the present application, such as application specific integrated circuits. (A SIC), Programmable Logic Device (P Optionally, the PLD is a complex programmable logic device. (C PLD), Field Programmable Gate Array (F PGA), Generic Array Logic (G AL), or any combination thereof.

[0239] The communication bus 1302 is configured to convey information between the aforementioned components. Optionally, the communication bus 1302 may be categorized as an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used to represent a bus in the figure, but this does not imply that there is only one bus or only one type of bus.

[0240] Optionally, the memory 1303 is a read-only memory (R OM), Random Access Memory (R AM), Electrically Erasable Programmable Read-Only Memory (E EPROM), optical disk (compact disk read-only memory (C The memory 1303 may be a medium that can be used to carry or store program code in the form of instructions or data structures, but is not limited to a medium such as, but not limited to, a hard disk drive (e.g., a hard disk drive (HD-ROM), a compact disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, another magnetic storage device, or any other medium that can be used to carry or store program code in the form of instructions or data structures and that can be accessed by a computer. The memory 1303 may exist independently and be connected to the processor 1301 via a communications bus 1302, or the memory 1303 may be integrated with the processor 1301.

[0241] The communication interface 1304 is configured to communicate with another device or a communication network by using any transceiver type device. The communication interface 1304 may include a wired communication interface or, optionally, a wireless communication interface. The wired communication interface is, for example, an Ethernet interface. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is, for example, a wireless local area network (WLAN) interface. (W LAN interfaces, cellular network communication interfaces, and combinations thereof.

[0242] In some embodiments, the memory 1303 is configured to store program code 1305 for implementing the solution of the present application. The processor 1301 can execute the program code 1305 stored in the memory 1303. The program code includes one or more software modules, and the headset can implement the noise cancellation method provided in the embodiment of FIG. 2 through the processor 1301 and the program code 1305 in the memory 1303.

[0243] All or part of the foregoing embodiments may be implemented by using software, hardware, firmware, or any combination thereof. When software is used to implement the foregoing embodiments, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the procedures or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, or digital subscriber line). (D SL)) or wirelessly (e.g., infrared, radio, or microwave). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. Available media may include magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., digital versatile disks, (DVD), semiconductor media (e.g., solid-state disks (S SD)) etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, that is, a non-transitory storage medium.

[0244] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, which, when executed by a processor, performs the steps of the aforementioned method.

[0245] An embodiment of the present application further provides a computer program product, which stores computer instructions that, when executed by a processor, perform the steps of the aforementioned method.

[0246] It should be understood that "at least one" referred to in this specification indicates one or more, and "multiple" indicates two or more. In the description of the embodiments of the present application, " / " indicates "or" unless otherwise specified. For example, A / B may indicate A or B. In this specification, "and / or" only describes an association relationship between related objects and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: when only A exists, when both A and B exist, and when only B exists. In addition, to clearly describe the technical solutions in the embodiments of the present application, terms such as "first" and "second" are used in the embodiments of the present application to distinguish between identical or similar items that basically provide the same function or purpose. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and terms such as "first" and "second" do not indicate clear distinctions.

[0247] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals in the embodiments of the present application are used with the authorization of the user or full authorization of all parties, and the capture, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0248] The foregoing description is only an embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, or improvement made without departing from the spirit and principle of the present application should fall within the protection scope of the present application.

Claims

1. 1. A noise cancellation method applied to a headset, the headset comprising at least one reference microphone, one error microphone, and a plurality of first speakers, the method comprising: determining a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; generating a plurality of groups of target anti-phase noises corresponding one-to-one to the plurality of first speakers based on the plurality of groups of target noise cancellation parameters, wherein a frequency band of each target anti-phase noise in the plurality of groups of target anti-phase noises covers a sound generation frequency band of the plurality of first speakers; performing noise cancellation via the plurality of first speakers by using the plurality of groups of target anti-phase noises; A method comprising:

2. the headset further comprises a plurality of feedforward (FF) filters corresponding one-to-one to the plurality of first speakers, and the plurality of groups of target noise cancellation parameters include k-th frame filter coefficients of the plurality of FF filters, where k is an integer greater than or equal to 1; determining a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; When k is equal to 1, determining initial filter coefficients of the plurality of FF filters as the k-th frame filter coefficients of the plurality of FF filters, or determining the k-th frame filter coefficients of the plurality of FF filters based on an initial noise canceling level and a mapping relationship between the noise canceling level and the FF filter coefficients; or and determining the k-th frame filter coefficients of the plurality of FF filters based on a (k-1)-th frame reference signal collected by the at least one reference microphone, a (k-1)-th frame error signal collected by the error microphone, and a target noise canceling level, when k is greater than 1. The method of claim 1 , comprising:

3. The step of determining the k-th frame filter coefficients of the plurality of FF filters based on a (k-1)-th frame reference signal collected by the at least one reference microphone, a (k-1)-th frame error signal collected by the error microphone, and a target noise canceling level includes: determining a (k-1)-th frame filter coefficient of a plurality of secondary paths (SPs) based on the target noise canceling level and a mapping relationship between the noise canceling level and filter coefficients of the SPs, the plurality of SPs being paths from the plurality of first speakers to the error microphone; determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)th frame reference signal collected by the at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficients of the plurality of SPs; The method of claim 2 , comprising:

4. The step of determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)th frame reference signal collected by the at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficients of the plurality of SPs includes: Until the k-th frame filter coefficients of each FF filter are determined, If the target FF filter is a first FF filter, determining the k-th frame filter coefficient of the target FF filter based on a (k-1)-th frame reference signal collected by a target reference microphone, the (k-1)-th frame error signal collected by the error microphone, and a (k-1)-th frame filter coefficient of a target SP, wherein the target reference microphone is a reference microphone corresponding to the target FF filter, and the target SP is a path from a first speaker corresponding to the target FF filter to the error microphone; or If the target FF filter is a non-first FF filter, determining the k-th frame filter coefficient of the target FF filter based on a (k-1)-th frame reference signal collected by a target reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficients of the plurality of SPs, and the k-th frame frequency response information and the (k-1)-th frame frequency response information of each FF filter before the target FF filter. determining a k-th frame filter coefficient of the target FF filter by using one of the plurality of FF filters as the target FF filter based on the operation The method of claim 3, comprising:

5. the headset further comprises a plurality of feedback (FB) filters corresponding one-to-one to the plurality of first speakers; The step of determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)th frame reference signal collected by the at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, and the (k-1)th frame filter coefficients of the plurality of SPs includes: determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)th frame reference signal collected by the at least one reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, and the (k-1)th frame filter coefficients of the plurality of FB filters; The method of claim 3, comprising:

6. The step of determining the k-th frame filter coefficients of the plurality of FF filters based on the (k-1)-th frame reference signal collected by the at least one reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficients of the plurality of SPs, and the (k-1)-th frame filter coefficients of the plurality of FB filters includes: Until the k-th frame filter coefficients of each FF filter are determined, If the target FF filter is a first FF filter, determining the k-th frame filter coefficient of the target FF filter based on the (k-1)th frame reference signal collected by a target reference microphone, the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficients of the plurality of SPs, and the (k-1)th frame filter coefficients of the plurality of FB filters, wherein the target reference microphone is a reference microphone corresponding to the target FF filter; or If the target FF filter is a non-first FF filter, determining the k-th frame filter coefficient of the target FF filter based on a (k-1)-th frame reference signal collected by a target reference microphone, the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficients of the plurality of SPs, the (k-1)-th frame filter coefficients of the plurality of FB filters, and the k-th frame frequency response information and the (k-1)-th frame frequency response information of each FF filter before the target FF filter. determining a k-th frame filter coefficient of the target FF filter by using one of the plurality of FF filters as the target FF filter based on the operation The method of claim 5 , comprising:

7. the headset further comprises a plurality of feedback (FB) filters corresponding one-to-one to the plurality of first speakers, and the plurality of groups of target noise cancellation parameters further include a k-th frame filter coefficient of the plurality of FB filters, where k is an integer greater than or equal to 1; determining a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; When k is equal to 1, determining initial filter coefficients of the plurality of FB filters as the k-th frame filter coefficients of the plurality of FB filters, or determining the k-th frame filter coefficients of the plurality of FB filters based on the initial noise canceling level and a mapping relationship between the noise canceling level and the FB filter coefficients; or determining the k-th frame filter coefficient of the plurality of FB filters based on the target noise canceling level, if k is greater than 1; 7. The method of claim 1, comprising:

8. The step of determining the k-th frame filter coefficient of the plurality of FB filters based on the target noise canceling level includes: Until the k-th frame filter coefficients of each FB filter are determined, determining the k-th frame filter coefficient of a target FB filter based on the target noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficient; or If the target FB filter is a first type FB filter, determining the k-th frame filter coefficient of the target FB filter based on the target noise canceling level and the mapping relationship between the noise canceling level and the FB filter coefficient; or if the target FB filter is a second type FB filter, determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the target noise canceling level. determining a k-th frame filter coefficient of the target FB filter by using one of the plurality of FB filters as the target FB filter based on the operation The method of claim 7, comprising:

9. The step of determining the k-th frame filter coefficient of the target FB filter based on the (k-1)-th frame error signal collected by the error microphone, the (k-1)-th frame filter coefficient of the target FB filter, and the target noise canceling level includes: determining a (k-1)-th frame filter coefficient of a target secondary path (SP) based on the target noise canceling level and the mapping relationship between the noise canceling level and the filter coefficient of the secondary path (SP), where the target SP is a path from a first speaker corresponding to the target FB filter to the error microphone; determining the k-th frame filter coefficient of the target FB filter based on the (k-1)th frame error signal collected by the error microphone, the (k-1)th frame filter coefficient of the target FB filter, and the (k-1)th frame filter coefficient of the target SP; The method of claim 8, comprising:

10. 10. The method according to claim 8, wherein a sound-producing frequency band of a first speaker corresponding to the first type of FB filter is higher than a sound-producing frequency band of a first speaker corresponding to the second type of FB filter.

11. determining a (k-1)th frame noise canceling level; obtaining noise canceling levels for m frames before the (k-1)th frame, where m is equal to or greater than 1 and less than k-1; determining the target noise canceling level based on the (k-1)th frame noise canceling level and the noise canceling levels in the m frames; 11. The method of claim 2, further comprising:

12. The step of determining a (k-1)th frame noise canceling level includes: If there is no valid downlink signal in the (k-1)th frame and the environment is not quiet, determining a noise canceling level for the (k-1)th frame based on reference filter coefficients of the plurality of FF filters and a mapping relationship between a noise canceling level and frequency response information of the FF filters. Including, When k is equal to 2, the reference filter coefficients are the initial filter coefficients of the corresponding FF filter; or when k is greater than 2, the reference filter coefficients are the filter coefficients of the corresponding FF filter that last satisfy a convergence stability condition before the k-th frame, or are the (k-1)-th frame filter coefficients of the corresponding FF filter; The method of claim 11.

13. determining the (k-1)th frame noise canceling level based on reference filter coefficients of the plurality of FF filters and a mapping relationship between a noise canceling level and frequency response information of the FF filters, determining reference frequency response information of the plurality of FF filters based on the reference filter coefficients of the plurality of FF filters; determining a noise canceling level that matches the reference frequency response information of the plurality of FF filters based on the mapping relationship between the noise canceling level and the frequency response information of the FF filters to obtain a plurality of reference noise canceling levels; determining the (k-1)th frame noise canceling level based on the plurality of reference noise canceling levels; 13. The method of claim 12, comprising:

14. The step of determining the (k-1)th frame noise canceling level based on the plurality of reference noise canceling levels includes: determining the (k-1)th frame noise canceling level based on an average value of the plurality of reference noise canceling levels; or determining the (k-1)th frame noise canceling level based on a reference noise canceling level having a maximum amount among the plurality of reference noise canceling levels; 14. The method of claim 13, comprising:

15. The step of determining a (k-1)th frame noise canceling level includes: If a valid downlink signal exists in the (k-1) frame, determining a noise canceling level for the (k-1) frame based on the valid downlink signal for the (k-1) frame, the (k-1) frame reference signal collected by the at least one reference microphone, and the (k-1) frame error signal collected by the error microphone. The method of claim 11 , comprising:

16. The method of claim 2 , wherein the filter coefficients of each FF filter include at least one biquad filter coefficient and one gain.

17. 1. A headset comprising at least one reference microphone, one error microphone, a plurality of first speakers, and a noise cancellation processor, The noise cancellation processor is configured to perform the steps of the method of any one of claims 1 to 16. Headset.

18. 20. The headset of claim 17, wherein the plurality of first speakers includes two first speakers formed by a single dual-diaphragm loudspeaker, or the plurality of first speakers includes a plurality of speakers with separate loudspeakers.

19. 19. The headset of claim 17 or 18, further comprising at least one second speaker, the at least one second speaker not participating in noise cancellation.

20. 1. A noise cancellation device for use in a headset comprising at least one reference microphone, one error microphone, and a plurality of first speakers, the noise cancellation device comprising: a noise cancellation parameter determination module configured to determine a plurality of groups of target noise cancellation parameters corresponding one-to-one to the plurality of first speakers; an anti-phase noise generation module configured to generate a plurality of groups of target anti-phase noises corresponding one-to-one to the plurality of first speakers based on the plurality of groups of target noise cancellation parameters, wherein a frequency band of each target anti-phase noise in the plurality of groups of target anti-phase noises covers a sound generation frequency band of the plurality of first speakers; a noise cancellation module configured to perform noise cancellation via the plurality of first speakers by using the plurality of groups of target anti-phase noises; An apparatus comprising:

21. 17. A computer-readable storage medium, the storage medium storing a computer program which, when executed by a processor, performs the steps of the method of any one of claims 1 to 16.

22. 17. A computer program product, the computer program product storing computer instructions which, when executed by a processor, cause the steps of the method of any one of claims 1 to 16 to be performed.

Citation Information

Patent Citations

  • Active noise reduction method and device

    CN113676804A

  • Noise reduction earphone, noise reduction method and device, storage medium and processor

    CN115278438A

  • Systems, methods, apparatus, and computer-readable media for adaptive active noise cancellation

    JP2012533091A

  • A Low Latency Multi-Driver Adaptive Noise Cancellation (anc) System for Personal Audio Devices

    JP2016510915A

  • Noise cancellation device and method

    JP2022528713A