Filter coefficient determination method and device, electronic equipment and storage medium
By decomposing the filter into sub-filters and incorporating constraint information, the problem of high complexity in determining filter coefficients is solved, resulting in faster convergence speed and more efficient echo cancellation, thereby improving call quality and speech recognition accuracy.
Patent Information
- Application Number
- CN202410620720.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-18
Smart Images

Figure CN120977326A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech signal processing, and in particular to a filter coefficient determination method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the continuous development of technology, audio and video communication and intelligent voice interaction are becoming more and more common. There is a problem of echo in these scenarios. The echo is caused by the fact that the sound played by a playing device is collected again by a voice collecting device after the impact response of the space. If echo cancellation is not performed, the voice signal transmitted to the far end or the voice recognition system will contain a large amount of echo interference, which greatly affects the call quality and voice recognition accuracy.
[0003] In related technologies, echo cancellation is performed by a filter, and the method for determining the filter coefficient is complex and slow. SUMMARY
[0004] The present application aims to at least partly solve one of the technical problems in the related art.
[0005] To this end, the present application proposes a filter coefficient determination method, device, electronic device, and storage medium. The target filter is decomposed into multiple sub-filters, the coefficients of the target filter are determined according to the coefficients of the sub-filters, and at the same time, constraint information is added when estimating the coefficients of the sub-filters, thereby reducing the complexity of filter coefficient determination and improving the convergence speed.
[0006] An embodiment of the present application proposes a filter coefficient determination method, comprising:
[0007] obtaining filter coefficients of multiple sub-filters at a historical time before a target time; wherein the multiple sub-filters are obtained by decomposing a target filter;
[0008] For each sub-filter, determining constraint information of the filter coefficients of the sub-filter according to the filter coefficients of the sub-filter;
[0009] determining error information of echo cancellation of the sub-filter on an input voice signal at the historical time according to the constraint information of the filter coefficients of the sub-filter at the historical time;
[0010] determining the filter coefficients of the sub-filter at the target time according to the error information of echo cancellation of the sub-filter at the historical time and the filter coefficients of the sub-filter at the historical time;
[0011] determining target filter coefficients of the target filter at the target time according to the filter coefficients of the multiple sub-filters at the target time.
[0012] Another aspect of the present application provides a device for determining filter coefficients, comprising:
[0013] a obtaining module, configured to obtain filter coefficients of a plurality of sub-filters at a historical time point before a target time point, wherein the plurality of sub-filters are obtained by decomposing a target filter;
[0014] a first determining module, configured to determine constraint information of the filter coefficients of each sub-filter according to the filter coefficients of the sub-filter;
[0015] a second determining module, configured to determine error information of the sub-filter at the historical time point for performing echo cancellation on an input speech signal according to the constraint information of the filter coefficients of the sub-filter at the historical time point;
[0016] a third determining module, configured to determine the filter coefficients of the sub-filter at the target time point according to the error information of the sub-filter at the historical time point for performing echo cancellation and the filter coefficients of the sub-filter at the historical time point;
[0017] a fourth determining module, configured to determine target filter coefficients of the target filter at the target time point according to the filter coefficients of the plurality of sub-filters at the target time point.
[0018] Another aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method of the foregoing aspect.
[0019] Another aspect of the present application provides a non-transitory computer readable storage medium, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the method of the foregoing aspect.
[0020] Another aspect of the present application provides a computer program product, having a computer program stored thereon, wherein the program is executable by a processor to implement the method of the foregoing aspect.
[0021] The method, device, electronic equipment and storage medium for determining filter coefficients provided in the application obtain filter coefficients of a plurality of sub-filters at a historical time before a target time, wherein the plurality of sub-filters are obtained by decomposing a target filter, for each sub-filter, determine constraint information of the filter coefficients of the sub-filter according to the filter coefficients of the sub-filter, determine error information of the sub-filter at the historical time for performing echo cancellation on an input speech signal according to the constraint information of the filter coefficients of the sub-filter at the historical time, determine filter coefficients of the sub-filter at the target time according to the error information of the sub-filter at the historical time for performing echo cancellation and the filter coefficients of the sub-filter at the historical time, determine target filter coefficients of the target filter at the target time according to the filter coefficients of the plurality of sub-filters at the target time, and determine the filter coefficients of the target filter by decomposing the target filter into the plurality of sub-filters, determine the filter coefficients of the target filter according to the filter coefficients of the sub-filters, and add constraint information when estimating the filter coefficients of the sub-filters, thereby reducing the complexity of determining the filter coefficients and improving the convergence speed.
[0022] Additional aspects and advantages of the application will be made apparent by the following description and the appended claims. BRIEF DESCRIPTION OF DRAWINGS
[0023] The above and / or additional aspects and advantages of the application will become apparent and be made clear to those skilled in the art from the following description and the appended claims, taken in conjunction with the accompanying drawings.
[0024] Figure 1 A flowchart of a method for determining filter coefficients provided by an embodiment of the application;
[0025] Figure 2 A schematic diagram of an echo cancellation scenario provided by an embodiment of the application;
[0026] Figure 3 A flowchart of another method for determining filter coefficients provided by an embodiment of the application;
[0027] Figure 4 A filter coefficient speed iteration diagram provided by an embodiment of the application;
[0028] Figure 5 A structural diagram of a device for determining filter coefficients provided by an embodiment of the application;
[0029] Figure 6 A structural diagram of an electronic device provided by an embodiment of the application. DETAILED DESCRIPTION
[0030] Embodiments of the present application are described below in detail with reference to the accompanying drawings, examples of which are shown in the drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.
[0031] The method for determining filter coefficients, the device for determining filter coefficients, the electronic device, and the storage medium of the embodiments of the present application are described below with reference to the accompanying drawings.
[0032] Figure 1 A flowchart of the method for determining filter coefficients provided by the embodiments of the present application.
[0033] The embodiments of the present application are exemplified by the method for determining filter coefficients configured in the device for determining filter coefficients, which can be applied to any electronic device, so that the electronic device can perform command processing functions.
[0034] The electronic device can be any device with computing capability, such as a personal computer, a mobile terminal, a server (or cloud), etc. The mobile terminal can be a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. hardware devices with various operating systems, touch screens, and / or display screens.
[0035] As shown in FIG. 1, the method can include the following steps: Figure 1
[0036] Step 101, obtaining filter coefficients of a plurality of sub-filters at a historical time before a target time.
[0037] The target time can be considered as any processing time in a speech scene, such as the current time. The historical time is a time before the target time, which can be one or more.
[0038] The plurality of sub-filters are obtained by decomposing a target filter. The target filter is an adaptive filter applied in an echo cancellation scene, used to estimate an impulse response signal of a speech signal propagation in the scene according to the filter coefficients, to output an estimated echo signal, and then to perform echo cancellation on the collected speech signal according to the echo signal. The adaptive filter can be a least mean squares (LMS) filter, a normalized least mean squares (NLMS) filter, a recursive least squares (RLS) filter, a Kalman filter, etc.
[0039] As an example,Figure 2 An echo cancellation scenario provided in an embodiment of the present application is shown in a schematic diagram as Figure 2 shown, a user A in a room is making an audio-video call or a smart voice interaction, a far-end voice signal played by a loudspeaker is r(n), for example, in an audio call scenario, r(n) is a voice signal of a user B who is making a call with user A, after being affected by an actual room impulse response w(n), the echo signal y(n) is obtained by a microphone, meanwhile, the actual signal x(n) collected by the microphone also contains a voice s(n) of a near-end speaker user A and a background noise v(n). An adaptive filtering algorithm is to estimate filter coefficients w'(n) by a target filter to determine an estimated room impulse response, so as to estimate an echo signal y'(n) in the microphone according to the filter coefficients w'(n) and the voice signal r(n), and then subtract the estimated echo y'(n) in the microphone from the actual signal x(n) collected by the microphone to achieve echo cancellation, therefore, efficient estimation of the filter coefficients w'(n) is required for echo cancellation, wherein n is a processing time. However, in an actual scenario, the coefficients of the target filter are relatively long, which greatly increases the calculation complexity, therefore, in an embodiment of the present application, the target filter can be decomposed into multiple sub-filters based on requirements, for example, into two sub-filters, as an implementation manner, a first length of the coefficients of the target filter is obtained, the coefficients of the first length of the target filter are decomposed into a sub-filter with a second length of coefficients and a sub-filter with a third length of coefficients, wherein the product of the second length and the third length is the first length. For example, the length of the coefficients of the target filter is L=1024, the target filter is decomposed into two sub-filters, which are referred to as filter g and filter h, wherein the length of the coefficients of filter g is 32, the length of the coefficients of filter h is 32, and 32*32=1024, that is, the product of the length of the coefficients of filter g and the length of the coefficients of filter h is the length of the coefficients of the target filter, so as to decompose the target filter with a relatively long length into two sub-filters with relatively short lengths to reduce the subsequent calculation complexity.
[0040] It needs to be understood that the coefficients of the filter are used to simulate the impulse response of the corresponding space to obtain an estimated impulse response.
[0041] In step 102, for each sub-filter, constraint information of the filter coefficients of the sub-filter is determined according to the filter coefficients of the sub-filter.
[0042] In an embodiment of the present application, for each sub-filter, the filter coefficients of the sub-filter are converted into a vector, the number of non-zero elements in the vector is taken as a zero norm, and the zero norm corresponding to the filter coefficients of the sub-filter is taken as the constraint information of the filter coefficients of the sub-filter.
[0043] Step 103, according to the constraint information of the filter coefficients of the historical time sub-filter, determine the error information of the echo cancellation of the input speech signal by the historical time sub-filter.
[0044] In the embodiment of the application, since the sparsity of the target filter, at least one of the decomposed sub-filters is also sparse, and the convergence speed of the sub-filter will be affected. In order to improve the calculation speed of each coefficient update, the constraint information is used to improve the operation speed. As an implementation manner, a cost function of the sub-filter can be determined, and the cost function is used to indicate the error information of the echo cancellation of the input speech signal by the sub-filter, that is, the cost function is used to indicate the size of the error, and then the constraint information is added to the cost function to obtain a cost function containing the constraint information. The cost function containing the constraint information is used as the error signal of the echo cancellation. Specifically, for each sub-filter, the estimated echo signal is determined according to the filter coefficients of the historical time sub-filter and the input speech signal of the historical time sub-filter, and the error signal of the echo cancellation of the sub-filter is determined according to the original speech signal collected by the microphone in the historical time scene, the estimated echo signal, and the constraint information corresponding to the sub-filter.
[0045] It should be understood that each sub-filter is obtained by decomposing a target filter, and each sub-filter is not a plurality of sub-filters in the entity, but is processed as a plurality of virtual sub-filters for coefficient determination in data processing when estimating the filter coefficients. Therefore, the input speech signal of each sub-filter is determined according to the original speech signal input to each sub-filter obtained by decomposition. As an implementation manner, the input speech signal of each sub-filter is determined according to the filter coefficients of each sub-filter and the original speech signal. Therefore, the input speech signal of each sub-filter can be different speech signals.
[0046] The cost function is, for example, a cost function based on Recursive Least Squares (RLS) according to the characteristics of the impulse response of the real space, and the filter coefficients have many coefficients close to 0. The constraint information of the zero norm (0 norm) constraint is added on the basis of the traditional RLS cost function, which can accelerate the convergence speed of the filter and improve the efficiency.
[0047] Step 104, according to the error information of the echo cancellation of the sub-filter at the historical time before the target time and the filter coefficients of the historical time sub-filter, determine the filter coefficients of the sub-filter at the target time.
[0048] In one implementation manner of the embodiment of the present application, the error information of the historical time sub-filter performing echo cancellation is determined, the error information is represented by a cost function of the historical time sub-filter, the gain value of the coefficient of the historical time sub-filter is obtained by solving the cost function of the historical time sub-filter, and the filter coefficient of the target time sub-filter is determined according to the gain value of the coefficient of the historical time sub-filter and the filter coefficient of the historical time sub-filter. Since the coefficient constraint is added in the cost function, and the length of the coefficient of the sub-filter obtained by decomposing the target filter is reduced, the iteration speed of the filter coefficient is improved.
[0049] In step 105, the filter coefficient of the target time target filter is determined according to the filter coefficients of the target time multiple sub-filters.
[0050] In one implementation manner of the embodiment of the present application, the Kronecker product of the filter coefficients of the target time multiple sub-filters is taken as the target filter coefficient of the target time target filter. Since the operation complexity and the iteration speed of the determination process of each sub-filter coefficient are improved, the filter coefficients of the target time multiple sub-filters are converted into corresponding vectors, the Kronecker product of the filter coefficients of the multiple sub-filters in the vector form is solved, the result of the Kronecker product is taken as the target filter coefficient of the target time target filter, and the calculation complexity of the target filter coefficient of the target filter is reduced and the iteration speed is improved.
[0051] The filter coefficient determination method of the embodiment of the present application obtains the filter coefficients of the historical time multiple sub-filters at the target time, wherein the multiple sub-filters are obtained by decomposing the target filter. For each sub-filter, the constraint information of the filter coefficient of the sub-filter is determined according to the filter coefficient of the sub-filter. The error information of the historical time sub-filter performing echo cancellation on the input speech signal is determined according to the constraint information of the filter coefficient of the historical time sub-filter. The filter coefficient of the target time sub-filter is determined according to the error information of the historical time sub-filter performing echo cancellation and the filter coefficient of the historical time sub-filter. The target filter coefficient of the target time target filter is determined according to the filter coefficients of the target time multiple sub-filters. By decomposing the target filter into multiple sub-filters, the coefficient of the target filter is determined according to the coefficient of the sub-filter, and the constraint information is added when estimating the coefficient of the sub-filter. The complexity of the filter coefficient determination is reduced, and the convergence speed is improved.
[0052] Based on the above embodiment, Figure 3 Another flowchart of a filter coefficient determination method provided by the embodiment of the present application is shown in FIG. 2. Figure 3 The method comprises the following steps:
[0053] In step 301, filter coefficients of the multiple sub-filters at the historical time before the target time are obtained.
[0054] In step 301, filter coefficients of the multiple sub-filters at the historical time before the target time are obtained.
[0055] In the embodiments of the present application, the target filter is taken as an example to be decomposed into two sub-filters, and the two sub-filters are referred to as sub-filters g and h for the sake of distinction.
[0056] The filter coefficients w(n) of the target filter have a length L, the length L is greater than a set threshold, and the length L satisfies sparsity, so that w(n) can be decomposed into:
[0057]
[0058] wherein g(n) and h(n) are sub-filters with lengths of L1 and L2 respectively, and the symbol is a Kronecker product, which is defined as follows:
[0059]
[0060] wherein L1×L2=L, and the Kronecker product has the following properties:
[0061]
[0062] wherein I L1 is a unit matrix with a size of L1×L1; I L2 is a unit matrix with a size of L2×L2. Generally, the diagonal elements of the unit matrix are 1, and the other positions are 0.
[0063] In step 302, for each sub-filter, constraint information of the filter coefficients of the sub-filter is determined according to the filter coefficients of the sub-filter, and error information of the sub-filter at the historical time for performing echo cancellation on the input speech signal is determined according to the constraint information of the filter coefficients of the sub-filter at the historical time.
[0064] The related explanations in the foregoing embodiments are also applicable to the present embodiment, and will not be repeated here.
[0065] In an implementation manner of the embodiments of the present application, the constraint information is a zero norm constraint, and the error information is indicated by a cost function, that is, the cost function of the sub-filter at the historical time is determined, that is, the error information of the sub-filter at the historical time for performing echo cancellation on the input speech signal is determined.
[0066] Wherein, for the sub-filter g, the error information of the sub-filter g performing echo cancellation on the input speech signal, i.e. the cost function, is determined as follows:
[0067]
[0068] Wherein, ||g(n)||0 is the zero norm of the filter coefficient g(n) of the sub-filter g in vector form, i.e. the number of non-zero elements in the filter coefficient g(n) in vector form, 0<γ<1 is a constant used to measure the importance of the sparsity constraint. Wherein, n is the time, K is the length of the observation data r h (n), for example, the historical time is 4, then K is 4, x(n) is the speech signal collected by the audio collection device, the audio collection device is, for example, a microphone; r h (n) is the input speech signal of the sub-filter h, which is determined according to the speech signal played by the device or the speech signal issued by the remote user, and is specifically as follows:
[0069]
[0070] Wherein, r T (n) is the transpose matrix of the speech signal played by the device or the speech signal issued by the remote user, and h(n) is the coefficient of the sub-filter h.
[0071] Step 303, the output speech signal of the historical time sub-filter is obtained.
[0072] In the embodiment of the application, the speech signal output by the historical time sub-filter is determined according to the product of the filter coefficient of the historical time sub-filter and the input speech signal at the historical time, which is the echo signal in the audio collection device estimated at the historical time, and the audio collection device is, for example, a microphone.
[0073] Step 304, the gain value of the historical time sub-filter is determined according to the error information of the historical time sub-filter performing echo cancellation.
[0074] In one implementation mode of the embodiment of the application, the cost function ξ{g(n)} of the historical time sub-filter g is derived with respect to the filter coefficient g(n) at the historical time, and the gradient is equal to 0, which can obtain:
[0075]
[0076] Wherein, is the zero norm of the vector g(n).
[0077] The process of solving the cost function of the above-mentioned sparse RLS sub-filter g can obtain:
[0078] Wherein, κ g(n) is the gain value of the sub-filter g, i.e. the Kalman gain, defined as follows:
[0079]
[0080] wherein, r g (n) is the speech signal input to the sub-filter g, which can be the speech signal played by the device or the speech signal uttered by the remote user. g -1 (n) is the speech signal input to the sub-filter g g (n) is the covariance matrix of the speech signal input to the sub-filter g
[0081] At step 305, the filter coefficient of the sub-filter g at the target time is determined according to the gain value of the sub-filter g at the historical time, the filter coefficient of the sub-filter g at the historical time and the output speech signal of the sub-filter g at the historical time.
[0082]
[0083] wherein g(n+1) is the filter coefficient of the sub-filter g at the target time, e g (n) is the difference signal determined according to the difference between the speech signal collected by the microphone at the historical time and the speech signal output by the sub-filter g, It is difficult to solve g(n), and an approximate expression of ||g(n)||0 can be used to solve g(n):
[0084]
[0085] wherein g i (n) is the i-th coefficient of g(n), and β is a zero attraction constant, which affects the strength and range of the zero attraction,
[0086] The derivative of g(n) with respect to g(n) is:
[0087]
[0088] wherein sgn(·) is a sign function. Then the filter coefficient g(n+1) of the sub-filter g at the target time is updated according to the filter coefficient g(n) at the historical time.
[0089] The formula of the filter coefficient g(n+1) of the sub-filter g is:
[0090]
[0091] wherein g L1-1 (n) is the L1-th coefficient in g(n);
[0092] Similarly, for the sub-filter h, the error information of the sub-filter h performing echo cancellation on the input speech signal is determined, that is, the cost function is as follows:
[0093]
[0094] wherein ||h(n)||0 is the zero norm of the filter coefficient h(n) of the sub-filter h in the form of a vector, that is, the number of non-zero elements in the vector.
[0095] wherein r g (n) is the input speech signal of the sub-filter g, which is determined according to the speech signal played by the device or the speech signal emitted by the user at the far end, and is specifically as follows:
[0096]
[0097] wherein r T (n) is the transposed matrix of the speech signal played by the device or the speech signal emitted by the user at the far end, and g(n) is the coefficient of the sub-filter g.
[0098] Similarly, the gain value of the sub-filter h, that is, the Kalman gain, can be determined, and the determination method of the gain value of the sub-filter g is the same in principle, which will not be described here. The gain value of the sub-filter h is as follows:
[0099] wherein r h (n) is the input speech signal of the sub-filter h; R h -1 (n) is the covariance matrix of the input speech signal r h (n) corresponding to the sub-filter h.
[0100] Further, the filter coefficient h(n+1) of the sub-filter h at the target time is determined, and the method of the sub-filter g is the same in principle, which will not be described here.
[0101] Step 306: The Kronecker product of the filter coefficients of the plurality of sub-filters at the target time is taken as the target filter coefficient of the target filter at the target time.
[0102] In an implementation manner of the embodiment of the application, the target filter coefficient of the target filter at the target time is By converting the estimation of the target filter into the estimation of two short filters, and adding the zero norm constraint to the sub-filter, the convergence speed of the algorithm is greatly accelerated.
[0103] Step 307: The speech signal to be processed collected by the audio collection device at the target time is acquired.
[0104] At step 308, the echo is cancelled from the to-be-processed voice signal according to the target filter coefficient estimated at the target time, to obtain a target voice signal at the target time.
[0105] In the embodiments of the present application, the echo is cancelled from the to-be-processed voice signal according to the target filter coefficient estimated at the target time, to obtain a voice signal after echo cancellation, and the voice signal after echo cancellation is converted from the frequency domain to the time domain, to obtain a target voice signal at the target time in the time domain, thereby reducing echo interference and improving the interactive experience of the user in the audio-video call and the intelligent voice scene.
[0106] The method for determining the filter coefficient in the embodiments of the present application comprises the following steps: obtaining the filter coefficients of a plurality of sub-filters at a historical time at a target time, wherein the plurality of sub-filters are obtained by decomposing the target filter; determining, for each sub-filter, the constraint information of the filter coefficient of the sub-filter according to the filter coefficient of the sub-filter; determining the error information of the sub-filter at the historical time for performing echo cancellation on the input voice signal according to the constraint information of the filter coefficient of the sub-filter at the historical time; determining the filter coefficient of the sub-filter at the target time according to the error information of the sub-filter at the historical time for performing echo cancellation and the filter coefficient of the sub-filter at the historical time; determining the target filter coefficient of the target filter at the target time according to the filter coefficients of the plurality of sub-filters at the target time; and decomposing the target filter into a plurality of sub-filters, determining the coefficient of the target filter according to the coefficients of the sub-filters, and adding the constraint information when estimating the coefficients of the sub-filters, thereby reducing the complexity of determining the filter coefficient and improving the convergence speed.
[0107] Based on the above embodiments, in the actual audio-video call and intelligent voice scene of the user, the room impulse response reaches the length of thousands, that is, the filter coefficient also reaches the length of thousands. The calculation complexity of the traditional RLS increases rapidly with the increase of the order of the filter coefficient, and the required calculation resource becomes very large, which makes it difficult for the RLS to be widely applied. The target filter is decomposed into two sub-filters by Kronecker product decomposition in the present application, the update iteration of the sub-filter is equivalent to the estimation of the long filter, which can greatly reduce the calculation complexity of the algorithm. For example, when the coefficient length of the target filter is L=1024, the coefficient lengths of the two sub-filters obtained by decomposition are L1=32 and L2=32, the filter estimation with the coefficient length of 1024 is converted into two filter estimations with the coefficient length of 32, and the calculation complexity is reduced.
[0108] Meanwhile, the sparsity of the actual room impulse response is very high, and the convergence of the traditional RLS will be affected in this case. The present application adopts the 0-norm constraint, which can accelerate the convergence speed of the algorithm. Figure 4A filter coefficient speed iteration schematic diagram provided by the embodiment of the present application indicates the error curve of the filter estimated by each adaptive filter and the real room impulse response with the change of iteration times, as shown in Figure 4 KPD-RLS is a filter using Kronecker product decomposition to decompose the target filter into two sub-filters without adding 0 norm constraint in the cost function, wherein KPD indicates that Kronecker product decomposition is used, l0-RLS is a filter without using Kronecker product decomposition by adding 0 norm constraint, and l0-KPD-RLS is the algorithm proposed in the present application, that is, 0 norm constraint is added in the cost function and Kronecker product decomposition is used. At iteration times 4000, the filter coefficients are changed, that is, the room impulse response is changed, and it can be seen that the estimation results of each filter have faster convergence speed and steady-state performance.
[0109] In order to realize the above-mentioned embodiment, the embodiment of the present application further proposes a filter coefficient determination device.
[0110] Figure 5 A structural schematic diagram of a filter coefficient determination device provided by the embodiment of the present application.
[0111] As Figure 5 shown, the device can include:
[0112] The acquisition module 51 is configured to acquire filter coefficients of a plurality of sub-filters at a historical time before a target time; wherein the plurality of sub-filters are obtained by decomposing a target filter.
[0113] The first determination module 52 is configured to determine constraint information of the filter coefficients of each sub-filter according to the filter coefficients of the sub-filter.
[0114] The second determination module 53 is configured to determine error information of echo cancellation of the input speech signal by the sub-filter at the historical time according to the constraint information of the filter coefficients of the sub-filter at the historical time.
[0115] The third determination module 54 is configured to determine the filter coefficients of the sub-filter at the target time according to the error information of echo cancellation by the sub-filter at the historical time and the filter coefficients of the sub-filter at the historical time.
[0116] The fourth determination module 55 is configured to determine target filter coefficients of the target filter at the target time according to the filter coefficients of the plurality of sub-filters at the target time.
[0117] Further, in an implementation form of the embodiment of the application, the third determining module 54 is further configured to:
[0118] obtain an output speech signal of the sub-filter at the historical time;
[0119] determine a gain value of the sub-filter at the historical time according to error information of echo cancellation performed by the sub-filter at the historical time;
[0120] determine the filter coefficient of the sub-filter at the target time according to the gain value of the sub-filter at the historical time, the filter coefficient of the sub-filter at the historical time, and the output speech signal of the sub-filter at the historical time.
[0121] In an implementation form of the embodiment of the application, the fourth determining module 55 is further configured to:
[0122] take the Kronecker product of the filter coefficients of the plurality of sub-filters at the target time as the target filter coefficient of the target filter at the target time.
[0123] In an implementation form of the embodiment of the application, the plurality of sub-filters are two, and the apparatus further comprises:
[0124] a decomposing module configured to obtain a first length of the filter coefficient of the target filter; decompose the filter coefficient of the first length to obtain a sub-filter with a second length of the filter coefficient and a sub-filter with a third length of the filter coefficient; wherein the product of the second length and the third length is the first length.
[0125] In an implementation form of the embodiment of the application, the apparatus further comprises:
[0126] a fifth determining module configured to determine a zero norm corresponding to the filter coefficient of the sub-filter; take the zero norm corresponding to the filter coefficient of the sub-filter as the constraint information of the filter coefficient of the sub-filter.
[0127] In an implementation form of the embodiment of the application, the apparatus further comprises:
[0128] a processing module configured to obtain a to-be-processed speech signal collected by an audio collection apparatus at a target time; perform echo cancellation on the to-be-processed speech signal according to the estimated target filter coefficient at the target time to obtain a target speech signal at the target time.
[0129] It should be noted that the foregoing explanation and description of the method embodiment are also applicable to the apparatus of the embodiment, which will not be described here again.
[0130] The filter coefficient determination apparatus of the embodiment of the present application obtains filter coefficients of a plurality of sub-filters at a historical time before a target time, wherein the plurality of sub-filters are obtained by decomposing a target filter, for each sub-filter, determines constraint information of the filter coefficients of the sub-filter according to the filter coefficients of the sub-filter, determines error information of the sub-filter at the historical time for performing echo cancellation on an input speech signal according to the constraint information of the filter coefficients of the sub-filter at the historical time, determines the filter coefficients of the sub-filter at the target time according to the error information of the sub-filter at the historical time for performing echo cancellation and the filter coefficients of the sub-filter at the historical time, determines target filter coefficients of the target filter at the target time according to the filter coefficients of the plurality of sub-filters at the target time, and determines the filter coefficients of the target filter according to the filter coefficients of the plurality of sub-filters, by decomposing the target filter into the plurality of sub-filters, and adds the constraint information when estimating the filter coefficients of the sub-filters, thereby reducing the complexity of determining the filter coefficients and improving the convergence speed.
[0131] To achieve the above-mentioned embodiments, the present application further provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method as described in the foregoing method embodiments.
[0132] To achieve the above-mentioned embodiments, the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the program is executable by a processor to implement the method as described in the foregoing method embodiments.
[0133] To achieve the above-mentioned embodiments, the present application further provides a computer program product having a computer program stored thereon, wherein the computer program is executable by a processor to implement the method as described in the foregoing method embodiments.
[0134] Figure 6 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a message transceiving device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0135] Reference Figure 6 , the electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0136] The processing component 802 generally controls the overall operations of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete the steps of the methods described above, in whole or in part. Moreover, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0137] The memory 804 is configured to store various types of data to support the operations of the electronic device 800. Examples of these data include instructions to operate any applications or methods on the electronic device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be realized by any type of volatile or non-volatile storage devices, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0138] The power component 806 provides power to the various components of the electronic device 800. The power component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0139] The multimedia component 808 includes a screen to provide an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. The front camera and / or the back camera can receive external multimedia data when the electronic device 800 is in an operating mode, such as a shooting mode or a video mode. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0140] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0141] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0142] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change of position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 814 can further include a light sensor, such as a CMOS or CCD image sensor, for use in an imaging application. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0143] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 4G, or 5G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcasting management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technology.
[0144] In exemplary embodiments, the electronic device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, for performing the above-described methods.
[0145] In exemplary embodiments, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to accomplish the above-described methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0146] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction.
[0147] In addition, the terms "first", "second", etc. are used only for the purpose of description and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly specified.
[0148] Any process or method descriptions or descriptions of the flow diagrams described herein or otherwise described in the specification can be understood as representing the modules, segments, or portions of code that include executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of the present application includes additional implementation in which the functions are performed in different orders, in substantially simultaneous fashion, or in reverse order, depending on the functionality involved, as would be understood by those skilled in the art.
[0149] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered as a sequence of executable instructions stored in a computer readable medium, which can be executed by an instruction execution system, apparatus or device, such as a computer-based system, a processor-based system, or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions, or a combination thereof. For the purposes of this specification, a "computer readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus or device. The computer readable medium can specifically be, but is not limited to, the following: an electronic connection (electronic apparatus) having one or more wires, a portable computer diskette (magnetic apparatus), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disk read-only memory (CDROM). In addition, the computer readable medium can even be paper or other suitable medium upon which the program can be printed, because the program can be electronically obtained, for example, by optically scanning the paper or other medium, then
[0150] It should be understood that portions of the application can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware, and in another embodiment, any of the following technologies, known in the art, or a combination thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0151] Those of ordinary skill in the art can understand that all or part of the steps carried out by the above-mentioned embodiment methods can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium. When the programs are executed, they include one of the steps of the method embodiments or a combination thereof.
[0152] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0153] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A method for determining filter coefficients, characterized in that, include: Obtain the filter coefficients of multiple sub-filters from historical moments prior to the target time; where the multiple sub-filters are obtained by decomposing the target filter; For each sub-filter, the constraint information of the filter coefficients of the sub-filter is determined based on the filter coefficients of the sub-filter; Based on the constraint information of the filter coefficients of the sub-filter at the historical moment, the error information of the sub-filter performing echo cancellation on the input speech signal at the historical moment is determined; Based on the error information of echo cancellation performed by the sub-filter at the historical time and the filter coefficients of the sub-filter at the historical time, the filter coefficients of the sub-filter at the target time are determined; The target filter coefficients of the target filter at the target time are determined based on the filter coefficients of the plurality of sub-filters at the target time.
2. The method as described in claim 1, characterized in that, The step of determining the filter coefficients of the sub-filter at the target time based on the error information of echo cancellation performed by the sub-filter at the historical time and the filter coefficients of the sub-filter at the historical time includes: Obtain the output speech signal of the sub-filter at the historical moment; Based on the error information of echo cancellation performed by the sub-filter at the historical moment, determine the gain value of the sub-filter at the historical moment; The filter coefficients of the sub-filter at the target time are determined based on the gain value of the sub-filter at the historical time, the filter coefficients of the sub-filter at the historical time, and the output speech signal of the sub-filter at the historical time.
3. The method as described in claim 1, characterized in that, The step of determining the target filter coefficients of the target filter at the target time based on the filter coefficients estimated by the plurality of sub-filters at the target time includes: The Kronecker product of the filter coefficients of the multiple sub-filters at the target time is taken as the target filter coefficient of the target filter at the target time.
4. The method according to any one of claims 1-3, characterized in that, The plurality of sub-filters are two in number, and the method further includes: Obtain the first length of the filter coefficients of the target filter; The filter coefficients of the first length are decomposed to obtain sub-filters with filter coefficients of the second length and sub-filters with filter coefficients of the third length; wherein the product of the second length and the third length is the first length.
5. The method according to any one of claims 1-3, characterized in that, The method further includes: Determine the zero norm corresponding to the filter coefficients of the sub-filter; The zero norm corresponding to the filter coefficients of the sub-filter is used as the constraint information for the filter coefficients of the sub-filter.
6. The method according to any one of claims 1-3, characterized in that, The method further includes: Acquire the audio signal to be processed collected by the audio acquisition device at the target time; Based on the target filter coefficients estimated at the target time, echo cancellation is performed on the speech signal to be processed to obtain the target speech signal at the target time.
7. A device for determining filter coefficients, characterized in that, include: The acquisition module is used to acquire the filter coefficients of multiple sub-filters in historical moments prior to the target time; wherein, the multiple sub-filters are obtained by decomposing the target filter; The first determining module is used to determine the constraint information of the filter coefficients of each sub-filter based on the filter coefficients of the sub-filter. The second determining module is used to determine the error information of the sub-filter performing echo cancellation on the input speech signal at the historical time based on the constraint information of the filter coefficients of the sub-filter at the historical time. The third determining module is used to determine the filter coefficients of the sub-filter at the target time based on the error information of echo cancellation performed by the sub-filter at the historical time and the filter coefficients of the sub-filter at the historical time. The fourth determining module is used to determine the target filter coefficient of the target filter at the target time based on the filter coefficients of the plurality of sub-filters at the target time.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Cited By
Mobile device and self-noise elimination method thereof
CN121838709A