Echo Cancellation Method, Device, Equipment and Storage Medium
By estimating and adjusting the delay when the voice interaction device of the computing device changes, and echo signals are eliminated using LMS and AEC algorithms, the problem of poor echo cancellation is solved, and more efficient echo cancellation is achieved.
Patent Information
- Application Number
- CN202111171723.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2039-03-18
AI Technical Summary
In the prior art, the use of default delay in echo cancellation processing leads to poor echo cancellation effect.
When the voice interaction device of the computing device changes, the delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal is estimated, and the echo signal is cancelled according to the estimated time delay, and the minimum mean square LMS algorithm and acoustic echo cancellation AEC algorithm are used for precise cancellation.
The echo cancellation effect is improved, and the problem of poor echo cancellation effect is avoided due to the use of default delay is ensured that the delay can be adjusted in time when the voice interactive device changes, which improves the accuracy of echo cancellation.
Smart Images

Figure CN113903351B_ABST
Abstract
Description
[0001] This application is a divisional application of the application with the application number 201910205707.9, the application date of March 18, 2019, and the invention title of "Echo Cancellation Method, Device, Equipment and Storage Medium". Technical Field
[0002] The present disclosure relates to the field of signal processing, and in particular, to an echo cancellation method, device, equipment and storage medium. Background Art
[0003] Currently, in speech recognition, echo cancellation processing, such as the Acoustic Echo Cancellation (AEC) algorithm, can be used to eliminate the echo in the collected speech signal.
[0004] In the prior art, the echo cancellation processing specifically eliminates the echo signal included in the speech signal collected by the microphone according to the time delay between the reference signal played and the echo signal corresponding to the reference signal collected by the microphone, so as to obtain the original signal sent by the speaker and avoid the echo caused by the echo signal being superimposed on the original signal. Usually, the time delay used in the echo cancellation processing is the default time delay, that is, based on the default time delay, the echo signal included in the speech signal collected by the microphone is eliminated.
[0005] However, in the prior art, there is a problem that the echo cancellation effect is poor due to the use of the default time delay in the echo cancellation processing. Summary of the Invention
[0006] Embodiments of the present disclosure provide an echo cancellation method, device, equipment and storage medium to solve the problem that the echo cancellation effect is poor due to the use of the default time delay in the prior art for echo cancellation processing.
[0007] In a first aspect, embodiments of the present disclosure provide an echo cancellation method, including:
[0008] When the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal;
[0009] The computing device cancels the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0010] In a possible implementation, if the connection object of the terminal computing device changes, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device.
[0011] In a possible implementation, if the computing device changes from being connected to the target device to not being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the target device includes the first voice interaction device, and the computing device includes the second voice interaction device;
[0012] Alternatively, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device includes the first voice interaction device, and the target device includes the second voice interaction device.
[0013] In a possible implementation, the target device is a vehicle.
[0014] In a possible implementation, if the computing device changes from being connected to a first target device to being connected to a second target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the first target device includes the first voice interaction device, and the second target device includes the second voice interaction device.
[0015] In a possible implementation, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal, including:
[0016] The computing device determines the time difference between each first time point in the multiple first time points and the second time point corresponding to each first time point according to the multiple first time points and the multiple second time points in one-to-one correspondence with the multiple first time points, obtaining multiple time differences. The first time point is the time point when the second voice interaction device plays the reference signal, and the second time point is the time point when the second voice device collects the echo signal corresponding to the reference signal played at the corresponding first time point;
[0017] The computing device determines the time delay between the reference signal and the echo signal according to the multiple time differences.
[0018] In a possible implementation, the computing device determines the time delay between the reference signal and the echo signal according to the multiple time differences, including:
[0019] The computing device determines the time delay between the reference signal and the echo signal according to the multiple time differences and a preset estimation algorithm.
[0020] In a possible implementation, the preset estimation algorithm is the least mean square (LMS) algorithm.
[0021] In a possible implementation, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, including:
[0022] The computing device determines whether the time delay is within a preset time delay range;
[0023] If the time delay is within the time delay range, the echo signal in the original signal collected by the second voice interaction device is eliminated according to the time delay;
[0024] If the time delay is not within the time delay range, the echo signal in the original signal collected by the second voice interaction device is eliminated according to the time delay within the time delay range.
[0025] In a possible implementation, the terminal computing device eliminates the echo signal in the original signal collected according to the estimated time delay, including:
[0026] The terminal computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay by using an acoustic echo cancellation (AEC) algorithm.
[0027] In a possible implementation, after the terminal computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, it further includes:
[0028] Performing speech recognition on the speech signal obtained after elimination to obtain a speech recognition result;
[0029] Performing subsequent processing according to the speech recognition result.
[0030] In a possible implementation, the subsequent processing includes wake-up processing and / or output processing.
[0031] In a second aspect, an embodiment of the present disclosure provides an echo cancellation device applied to a computing device, including:
[0032] An estimation module, configured to estimate the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device;
[0033] An elimination module, configured to eliminate the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0034] In a possible implementation, if the connection object of the terminal computing device changes, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device.
[0035] In a possible implementation, if the computing device changes from being connected to a target device to not being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the target device includes the first voice interaction device, and the computing device includes the second voice interaction device;
[0036] Alternatively, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device includes the first voice interaction device, and the target device includes the second voice interaction device.
[0037] In a possible implementation, if the computing device changes from being connected to a first target device to being connected to a second target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the first target device includes the first voice interaction device, and the second target device includes the second voice interaction device.
[0038] In a possible implementation, the estimation module is specifically configured to:
[0039] According to a plurality of first time points and a plurality of second time points corresponding to the plurality of first time points one by one, determine the time difference between each first time point in the plurality of first time points and the second time point corresponding to each first time point, to obtain a plurality of time differences, where the first time point is the time point when the second voice interaction device plays a reference signal, and the second time point is the time point when the second voice device collects an echo signal corresponding to the reference signal played at the corresponding first time point;
[0040] According to the plurality of time differences, determine the time delay between the reference signal and the echo signal.
[0041] In a possible implementation, the estimation module is used to determine the time delay between the reference signal and the echo signal according to the plurality of time differences, specifically including:
[0042] According to the plurality of time differences and a preset estimation algorithm, determine the time delay between the reference signal and the echo signal.
[0043] In a possible implementation, the preset estimation algorithm is the least mean square (LMS) algorithm.
[0044] In a possible implementation, the cancellation module is specifically configured to:
[0045] Determine whether the time delay is within a preset time delay range;
[0046] If the time delay is within the time delay range, cancel the echo signal in the original signal collected by the second voice interaction device according to the time delay;
[0047] If the time delay is not within the time delay range, cancel the echo signal in the original signal collected by the second voice interaction device according to the time delay within the time delay range.
[0048] In a possible implementation, the cancellation module cancels the echo signal in the original signal collected by the second voice interaction device according to the time delay, which specifically includes:
[0049] According to the estimated time delay, use the acoustic echo cancellation (AEC) algorithm to cancel the echo signal in the original signal collected by the second voice interaction device.
[0050] In a possible implementation, the device further includes: a response module;
[0051] The response module is configured to: perform speech recognition on the voice signal obtained after cancellation to obtain a speech recognition result; and perform subsequent processing according to the speech recognition result.
[0052] In a possible implementation, the subsequent processing includes wake-up processing and / or output processing.
[0053] In a third aspect, an embodiment of the present disclosure provides an echo cancellation device, including:
[0054] A processor and a memory for storing computer instructions; the processor runs the computer instructions to execute the method according to any one of the first aspects above.
[0055] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the echo cancellation device, enabling the echo cancellation device to execute the method according to any one of the first aspects above.
[0056] The echo cancellation method, device, equipment and storage medium provided by the embodiments of the present disclosure estimate the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device, and cancel the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay. When the voice interaction device used by the computer device changes, it can timely estimate the time delay of the changed voice interaction device, and cancel the echo signal in the original signal collected by the changed voice interaction device based on the estimated time delay. This can not only avoid the problem of poor echo cancellation effect caused by using the default time delay, but also avoid the problem of poor echo cancellation effect caused by inaccurate time delay when using the time delay of the voice interaction device before the change to cancel the echo signal in the original signal collected by the changed voice interaction device, thus improving the echo cancellation effect.
[0057] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings according to these drawings.
[0059] Figure 1 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 1 ;
[0060] Figure 2 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 2 ;
[0061] Figure 3 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 3 ;
[0062] Figure 4 Schematic flow chart of the first embodiment of the echo cancellation method provided by the embodiments of the present disclosure;
[0063] Figure 5 Schematic flow chart of the second embodiment of the echo cancellation method provided by the embodiments of the present disclosure;
[0064] Figure 6 Schematic flowchart of the third embodiment of the echo cancellation method provided by the embodiments of the present disclosure;
[0065] Figure 7 Schematic structural diagram of the first embodiment of the echo cancellation device provided by the embodiments of the present disclosure;
[0066] Figure 8 Schematic structural diagram of the second embodiment of the echo cancellation device provided by the embodiments of the present disclosure. Detailed implementation manners
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained based on the embodiments in the present disclosure fall within the scope of protection of the present disclosure.
[0068] Figure 1 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 1 , such as Figure 1 shown. In this application scenario, it may include a computing device 11, and the computing device 11 may include at least two voice interaction devices, such as Figure 1 voice interaction device a and voice interaction device b in
[0069] Figure 2 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 2 , such as Figure 2 shown. In this application scenario, it may include a computing device 11 and a first target device 12. Among them, the computing device 11 may include at least one voice interaction device, and the first target device 12 may include at least one voice interaction device, such as Figure 1The computing device 11 includes a voice interaction device a, and the first target device 12 includes a voice interaction device b. The computing device 11 can perform voice interaction with the user using the voice interaction device a of the computing device 11 or the voice interaction device b of the first target device 12. Specifically, the computing device 11 can use the voice interaction device a of the computing device 11 to collect voice and use the voice interaction device b of the computing device 11 to play voice, such as playing music, playing navigation, etc.; or, the computing device 11 can use the voice interaction device b of the first target device 12 to collect voice and use the voice interaction device b of the first target device 12 to play voice.
[0070] Figure 3 Schematic diagram of the application scenario of the echo cancellation method provided by the embodiments of the present disclosure Figure 3 , such as Figure 3 As shown, this application scenario may include a computing device 11, a first target device 12, and a second target device 13. Among them, the first target device 12 may include at least one voice interaction device, and the second target device 13 may include at least one voice interaction device. For example Figure 1 in the first target device 12 includes a voice interaction device a, and the second target device 13 includes a voice interaction device b. The computing device 11 can perform voice interaction with the user using the voice interaction device a of the first target device 12 or the voice interaction device b of the second target device 13. Specifically, the computing device 11 can use the voice interaction device a of the first target device 12 to collect voice and use the voice interaction device b of the first target device 12 to play voice, such as playing music, playing navigation, etc.; or, the computing device 11 can use the voice interaction device b of the second target device 13 to collect voice and use the voice interaction device b of the second target device 13 to play voice.
[0071] It can be understood that the above three application scenarios can be combined. An application scenario may include a computing device 11, a first target device 12, and a second target device 13. Among them, the computing device 11 may include at least two voice interaction devices, and the first target device 12 and the second target device 13 may each include one voice interaction device. Among them, the computing device 11 can use one voice interaction device of the computing device 11 to collect voice and use this voice interaction device of the computing device 11 to play voice; or, the computing device 11 can use another voice interaction device of the computing device 11 to collect voice and use this another voice interaction device of the computing device 11 to play voice; or, the computing device 11 can use the voice interaction device of the first target device 12 to collect voice and use the voice interaction device of the first target device 12 to play voice; the computing device 11 can use the voice interaction device of the second target device 13 to collect voice and use the voice interaction device of the second target device 13 to play voice.
[0072] It should be noted that the voice interaction device in the embodiments of the present disclosure can be any physical device capable of collecting and playing voices.
[0073] It should be noted that the computing device 11 can specifically be a device capable of playing and collecting voices through the voice interaction device and having a certain computing ability (for example, estimating the time delay). The specific type of the computing device is not limited in the present disclosure. For example, it can be a mobile phone, a tablet computer, a wearable device, etc.
[0074] It should be noted that for Figure 2 and Figure 3 the connection manner between the computing device and the voice interaction device of the target device, the present disclosure can make no limitation.
[0075] Figure 4 FIG. is a schematic flowchart of the first embodiment of the echo cancellation method provided by the embodiments of the present disclosure. The method of this embodiment can be executed by a computing device. As Figure 4 shown, the method of this embodiment can include:
[0076] Step 401, when the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal.
[0077] In this step, the first voice interaction device can be understood as the above-mentioned voice interaction device a, and the second voice interaction device can be understood as the above-mentioned voice interaction device b; or, the first voice interaction device can be understood as the above-mentioned voice interaction device b, and the second voice interaction device can be understood as the above-mentioned voice interaction device a. The voice interaction device used by the computing device can be understood as the voice interaction device used by the computing device to play and collect voices, and the user can perform voice interaction with the computing device through this voice interaction device.
[0078] For Figure 1 the application scenario shown, when the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, for example, the voice interaction device used by the computing device 11 changes from the voice interaction device a of the computing device 11 to the voice interaction device b of the computing device 11. At this time, the voice interaction device a of the computing device 11 can be understood as the first voice interaction device, and the voice interaction device b of the computing device 11 can be understood as the second voice interaction device.
[0079] For Figure 2In the application scenario shown, the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. For example, the voice interaction device used by computing device 11 changes from voice interaction device a of computing device 11 to voice interaction device b of the first target device 12. At this time, voice interaction device a of computing device 11 can be understood as the first voice interaction device, and voice interaction device b of the first target device 12 can be understood as the second voice interaction device.
[0080] For Figure 3 In the application scenario shown, the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. For example, the voice interaction device used by computing device 11 changes from voice interaction device a of the first target device 12 to voice interaction device b of the second target device 13. At this time, voice interaction device a of the first target device 12 can be understood as the first voice interaction device, and voice interaction device b of the second target device 13 can be understood as the second voice interaction device.
[0081] Among them, the voice signal played by the computing device using the voice interaction device can be called the reference signal, and the voice signal collected by the computing device using this voice interaction device can be called the original signal. It can be understood that after the reference signal is played by the computing device, the played sound can be collected by this voice interaction device, that is, the collected original signal can include the voice signal played by the reference signal of this computing device.
[0082] Since the hardware structures of different voice interaction devices are different, the time delay between the reference signal played by the computing device and the echo signal corresponding to the collected reference signal may be different for different voice interaction devices. Here, by estimating the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device, it is possible to estimate in a timely manner the time delay between the reference signal played by the changed second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes.
[0083] It can be understood that during the process of the computing device playing the reference signal, when the user speaks, the collected original signal can also include the user's voice signal.
[0084] It should be noted that the present disclosure may not limit the specific manner in which the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal.
[0085] It should be noted that the embodiments of the present disclosure may not limit the specific manner in which the computing device determines that the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. For example, the computing device may monitor the voice interaction device used to determine whether the voice interaction device used has changed, that is, whether it has changed from the first voice interaction device to the second voice interaction device.
[0086] Step 402, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0087] In this step, the embodiments of the present disclosure may not limit the specific manner of eliminating the echo signal in the original signal collected by the second voice interaction device according to the time delay estimated in step 401. For example, the reference signal may be moved according to the estimated time delay, and the echo signal in the original signal collected by the second voice interaction device may be eliminated according to the collected original signal and the moved reference signal.
[0088] Here, since step 401 estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device, the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal can be used in step 402 to eliminate the echo signal in the original signal collected by the second voice interaction device, avoiding the problem of poor echo cancellation effect caused by inaccurate time delay when the echo signal in the original signal collected by the second voice interaction device is eliminated using the time delay between the reference signal played by the first voice interaction device and the echo signal corresponding to the collected reference signal after the voice interaction device changes from the first voice interaction device to the second voice interaction device.
[0089] The echo cancellation method provided in this embodiment estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. Then, according to the estimated time delay, the echo signal in the original signal collected by the second voice interaction device is cancelled. This realizes that when the voice interaction device used by the computer device changes, the time delay of the changed voice interaction device can be estimated in a timely manner, and the echo signal in the original signal collected by the changed voice interaction device is cancelled based on the estimated time delay. This not only avoids the problem of poor echo cancellation effect caused by using the default time delay, but also avoids the problem of poor echo cancellation effect caused by inaccurate time delay when using the time delay of the voice interaction device before the change (i.e., the first voice interaction device) to cancel the echo signal in the original signal collected by the changed voice interaction device (i.e., the second voice interaction device), thus improving the echo cancellation effect.
[0090] Figure 5 This is a schematic flowchart of Embodiment 2 of the echo cancellation method provided by the present disclosure. Based on the embodiment shown Figure 5 an optional implementation manner of estimating the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device changes is mainly described.
[0091] Step 501: Determine whether the connection object of the computing device has changed.
[0092] In this step, if the connection object of the computing device has changed, it can indicate that the voice interaction device has changed, that is, the voice interaction device used by the computing device has changed from the first voice interaction device to the second voice interaction device. If the connection object of the computing device has not changed, it can indicate that the voice interaction device has not changed, that is, the voice interaction device used by the computing device has not changed from the first voice interaction device to the second voice interaction device.
[0093] Among them, the first voice interaction device can be understood as the voice interaction device used by the computing device before the change of the voice interaction device. The second voice interaction device can be understood as the voice interaction device used by the computing device after the change of the voice interaction device.
[0094] Optionally, the change of the connection object of the computing device can specifically be the change between two states: the computing device is connected to the target device and the computing device is not connected to the target device.
[0095] Specifically, if the computing device changes from being connected to the target device to not being connected to the target device, the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. The target device includes the first voice interaction device, and the computing device includes the second voice interaction device. Or, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. The computing device includes the first voice interaction device, and the target device includes the second voice interaction device.
[0096] For example, as Figure 2 shown, when the computing device 11 is connected to the first target device 12, the computing device 11 can use the voice interaction device b of the first target device 12 to perform voice interaction with the user. When the computing device 11 is not connected to the first target device 12, the computing device 11 can use the voice interaction device a of the computing device 11 to perform voice interaction with the user. Therefore, when the connection state between the computing device 11 and the first target device 12 changes, it can indicate that the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. Specifically, when the computing device 11 changes from being connected to the first target device 12 to not being connected to the first target device 12, the voice interaction device b can be regarded as the first voice interaction device, and the voice interaction device a can be regarded as the second voice interaction device. When the computing device 11 changes from not being connected to the first target device 12 to being connected to the first target device 12, the voice interaction device a can be regarded as the first voice interaction device, and the voice interaction device b can be regarded as the second voice interaction device.
[0097] It should be noted that the target device can specifically be a device to which the computing device 11 can establish a connection and can control a part of its hardware, and this part of the hardware includes a voice interaction device. Exemplarily, the target device can be a vehicle. At this time, the computing device can be a computing device that supports a specific function, and the specific function is that the computing device can establish a connection with the target device and can control a part of the hardware of the target device.
[0098] Or, optionally, the change in the connection object of the computing device can specifically be a change between two states: the computing device is connected to one target device and the computing device is connected to another target device. Specifically, if the computing device changes from being connected to the first target device to being connected to the second target device, the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. The first target device includes the first voice interaction device, and the second target device includes the second voice interaction device.
[0099] For example, as Figure 3As shown, when the computing device 11 is connected to the first target device 12, the computing device 11 can use the voice interaction device a of the first target device 12 to interact with the user by voice; when the computing device 11 is connected to the second target device 13, the computing device 11 can use the voice interaction device b of the second target device 13 to interact with the user by voice. Therefore, when the connection state between the computing device 11 and the first target device 12 and the second target device 13 changes, it can indicate that the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device. Specifically, when the computing device 11 changes from being connected to the first target device 12 to being connected to the second target device 13, the voice interaction device a can be regarded as the first voice interaction device, and the voice interaction device b can be regarded as the second voice interaction device; when the computing device 11 changes from being connected to the second target device 13 to being connected to the first target device 12, the voice interaction device b can be regarded as the first voice interaction device, and the voice interaction device a can be regarded as the second voice interaction device.
[0100] Wherein, if the connection object of the computing device changes, step 502 is executed; if the connection object of the computing device does not change, the process ends.
[0101] Step 502, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal.
[0102] In this step, the second voice interaction device can be understood as the voice interaction device currently used by the computing device. Optionally, the time delay can be determined through the following steps:
[0103] Step A, the computing device determines the time difference between each first time point and the second time point corresponding to each first time point among the multiple first time points according to the multiple first time points and the multiple second time points in one-to-one correspondence with the multiple first time points, to obtain multiple time differences. The first time point is the time point when the second voice interaction device plays the reference signal, and the second time point is the time point when the second voice device collects the echo signal corresponding to the reference signal played at the corresponding first time point.
[0104] Here, to avoid the problem that the determined time delay is inaccurate due to the inaccuracy of a single time difference, optionally, multiple time differences can be obtained based on multiple first time points and multiple second time points. For example, when the computing device plays the voice signal x (which can be understood as a reference signal), it records the time point 1 when playing the voice signal x (which can be understood as the first time point), collects the original signal, and records the time point 2 when the original signal is collected. If the voice signal x is included in the original signal, then the time point 2 is the second time point corresponding to the time point 1. Further, the time difference between the time point 2 and the time point 1 can be obtained. For another example, when the computing device plays the voice signal y (which can be understood as a reference signal), it records the time point 3 when playing the voice signal y (which can be understood as the first time point), collects the original signal, and records the time point 4 when the original signal is collected. If the voice signal y is included in the original signal, then the time point 4 is the second time point corresponding to the time point 3. Further, the time difference between the time point 4 and the time point 3 can be obtained.
[0105] It should be noted that the present disclosure does not limit the specific manner in which the reference signal is included in the collected original signal.
[0106] Step B, the computing device determines the time delay between the reference signal and the echo signal according to the multiple time differences.
[0107] Here, specifically, mathematical calculations can be performed on the multiple time differences to obtain the time delay between the reference signal and the echo signal. For example, the multiple time differences can be averaged to obtain the time delay. Optionally, when obtaining the time delay based on the time difference, a certain estimation algorithm can be used. Further optionally, step B may specifically include: the computing device determines the time delay between the reference signal and the echo signal according to the multiple time differences and a preset estimation algorithm.
[0108] Exemplarily, the preset estimation algorithm is the Least-Mean-Square (LMS) algorithm. Here, by setting the preset estimation algorithm as the LMS algorithm, the time delay is determined in a machine learning manner based on multiple time differences, providing the accuracy of time delay determination.
[0109] Step 503, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0110] In this step, optionally, the AEC algorithm can be used to eliminate the echo signal in the collected original signal. Specifically, step 503 may include: the computing device uses the acoustic echo cancellation AEC algorithm to eliminate the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0111] Considering that after a certain AEC algorithm is determined, the applicable time delay range is fixed. Therefore, to avoid the problem of poor echo cancellation effect caused by the determined time delay being outside this certain time delay range, optionally, step 503 may specifically include: the computing device determines whether the time delay is within a preset time delay range; if the time delay is within the time delay range, the echo signal in the original signal collected by the second voice interaction device is cancelled according to the time delay; if the time delay is not within the time delay range, the echo signal in the original signal collected by the second voice interaction device is cancelled according to the time delay within the time delay range.
[0112] The echo cancellation method provided in this embodiment determines whether the connection object of the computing device has changed. If the connection object of the computing device has changed, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal. The computing device cancels the echo signal in the collected original signal according to the estimated time delay, realizing that the change in the connection object of the computing device represents that the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device.
[0113] Figure 6 It is a schematic flowchart of the third embodiment of the echo cancellation method provided by the present disclosure. On the basis of the above embodiment, this embodiment mainly describes an optional implementation manner after echo cancellation. As Figure 6 shown, the method of this embodiment may include:
[0114] Step 601, when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal.
[0115] It should be noted that step 601 is similar to step 401 and will not be elaborated here.
[0116] Step 602, the computing device cancels the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0117] It should be noted that step 602 is similar to step 402 and will not be elaborated here.
[0118] Step 603, perform speech recognition on the voice signal obtained after cancellation to obtain a speech recognition result.
[0119] In this step, the voice recognition result can be, for example, "power on", "weather", etc. It should be noted that the present disclosure does not limit the specific manner of performing voice recognition on the voice signal obtained after cancellation.
[0120] Since steps 601 and 602 can improve the echo cancellation effect, the accuracy of the voice signal based on which step 603 performs voice recognition is higher, thereby improving the accuracy of the voice recognition result.
[0121] Step 604, perform subsequent processing according to the voice recognition result.
[0122] In this step, after obtaining the voice recognition result, certain processing can be performed based on the voice recognition result. Here, the present disclosure does not limit the type of processing. Exemplarily, the subsequent processing may include wake-up processing and / or output processing.
[0123] Among them, for wake-up processing, exemplarily, it can be determined whether the voice recognition result is the same as a preset wake-up instruction. If the voice recognition result is the same as the preset result, the application program corresponding to the preset wake-up instruction of the computing device is woken up. For output processing, exemplarily, the voice recognition result can be output in the text box of the input interface.
[0124] The echo cancellation method provided in this embodiment, when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal. The computing device cancels the echo signal in the collected original signal according to the estimated time delay, performs voice recognition on the voice signal obtained after cancellation to obtain a voice recognition result, and performs subsequent processing according to the voice recognition result, achieving an improvement in the accuracy of the voice recognition result on the basis of improving the echo cancellation effect, thereby improving the user experience.
[0125] Figure 7 It is a schematic structural diagram of Embodiment 1 of the echo cancellation device provided by the present disclosure. The device provided in this embodiment can be applied to the above method embodiment to implement the functions of its computing device. As Figure 7 shown, the device of this embodiment may include: an estimation module 701 and a cancellation module 702.
[0126] Among them, the estimation module 701 is configured to estimate the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal when the voice interaction device used by the computing device changes from the first voice interaction device to the second voice interaction device;
[0127] An elimination module 702 is configured to eliminate the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0128] In a possible implementation, if the connection object of the computing device changes, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device.
[0129] In a possible implementation, if the computing device changes from being connected to a target device to not being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, where the target device includes the first voice interaction device and the computing device includes the second voice interaction device;
[0130] Alternatively, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, where the computing device includes the first voice interaction device and the target device includes the second voice interaction device.
[0131] In a possible implementation, the target device is a vehicle.
[0132] In a possible implementation, if the computing device changes from being connected to a first target device to being connected to a second target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, where the first target device includes the first voice interaction device and the second target device includes the second voice interaction device.
[0133] In a possible implementation, the estimating module 701 is specifically configured to:
[0134] According to a plurality of first time points and a plurality of second time points corresponding to the plurality of first time points one by one, determine the time difference between each first time point in the plurality of first time points and the second time point corresponding to each first time point, to obtain a plurality of time differences, where the first time point is the time point when the second voice interaction device plays a reference signal, and the second time point is the time point when the second voice device collects the echo signal corresponding to the reference signal played at the corresponding first time point;
[0135] According to the plurality of time differences, determine the time delay between the reference signal and the echo signal.
[0136] In a possible implementation, the estimating module 701 is configured to determine the time delay between the reference signal and the echo signal according to the plurality of time differences, specifically including:
[0137] Determine the time delay between the reference signal and the echo signal according to the multiple time differences and a preset estimation algorithm.
[0138] In a possible implementation, the preset estimation algorithm is the least mean square (LMS) algorithm.
[0139] In a possible implementation, the elimination module 702 is specifically configured to:
[0140] Determine whether the time delay is within a preset time delay range;
[0141] If the time delay is within the time delay range, eliminate the echo signal in the original signal collected by the second voice interaction device according to the time delay;
[0142] If the time delay is not within the time delay range, eliminate the echo signal in the original signal collected by the second voice interaction device according to the time delay within the time delay range.
[0143] In a possible implementation, the elimination module 702 eliminates the echo signal in the original signal collected by the second voice interaction device according to the time delay, which specifically includes:
[0144] According to the estimated time delay, use the acoustic echo cancellation (AEC) algorithm to eliminate the echo signal in the original signal collected by the second voice interaction device.
[0145] In a possible implementation, the device further includes: a response module 703;
[0146] The response module 703 is configured to: perform speech recognition on the voice signal obtained after elimination to obtain a speech recognition result; and perform subsequent processing according to the speech recognition result.
[0147] In a possible implementation, the subsequent processing includes wake-up processing and / or output processing.
[0148] The device in this embodiment can be used to execute the technical solutions of the foregoing method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0149] Figure 8 FIG. is a schematic structural diagram of Embodiment 2 of the echo cancellation device provided by an embodiment of the present disclosure. As Figure 8 shown, the device may include: a processor 801 and a memory 802 for storing computer instructions.
[0150] Wherein, the processor 801 runs the computer instructions to execute the following method:
[0151] When the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal;
[0152] The computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0153] In a possible implementation, if the connection object of the computing device changes, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device.
[0154] In a possible implementation, if the computing device changes from being connected to a target device to not being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The target device includes the first voice interaction device, and the computing device includes the second voice interaction device;
[0155] Alternatively, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The computing device includes the first voice interaction device, and the target device includes the second voice interaction device.
[0156] In a possible implementation, the target device is a vehicle.
[0157] In a possible implementation, if the computing device changes from being connected to a first target device to being connected to a second target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The first target device includes the first voice interaction device, and the second target device includes the second voice interaction device.
[0158] In a possible implementation, the computing device estimates the time delay between the played reference signal and the echo signal corresponding to the collected reference signal, including:
[0159] The computing device determines the time difference between each first time point and the second time point corresponding to each first time point among the multiple first time points according to the multiple first time points and the multiple second time points in one-to-one correspondence with the multiple first time points, and obtains multiple time differences. The first time point is the time point when the second voice interaction device plays the reference signal, and the second time point is the time point when the second voice device collects the echo signal corresponding to the reference signal played at the corresponding first time point;
[0160] The computing device determines the time delay between the reference signal and the echo signal according to the plurality of time differences.
[0161] In a possible implementation, the computing device determines the time delay between the reference signal and the echo signal according to the plurality of time differences, including:
[0162] The computing device determines the time delay between the reference signal and the echo signal according to the plurality of time differences and a preset estimation algorithm.
[0163] In a possible implementation, the preset estimation algorithm is the least mean square (LMS) algorithm.
[0164] In a possible implementation, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, including:
[0165] The computing device determines whether the time delay is within a preset time delay range;
[0166] If the time delay is within the time delay range, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the time delay;
[0167] If the time delay is not within the time delay range, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the time delay within the time delay range.
[0168] In a possible implementation, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, including:
[0169] The computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay by using an acoustic echo cancellation (AEC) algorithm.
[0170] In a possible implementation, after the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, it further includes:
[0171] Performing speech recognition on the voice signal obtained after elimination to obtain a speech recognition result;
[0172] Performing subsequent processing according to the speech recognition result.
[0173] In a possible implementation, the subsequent processing includes wake-up processing and / or output processing.
[0174] An embodiment of the present disclosure further provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an echo cancellation device, the echo cancellation device can execute an echo cancellation method, which includes:
[0175] When the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal. The voice interaction device is used for the user to perform voice interaction with the computing device;
[0176] The computing device cancels the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay.
[0177] In a possible implementation, if the connection object of the computing device changes, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device.
[0178] In a possible implementation, if the computing device changes from being connected to a target device to not being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The target device includes the first voice interaction device, and the computing device includes the second voice interaction device;
[0179] Or, if the computing device changes from not being connected to the target device to being connected to the target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The computing device includes the first voice interaction device, and the target device includes the second voice interaction device.
[0180] In a possible implementation, the target device is a vehicle.
[0181] In a possible implementation, if the computing device changes from being connected to a first target device to being connected to a second target device, the voice interaction device used by the computing device changes from a first voice interaction device to a second voice interaction device. The first target device includes the first voice interaction device, and the second target device includes the second voice interaction device.
[0182] In a possible implementation, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal, including:
[0183] The computing device determines, according to a plurality of first time points and a plurality of second time points that are in one-to-one correspondence with the plurality of first time points, a time difference between each first time point in the plurality of first time points and the second time point corresponding to each first time point, to obtain a plurality of time differences, where the first time point is the time point when the second voice interaction device plays a reference signal, and the second time point is the time point when the second voice device collects an echo signal corresponding to the reference signal played at the corresponding first time point;
[0184] The computing device determines a time delay between the reference signal and the echo signal according to the plurality of time differences.
[0185] In a possible implementation, the computing device determines a time delay between the reference signal and the echo signal according to the plurality of time differences, including:
[0186] The computing device determines a time delay between the reference signal and the echo signal according to the plurality of time differences and a preset estimation algorithm.
[0187] In a possible implementation, the preset estimation algorithm is the least mean square (LMS) algorithm.
[0188] In a possible implementation, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, including:
[0189] The computing device determines whether the time delay is within a preset time delay range;
[0190] If the time delay is within the time delay range, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the time delay;
[0191] If the time delay is not within the time delay range, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the time delay within the time delay range.
[0192] In a possible implementation, the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, including:
[0193] The computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay by using an acoustic echo cancellation (AEC) algorithm.
[0194] In a possible implementation, after the computing device eliminates the echo signal in the original signal collected by the second voice interaction device according to the estimated time delay, it further includes:
[0195] Perform speech recognition on the obtained voice signal after elimination to obtain a speech recognition result;
[0196] Perform subsequent processing according to the speech recognition result.
[0197] In a possible implementation, the subsequent processing includes wake-up processing and / or output processing.
[0198] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, and are not intended to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. An echo cancellation method, comprising: Determining whether the connection object of the computing device has changed; If so, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal, and the second voice interaction device is the voice interaction device currently used by the computing device; The computing device determines whether the time delay is within a preset time delay range; If the time delay is within the time delay range, the echo signal in the original signal collected by the second voice interaction device is cancelled according to the time delay; If the time delay is not within the time delay range, the echo signal in the original signal collected by the second voice interaction device is cancelled according to the time delay within the time delay range.
2. The method according to claim 1, wherein The computing device estimating the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal includes: The computing device determines the time difference between each first time point and the second time point corresponding to each first time point among the plurality of first time points according to a plurality of first time points and a plurality of second time points in one-to-one correspondence with the plurality of first time points, obtaining a plurality of time differences, where the first time point is the time point when the second voice interaction device plays the reference signal, and the second time point is the time point when the second voice interaction device collects the echo signal corresponding to the reference signal played at the corresponding first time point; The computing device determines the time delay between the reference signal and the echo signal according to the plurality of time differences.
3. The method according to claim 2, wherein The computing device determining the time delay between the reference signal and the echo signal according to the plurality of time differences includes: The computing device determines the time delay between the reference signal and the echo signal according to the plurality of time differences and a preset estimation algorithm.
4. The method according to any one of claims 1 to 3, the method further comprising: After determining that the connection object of the computing device has changed, the computing device determines the second voice interaction device as the voice interaction device currently used.
5. The method according to claim 4, wherein The computing device determining that the voice interaction device currently used is the second voice interaction device includes at least one of the following: If the computing device plays and collects voices through the second voice interaction device, the computing device determines the second voice interaction device as the voice interaction device currently used; Or, If the computing device changes from being connected to a target device to not being connected to the target device, the computing device determines the second voice interaction device as the voice interaction device currently used, where the target device includes a first voice interaction device, and the computing device includes the second voice interaction device; Or, If the computing device changes from not being connected to the target device to being connected to the target device, the computing device determines the second voice interaction device as the voice interaction device currently used, where the computing device includes a first voice interaction device, and the target device includes the second voice interaction device; Or, If the connection object of the computing device changes from being connected to a first target device to being connected to a second target device, the computing device determines the second voice interaction device as the currently used voice interaction device. The first target device includes a first voice interaction device, and the second target device includes the second voice interaction device.
6. An echo cancellation device, comprising: An estimation module, configured to determine whether the connection object of the computing device has changed; If so, the computing device estimates the time delay between the reference signal played by the second voice interaction device and the echo signal corresponding to the collected reference signal. The second voice interaction device is the currently used voice interaction device of the computing device; An elimination module, configured to determine whether the time delay is within a preset time delay range; If the time delay is within the time delay range, the echo signal in the original signal collected by the second voice interaction device is eliminated according to the time delay; If the time delay is not within the time delay range, the echo signal in the original signal collected by the second voice interaction device is eliminated according to the time delay within the time delay range.
7. The device according to claim 6, wherein, The estimation module is specifically configured to: According to a plurality of first time points and a plurality of second time points corresponding to the plurality of first time points one by one, determine the time difference between each first time point in the plurality of first time points and the second time point corresponding to each first time point, obtaining a plurality of time differences. The first time point is the time point when the second voice interaction device plays the reference signal, and the second time point is the time point when the second voice interaction device collects the echo signal corresponding to the reference signal played at the corresponding first time point; According to the plurality of time differences, determine the time delay between the reference signal and the echo signal.
8. The device according to claim 7, wherein, The estimation module is specifically configured to: According to the plurality of time differences and a preset estimation algorithm, determine the time delay between the reference signal and the echo signal.
9. The device according to any one of claims 6 to 8, wherein the estimation module is further configured to: Determine the second voice interaction device as the currently used voice interaction device.
10. The device according to claim 9, wherein, The estimation module is specifically configured to perform at least one of the following: If the computing device plays and collects voices through the second voice interaction device, determine the second voice interaction device as the currently used voice interaction device; or, If the connection of the computing device changes from being connected to a target device to not being connected to the target device, determine the second voice interaction device as the currently used voice interaction device. The target device includes a first voice interaction device, and the computing device includes the second voice interaction device; Or, If the connection of the computing device changes from not being connected to the target device to being connected to the target device, determine the second voice interaction device as the currently used voice interaction device. The computing device includes a first voice interaction device, and the target device includes the second voice interaction device; or, If the computing device changes from being connected to a first target device to being connected to a second target device, the second voice interaction device is determined as the currently used voice interaction device, the first target device includes a first voice interaction device, and the second target device includes the second voice interaction device.
11. An echo cancellation device, comprising: a processor and a memory for storing computer instructions; The processor runs the computer instructions to execute the method according to any one of claims 1-5.
12. A computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an echo cancellation device, enabling the echo cancellation device to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Time delay estimation method for indoor echo cancellation of set-top box
CN107785026A
An echo cancellation method for improving VOIP call quality
CN109040501A