Echo cancellation method for voice interaction system, and electronic device and storage medium
By using a reference audio signal modulated with a parity bitstream and dynamically adjusting digital frequency modulation parameters in a voice interaction system, the problem of poor echo cancellation in complex environments is solved, achieving higher reliability and accuracy.
Patent Information
- Application Number
- CN202111559446.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-12-20
AI Technical Summary
Existing echo cancellation methods suffer from delay estimation errors in complex environments, resulting in poor echo cancellation performance.
The reference audio signal modulated by the check bitstream is used for testing. The digital frequency modulation parameters are dynamically adjusted to obtain accurate echo delay. The anti-interference and anti-channel loss performance of digital frequency modulation is utilized to simplify the echo delay algorithm.
This improves the reliability and accuracy of echo cancellation methods in complex environments and enhances the echo cancellation effect.
Smart Images

Figure CN114242101B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of audio signal processing, and more particularly, to an echo cancellation method for a voice interaction system, and an electronic device and a storage medium. BACKGROUND
[0002] In a voice interaction scenario such as a mobile phone, a telephone conference, etc., multiple users use microphones to pick up near-end speech and use loudspeakers to play far-end speech. The microphone of a near-end user not only picks up the user's own speech, but also picks up the far-end user's speech played by the loudspeaker and sends it back to the far-end user. The far-end user not only hears the near-end user's speech, but also hears his / her own speech. Acoustic echo seriously affects the user's voice interaction experience.
[0003] Acoustic echo is a phenomenon that the sound played by a loudspeaker is picked up by a microphone and sent back to the opposite end. Acoustic echo is further divided into direct echo and indirect echo. Direct echo refers to the sound played by a loudspeaker directly entering a microphone without any reflection. The delay of direct echo is short, and is related to the far-end speaker's speech energy, the distance, angle, loudspeaker playing volume, and microphone pickup sensitivity between the loudspeaker and the microphone, etc. Indirect echo refers to a set of echoes generated when the sound played by a loudspeaker is reflected one or more times through different echo paths and then enters a microphone.
[0004] Acoustic echo cancellation is to subtract an echo signal from a speech signal picked up by a microphone. Referring to Figure 1 , an echo cancellation system includes modules such as delay estimation, linear echo cancellation, double-talk detection, and residual echo cancellation. The difference between acoustic echo and original speech not only includes distortion caused by the non-linear characteristics of the near-end user's loudspeaker, but also includes the response of the near-end user's room system. An echo cancellation algorithm mainly uses an adaptive filter to simulate an echo path, and makes the impulse response as close as possible to the actual echo path, so as to obtain an estimated value of the echo signal, and then subtracts the estimated value from the near-end picked speech signal to achieve echo cancellation. Acoustic echo cancellation is an indispensable module in a voice interaction scenario.
[0005] In a voice interaction system, the echo cancellation methods that have been adopted include real-time echo cancellation and tuning echo cancellation. In the real-time echo cancellation method, a delay parameter is obtained according to the comparison between the near-end signal and the reference signal of real-time voice communication. In the tuning echo cancellation method, the actual environment is tuned before real-time voice communication to obtain a delay parameter. Compared with real-time echo cancellation, the delay parameter obtained by using tuning echo cancellation is not only more accurate, but also does not need to consume time to calculate the delay parameter in the voice communication stage, and the processing speed of the audio signal is faster, so that a better echo cancellation effect can be obtained.
[0006] For the echo cancellation algorithm of deep learning, the delay estimation is an important factor affecting the echo cancellation effect. The existing echo cancellation method still has the problem of error in delay estimation in a complex environment with too large environmental noise and / or too long delay time, resulting in poor echo cancellation effect. SUMMARY
[0007] In view of the above problems, the purpose of the present disclosure is to provide an echo cancellation method for a voice interaction system, an electronic device and a storage medium, wherein a reference audio signal modulated by a check bit stream is used for testing, and the digital frequency modulation parameters are dynamically adjusted according to the evaluation parameters, so as to improve the reliability of the echo cancellation method in a complex environment and improve the echo cancellation effect.
[0008] According to a first aspect of the present disclosure, an echo cancellation method for a voice interaction system is provided, comprising: tuning the voice interaction system to obtain an echo delay; and processing a picked-up audio signal in the voice interaction system to remove echo using the echo delay, wherein the tuning comprises: digitally frequency modulating using a check bit stream to generate a reference audio signal; testing using the reference audio signal to obtain an echo delay; and dynamically adjusting the digital frequency modulation parameters according to the evaluation parameters and repeating the testing step until the evaluation parameters are qualified.
[0009] Preferably, it further comprises: in the driving circuit of the loudspeaker, collecting the driving signal of the loudspeaker to obtain the reference audio signal; and in the signal processing circuit of the microphone, collecting the picked-up signal of the microphone to obtain the picked-up audio signal.
[0010] Preferably, the step of obtaining the echo delay comprises: demodulating the reference audio signal to obtain the time position of the check bit stream in the reference audio signal as a starting time; demodulating the picked-up audio signal to obtain the time position of the check bit stream in the picked-up audio signal as an arrival time; and obtaining the echo delay according to the difference between the starting time and the arrival time.
[0011] Preferably, the step of obtaining the echo delay comprises: estimating the starting time according to the time when the loudspeaker plays the real-time generated audio data; demodulating the picked-up audio signal to obtain the time position of the check bit stream in the picked-up audio signal as an arrival time; and obtaining the echo delay according to the difference between the starting time and the arrival time.
[0012] Preferably, at least one of the reference audio signal and the picked-up audio signal is demodulated to obtain a received bit stream, and the time position of the check bit stream in the corresponding audio signal is obtained according to the similarity between the received bit stream and the check bit stream.
[0013] Preferably, the step of dynamically adjusting the digital frequency modulation parameter according to the evaluation parameter comprises: comparing the evaluation parameter with a corresponding threshold value; and changing the value of the modulation parameter according to the comparison result.
[0014] Preferably, the evaluation parameter comprises at least one of the time length of the echo delay, the bit error rate of the audio signal, the signal-to-noise ratio of the audio signal, and the reverb time of the audio signal.
[0015] Preferably, the time length of the calibration bit stream is increased when the echo delay is greater than a corresponding threshold value.
[0016] Preferably, increasing the time length of the calibration bit stream comprises: increasing the data length of the calibration bit stream, and / or, decreasing the baud rate of the modulation signal.
[0017] Preferably, the baud rate of the modulation signal is decreased when the bit error rate of the audio signal is greater than a corresponding threshold value.
[0018] Preferably, the reference audio signal is demodulated to obtain a first bit stream, the picked-up audio signal is demodulated to obtain a second bit stream, and a similarity calculation is performed on the first bit stream and the second bit stream to obtain the bit error rate of the audio signal.
[0019] Preferably, at least one of the carrier frequency and the digital frequency is decreased when the signal-to-noise ratio of the audio signal is less than a corresponding threshold value.
[0020] Preferably, the signal-to-noise ratio of the audio signal is obtained by calculating the ratio of the modulation signal strength and the blank signal strength of the picked-up audio signal.
[0021] Preferably, before the digital frequency modulation is performed on the calibration bit stream to generate the reference audio signal, the tuning further comprises: performing a test on an analog audio signal to obtain an echo delay; and dynamically adjusting the digital frequency modulation parameter according to an evaluation parameter.
[0022] According to a second aspect of the present application, an electronic device is provided, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the above method.
[0023] According to a third aspect of the present application, a computer readable storage medium is provided, wherein the computer readable storage medium stores a computer program or instructions, and the computer program or instructions, when executed by a processor, implement the steps of the above method.
[0024] In this embodiment, a digital frequency modulation signal is used as a reference audio signal. The echo delay is obtained through testing using this reference audio signal, and the digital frequency modulation parameters are dynamically adjusted based on evaluation parameters. On one hand, this echo cancellation method utilizes the anti-interference and anti-channel loss properties of digital frequency modulation, thus it can be applied to the complex environments of voice interaction systems and improves reliability. On the other hand, by dynamically adjusting the digital frequency modulation parameters based on evaluation parameters, this echo cancellation method can improve the accuracy of the echo delay measurement and enhance the echo cancellation effect in the complex environments of voice interaction systems. Attached Figure Description
[0025] Figure 1 A schematic block diagram of an echo cancellation system in a voice interaction scenario is shown.
[0026] Figure 2 A flowchart of an echo cancellation method according to the prior art is shown.
[0027] Figure 3 The waveforms of the reference signal and the pickup signal are shown during the tuning phase of echo cancellation.
[0028] Figure 4 A flowchart of an echo cancellation method according to a first embodiment of the present disclosure is shown.
[0029] Figure 5 Show Figure 4 The detailed steps for obtaining the echo delay in the echo cancellation method shown are as follows.
[0030] Figure 6 The waveform of the modulated signal obtained by digitally frequency modulating the carrier using a parity bit stream is shown.
[0031] Figure 7 The data structure for the parity bit stream is shown.
[0032] Figure 8 A schematic block diagram of a demodulator that performs coherent demodulation of a modulated signal is shown.
[0033] Figure 9 Show Figure 4 The detailed steps for dynamically adjusting digital frequency modulation parameters in the echo cancellation method shown are as follows.
[0034] Figure 10 A schematic block diagram of an electronic device for echo cancellation according to a third embodiment of the present disclosure is shown. Detailed Implementation
[0035] For the purposes of promoting an understanding of the principles of the present disclosure, reference will now be made to the embodiments illustrated in the drawings. The present disclosure can be implemented in various forms and is not limited to the embodiments described herein. Rather, the embodiments are provided as examples of the disclosure.
[0036] In the following description, the term "user" refers to any end user interacting with the voice interactive system, unless otherwise specified.
[0037] The inventor has noticed that the existing voice interactive system has poor echo cancellation effect in a network environment with large delay, mainly because of the error in the echo delay estimation between the reference audio signal and the picked-up audio signal.
[0038] The inventor proposes to modulate the bit stream into the reference audio sound to form a special modulated signal, and use the special modulated signal to pre-tone to more accurately estimate the echo delay T between the reference audio signal and the picked-up audio signal. This not only simplifies the algorithm for echo delay estimation, but also obtains accurate echo delay estimation value in any delay length range, thereby improving the echo cancellation effect in a network environment with large delay.
[0039] Figure 2 A flowchart of an echo cancellation method according to the prior art is shown. The echo cancellation method includes steps S01 to S03 performed in the tone phase, and step S04 performed in the voice communication phase.
[0040] The user's voice interactive system includes, for example, a speaker, a microphone arranged in the same room, and an audio processing system arranged in or on the cloud. Preferably, the steps of the tone phase are performed when the user's voice interactive system is powered on, to detect audio parameters that change with the surrounding environment.
[0041] In step S01, the speaker is used to play an analog audio signal. The reference audio is, for example, a white noise signal, or a single-frequency high-frequency signal, for example, a frequency greater than the audible frequency range of the human ear. When the analog audio signal starts to be played, the start time T1 is recorded.
[0042] In step S02, the microphone is used to pick up the analog audio signal. The picked-up sound includes the sound directly reaching the microphone from the speaker's playing sound and the echo reaching the microphone after one or more reflections via the echo path.
[0043] In step S03, the echo delay is calculated by similarity calculation between the analog audio signal and the picked-up audio signal. For example, the similarity calculation is performed on the spectral energy of the analog audio signal and the picked-up audio signal. Alternatively, the picked-up audio signal of the microphone is subjected to a cyclic discrete Fourier transform (FFT), and when the frequency domain in the FFT calculation result contains the frequency value of the analog audio signal, it is considered that the picked-up audio signal contains the echo of the analog audio signal.
[0044] The arrival time T2 of the analog audio signal received by the microphone can be obtained by the similarity calculation. The echo delay t of the audio interaction system is represented as: t = T2 - T1.
[0045] In step S04, the echo delay is used as an audio parameter, and an adaptive filter is used to process the picked-up audio signal of the microphone to eliminate the echo. For example, the user voice audio signal is the difference between the picked-up audio signal of the microphone and the estimated value of the echo audio signal.
[0046] However, due to the time-varying characteristics of the voice signal and the random characteristics of the noise, there is a possibility of error in estimating the echo delay t based on the similarity calculation.
[0047] Referring to Figure 3 When the echo delay t is less than or equal to the adaptive filter length τ, as shown in curve b, the analog audio signal a and the picked-up audio signal b have correlation, and the echo delay t of the voice interaction system can be efficiently estimated by calculating the signal correlation. After removing the audio signal before the echo delay t, the picked-up audio signal is approximately aligned with the analog audio signal. Therefore, when the echo delay is less than or equal to the adaptive filter length τ, the voice interaction system can efficiently work to remove the echo.
[0048] Further, when the echo delay t is greater than the adaptive filter length τ, as shown in curve b', at this time, within the adaptive filter length τ, the analog audio signal a and the picked-up audio signal b' have no correlation, and the similarity calculation of the echo delay t of the voice interaction system will produce an error. Based on the estimated value of the error echo delay t, the picked-up audio signal cannot be processed to align with the analog audio signal. Therefore, when the echo delay is greater than the adaptive filter length τ, the voice interaction system cannot effectively remove the echo.
[0049] Figure 4 A flowchart of an echo cancellation method according to the first embodiment of the present disclosure is shown. The echo cancellation method includes steps S11 to S14 performed in the tuning phase, and step S15 performed in the voice communication phase.
[0050] A user's voice interaction system comprises for example loudspeakers, microphones arranged in the same room, and an audio processing system arranged in or at the cloud. Preferably, the steps of the tuning phase are performed at start-up of the user's voice interaction system to detect audio parameters varying with the surrounding environment.
[0051] In step Sll, digital frequency modulation parameters are set.
[0052] Digital frequency modulation is a form of modulation in which the variation of the carrier frequency is controlled by a baseband digital signal to transmit digital information. In the present embodiment, the digital frequency modulation parameters comprise at least one of the carrier frequency, the digital frequency, the baud rate of the baseband digital signal, and the data length.
[0053] In step S12, a reference audio signal is generated according to the digital frequency modulation parameters.
[0054] Digital frequency modulation comprises for example the frequency keying method and the direct frequency modulation method. In the frequency keying method, two independent oscillators generating sinusoidal oscillations are selected by an electronic switch controlled by the digital baseband signal. The selected high frequency oscillation signal is the digital modulated signal. In the direct frequency modulation method, the oscillation frequency of a carrier frequency oscillator is directly controlled by the digital baseband signal.
[0055] Referring to Fig. 2, in the digital frequency modulation method, a sinusoidal carrier wave is used. A predetermined check bit stream is used as the baseband digital signal to control the carrier frequency. Two frequencies Fcl and fc2 in the sinusoidal carrier wave are used to represent binary digits 1 and 0 respectively. The frequencies Fcl and fc2 are for example 2200 Hz and 1200 Hz respectively, and the bit rate is for example 1200 bps. Figure 6 Referring to Fig. 3, the check bit stream comprises for example synchronization data and check data. The synchronization data comprises for example a 600-bit synchronization flag, which is for example composed of alternating binary digits 1 and 0. The check data comprises for example 650 bits of binary digits, which comprises in turn a start flag, a message string, a check character, and an end flag.
[0056] Figure 7 Referring to Fig. 4, the check bit stream comprises for example synchronization data and check data. The synchronization data comprises for example a 600-bit synchronization flag, which is for example composed of alternating binary digits 1 and 0. The check data comprises for example 650 bits of binary digits, which comprises in turn a start flag, a message string, a check character, and an end flag.
[0057] In the test data, the start flag is composed of 200 consecutive binary digits "1", the end flag is composed of 40 consecutive binary digits "1", and the message string is composed of 400 consecutive binary digits, including 40 ASCII characters "0123456789". Each ASCII character occupies 10 bits, the start bit of each ASCII character is "1", the middle eight bits are information, and the end bit is "0". The check character is composed of 10 consecutive binary digits, in which the start bit is "1", the middle eight bits are the bit value of the check character, and the end bit is "0". When the sum of all data (including the check character) of the check bit stream modulo 256 is 00, it is proved that the received data is completely correct.
[0058] In step S13, the echo delay is obtained by testing with the reference audio signal.
[0059] In this step, the speaker plays the sound of the reference audio, the pickup signal of the microphone is collected to obtain the pickup audio signal, and the echo delay is calculated according to the time position of the check bit stream in the reference audio signal and the pickup audio signal.
[0060] In step S14, the evaluation parameter is calculated according to the test data, and it is judged whether the evaluation parameter is qualified.
[0061] In this step, the evaluation parameter includes at least one of the delay length, the bit error rate of the audio signal, and the signal-to-noise ratio.
[0062] In this step, if the evaluation parameter is not qualified, return to step S11, repeat steps S11 to S14, and reset the digital frequency modulation parameter to obtain a new echo delay and evaluation parameter calculation value.
[0063] In this step, if the evaluation parameter is qualified, the calculation value of the echo delay is taken as the measurement value, and step S15 is further executed.
[0064] In step S15, the echo delay is taken as the audio parameter, and the adaptive filter is used to process the pickup audio signal of the microphone to eliminate the echo. For example, the user voice audio signal is the difference between the pickup audio signal of the microphone and the estimated value of the echo audio signal.
[0065] In the embodiment, the digital frequency modulation signal is used as the reference audio signal, the reference audio signal is used for testing to obtain the echo delay, and the digital frequency modulation parameter is dynamically adjusted according to the evaluation parameter. On the one hand, the echo cancellation method utilizes the anti-interference performance and the anti-channel loss performance of the digital frequency modulation, and thus can be applied to a complex environment of a voice interaction system and improve the reliability. On the other hand, the echo cancellation method dynamically adjusts the digital frequency modulation parameter according to the evaluation parameter, and thus can improve the accuracy of the measurement value of the echo delay in the complex environment of the voice interaction system and improve the echo cancellation effect.
[0066] In the embodiment, the reference audio signal used for testing is a digital frequency modulation signal, and the evaluation parameter is obtained based on the testing data of the reference audio signal. In an alternative embodiment, the analog audio signal and the reference audio signal are combined, and the analog audio signal is used for testing to obtain the evaluation parameter, and the reference audio signal is further used for testing to obtain the evaluation parameter, according to the steps S01 to S03 shown in the figure. Figure 2
[0067] Figure 5 The detailed steps of obtaining the echo delay in the echo cancellation method shown in the figure. Figure 4
[0068] In step S21, the reference audio is played by the loudspeaker. As described above, the reference audio is, for example, a sine wave signal subjected to digital frequency modulation.
[0069] In the embodiment, the reference audio is, for example, an audio file generated and stored in advance, and the starting time T1 is obtained in the audio signal processing step described below. In an alternative embodiment, the reference audio is, for example, audio data generated in real time, and the starting time T1 is recorded when the reference audio starts to be played.
[0070] In step S22, the playback signal of the reference audio signal and the picked-up audio signal are obtained. For example, the driving signal in the driving circuit of the loudspeaker is collected to obtain the playback signal of the reference audio signal, and the picked-up signal of the microphone in the signal processing circuit of the microphone is collected to obtain the picked-up audio signal. The sound picked up by the microphone includes the sound directly reaching the microphone from the playback sound of the loudspeaker and the echo reaching the microphone after one or more reflections via the echo path.
[0071] In steps S231 and S232, the first bit stream A and the second bit stream B are demodulated from the reference audio signal and the picked-up audio signal, respectively.
[0072] The circuit structure and working principle of the demodulator for coherent demodulation of the modulated signal are known. See, for example, the patent application CN201110339593.7. Figure 8 The demodulator 100 comprises band-pass filters 111 and 112, multipliers 113 and 114, low-pass filters 115 and 116, and a sample decision device 118. The center frequency fcl of the band-pass filter 111 corresponds to binary digit 1, and the center frequency fc2 of the band-pass filter 112 corresponds to binary digit 0. The band-pass filters 111 and 112 divide the modulated signal into two signals, a first signal corresponding to binary digit 1 and a second signal corresponding to binary digit 0. The multiplier 113 multiplies the first signal with the coherent reference signal, and the low-pass filter 115 extracts the time-varying amplitude and phase of the first signal. The multiplier 114 multiplies the second signal with the coherent reference signal, and the low-pass filter 116 extracts the time-varying amplitude and phase of the second signal. The sample decision device 118 obtains the sample signals of the first signal amplitude and the second signal amplitude at the same phase, and compares the first signal amplitude and the second signal amplitude to determine the value of the binary digit at the corresponding phase.
[0073] In step S231, the first bit stream A is demodulated from the reference audio signal. In step S241, the second bit stream B is demodulated from the picked-up audio signal.
[0074] In the present embodiment, the reference audio signal and the picked-up audio signal are both analog signals collected in real time. Due to the delay and signal distortion of the signal processing circuit, and the influence of factors such as the echo path difference of the environment and environmental noise interference, the first bit stream A demodulated from the reference audio signal is not completely consistent with the second bit stream B demodulated from the picked-up audio signal, however, both the first bit stream A and the second bit stream B contain the check bit stream V.
[0075] In step S241, the time position of the first bit stream A in the reference audio signal is obtained according to the similarity of the first bit stream A and the check bit stream. In step S242, the time position of the second bit stream B in the picked-up audio signal is obtained according to the similarity of the second bit stream B and the check bit stream.
[0076] The similarity of the first bit stream A and the check bit stream V is calculated, the time position of the first bit stream A in the reference audio signal under the most similar condition is obtained, and the starting time Tl is obtained.
[0077] The similarity of the second bit stream B and the check bit stream V is calculated, the time position of the second bit stream B in the picked-up audio signal under the most similar condition is obtained, and the arrival time T2 is obtained.
[0078] In step S25, the echo delay is calculated according to the difference between the first time position and the second time position.
[0079] In this step, the echo delay t of the audio interaction system is represented as: t = T2 - Tl.
[0080] In this embodiment, digital frequency modulation is employed to modulate the parity bitstream into the reference audio sound, forming a special modulation signal. During the echo cancellation tuning stage, the playback signal and the picked-up audio signal of the reference audio signal are acquired. After demodulation, the time position of the parity bitstream in the playback signal and the picked-up audio signal of the reference audio signal are obtained, thereby allowing the calculation of the echo delay of the audio interaction system. Due to the anti-interference and anti-channel loss performance of digital frequency modulation, this echo cancellation method can be applied to the complex environment of voice interaction systems and improve reliability.
[0081] Furthermore, the demodulator in this echo cancellation method primarily performs multiplication calculations, eliminating the need for similarity calculations of the audio signal's spectral energy or Discrete Fourier Transform (FFT). This simplifies the echo delay algorithm. The longer the actual delay, the greater the computational simplification.
[0082] Furthermore, the echo delay is calculated by referencing the time position of the audio signal and the parity bitstream of the picked-up audio signal. The accuracy of this time position depends on the bit rate of the parity bitstream; therefore, the time accuracy of the echo delay also depends on the bit rate of the parity bitstream. At a bit rate of, for example, 1200 bps, the time accuracy is approximately 0.84 ms (1000 ms / 1200 bit). Higher baud rates can achieve even higher time accuracy. Therefore, this echo cancellation method can improve the time accuracy of echo delay calculation. For audio interaction systems with small delays, this echo cancellation method can also calculate accurate echo delays, thereby improving the echo cancellation effect.
[0083] Furthermore, after calculating the accurate echo delay, the length τ of the adaptive filter can be significantly reduced, thereby lowering the difficulty of adaptation and decreasing the computational load. For network environments with large delays, this echo cancellation method can also calculate the accurate echo delay, thus improving the echo cancellation effect.
[0084] Figure 9 Show Figure 4 The detailed steps of dynamically adjusting digital frequency modulation parameters in the echo cancellation method are shown. In the step of dynamically adjusting digital frequency modulation parameters, multiple modulation parameters are adjusted based on the values of multiple evaluation parameters.
[0085] In step S31, a reference audio signal modulated with a calibrated bitstream is used for testing to calculate the echo delay t1 (see...). Figure 5 Steps S21 to S25 are shown.
[0086] In step S32, the calculated echo delay t1 is compared with a preset time threshold Thrt to determine whether the echo delay t1 is too long.
[0087] If the echo delay t1 is greater than or equal to the time threshold Tht, step S33 is performed, in which the time length of the calibration bit stream V is increased. For example, the method for increasing the time length of the calibration bit stream V includes increasing the data length of the calibration bit stream V, and / or decreasing the baud rate of the modulation signal. Then, steps S31 and S32 are repeated, and the echo delay t2 is recalculated and determined whether it is too long.
[0088] If the echo delay t1 or t2 is less than the time threshold Tht, step S34 is continued.
[0089] In step S34, the error rate e1 is calculated according to the test data, and the calculated error rate e1 is compared with the preset error rate threshold The to determine whether the error rate e1 is too high.
[0090] In this step, the reference audio signal modulated by the calibration bit stream is used for testing. After the time position of the first bit stream in the reference audio signal and the time position of the second bit stream in the picked-up audio signal are obtained, the first bit stream A and the second bit stream B are obtained based on the time positions, respectively. Similarity calculation is performed on the first bit stream A and the second bit stream B to obtain the error rate e1 of the audio signal.
[0091] If the error rate e1 is greater than or equal to the time threshold The, step S35 is performed, in which the baud rate of the modulation signal is decreased. Then, steps S31 and S34 are repeated, and the error rate e2 is recalculated and determined whether it is too high.
[0092] If the error rate e1 or e2 is less than the time threshold The, step S36 is continued.
[0093] In step S36, the signal-to-noise ratio s1 is calculated according to the test data, and the calculated signal-to-noise ratio s1 is compared with the preset signal-to-noise ratio threshold Ths to determine whether the signal-to-noise ratio s1 is too high.
[0094] In this step, the reference audio signal modulated by the calibration bit stream is used for testing. After the picked-up audio signal is obtained, the ratio of the modulation signal strength and the blank signal strength of the picked-up audio signal is calculated to obtain the signal-to-noise ratio s1 of the audio signal.
[0095] If the signal-to-noise ratio s1 is less than the signal-to-noise ratio threshold Ths, step S37 is performed, in which at least one of the carrier frequency and the digital frequency is decreased. Then, steps S31 and S36 are repeated, and the signal-to-noise ratio s2 is recalculated and determined whether it is too low.
[0096] If the signal-to-noise ratio s1 or s2 is greater than or equal to the signal-to-noise ratio threshold Ths, step S38 is continued.
[0097] In step S38, the calculated value of the echo delay is saved as an audio parameter of the echo cancellation algorithm.
[0098] In the embodiment, the detailed steps of dynamically adjusting the digital frequency modulation parameters according to the evaluation parameters of the audio signal in the tuning phase are described, wherein the evaluation parameters include the time length of the echo delay, the bit error rate and the signal-to-noise ratio of the audio signal, and the modulation parameters include the time length of the calibration bit stream, the baud rate of the modulation signal, the carrier frequency and the digital frequency. However, the present disclosure is not limited thereto. According to the complexity of the environment of the audio interaction system, the evaluation parameters can include one or more of the time length of the echo delay, the bit error rate and the signal-to-noise ratio of the audio signal. Further, the evaluation parameters can also include additional parameters, such as the reverberation time of the audio signal.
[0099] The embodiment of the present disclosure also provides an electronic device 1300, as shown in the figure, comprising a memory 1310 and a processor 1320, and a program stored in the memory 1310 and executable on the processor 1320, which, when executed by the processor 1320, can implement the processes of each of the embodiments of the above-mentioned echo cancellation method and achieve the same technical effects. To avoid repetition, details are not described here. Of course, the electronic device can also include a power supply component 1330, a network interface 1340 and an input / output interface 1350 and other auxiliary sub-devices. Figure 10
[0100] Those skilled in the art can understand that all or part of the steps of the various methods of the above-mentioned embodiments can be completed by instructions, or by related hardware controlled by the instructions, which can be stored in a computer-readable readable storage medium and loaded and executed by a processor. For this purpose, the embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program or instructions, which, when executed by a processor, can implement the processes of each of the embodiments of the above-mentioned echo cancellation method. The computer-readable storage medium, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk and various media that can store program codes.
[0101] Due to the instructions stored in the readable storage medium, the steps in any of the echo cancellation methods provided by the embodiments of the present disclosure can be executed, so that the beneficial effects of any of the echo cancellation methods provided by the embodiments of the present disclosure can be achieved. Details are described in the previous embodiments, which are not described here. The specific implementation of each operation can be referred to the previous embodiments, which are not described here.
[0102] It should be noted that in describing the various embodiments in this specification, the focus is on the differences from other embodiments, while the same or similar parts between the various embodiments can be understood by referring to each other. For the system embodiments, since they are basically similar to the method embodiments, the relevant parts can be referred to the description of the method embodiments.
[0103] Furthermore, it should be noted that in the apparatus and method of this disclosure, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of this disclosure. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of this disclosure can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof, which can be achieved by those skilled in the art using their basic programming skills after reading the description of this disclosure.
[0104] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating this disclosure and are not intended to limit the implementation. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of this disclosure.
Claims
1. An echo cancellation method for a voice interaction system, comprising: tuning a voice interaction system to obtain an echo delay; and processing a picked-up audio signal in the voice interaction system with the echo delay to remove echo, wherein the tuning comprises: generating a reference audio signal by digitally frequency modulating with a check bit stream, the digitally frequency modulating comprising controlling a carrier frequency with the check bit stream as a baseband digital signal, the carrier having a first frequency when the check bit stream is a digital one and a second frequency when the check bit stream is a digital zero; testing with the reference audio signal to obtain the echo delay according to a time position of the check bit stream in the reference audio signal and the picked-up audio signal; and dynamically adjusting digitally frequency modulating parameters according to an evaluation parameter and repeating the testing step until the evaluation parameter is qualified, wherein if the evaluation parameter is not qualified, repeating the testing step, resetting the digitally frequency modulating parameters to obtain a new echo delay and an evaluation parameter calculation value, and if the evaluation parameter is qualified, taking the calculation value of the echo delay as a measurement value. obtaining the time position of the check bit stream in the reference audio signal comprises: demodulating the reference audio signal based on multiplication calculation to obtain a first bit stream, calculating a similarity of the first bit stream with the check bit stream, and obtaining a time position of the first bit stream in the reference audio signal under a most similar condition; obtaining the time position of the check bit stream in the picked-up audio signal comprises: demodulating the picked-up audio signal based on multiplication calculation to obtain a second bit stream, calculating a similarity of the second bit stream with the check bit stream, and obtaining a time position of the second bit stream in the reference audio signal under a most similar condition; the evaluation parameter comprises at least one of a time length of the echo delay, a bit error rate of the audio signal, a signal-to-noise ratio of the audio signal, and a reverberation time of the audio signal, and the digitally frequency modulating parameter comprises at least one of a time length of the check bit stream, a baud rate of a modulating signal, a carrier frequency, and a digital frequency, and dynamically adjusting the digitally frequency modulating parameter according to the evaluation parameter comprises: comparing the calculated echo delay with a preset time threshold, and increasing the time length of the check bit stream if the echo delay is greater than or equal to the preset time threshold; calculating a bit error rate according to test data if the echo delay is less than the preset time threshold, comparing the calculated bit error rate with a preset bit error rate threshold, and decreasing the baud rate of the modulating signal if the bit error rate is greater than or equal to the preset bit error rate threshold; calculating a signal-to-noise ratio according to test data if the bit error rate is less than the preset bit error rate threshold, comparing the calculated signal-to-noise ratio with a preset signal-to-noise ratio threshold, and decreasing at least one of the carrier frequency and the digital frequency if the signal-to-noise ratio is less than the preset signal-to-noise ratio threshold.
2. The echo cancellation method of claim 1, further comprising: In a driving circuit of the loudspeaker, a driving signal of the loudspeaker is collected to obtain the reference audio signal; and In a signal processing circuit of the microphone, a picked-up signal of the microphone is collected to obtain the picked-up audio signal.
3. The echo cancellation method of claim 1, wherein, The step of obtaining the echo delay comprises: demodulating the reference audio signal to obtain a time position of the check bit stream in the reference audio signal as a start time; demodulating the picked-up audio signal to obtain a time position of the check bit stream in the picked-up audio signal as an arrival time; and obtaining the echo delay according to a difference between the start time and the arrival time.
4. The echo cancellation method of claim 1, wherein, The step of obtaining the echo delay comprises: estimating a start time according to a time when the loudspeaker plays real-time generated audio data; demodulating the picked-up audio signal to obtain a time position of the check bit stream in the picked-up audio signal as an arrival time; and obtaining the echo delay according to a difference between the start time and the arrival time.
5. The echo cancellation method of claim 1, wherein, Increasing the time length of the calibration bit stream comprises: increasing a data length of the calibration bit stream, and / or, decreasing a baud rate of a modulating signal.
6. The echo cancellation method of claim 1, wherein, Performing a similarity calculation on the first bit stream and the second bit stream to obtain an error code rate of the audio signal.
7. The echo cancellation method of claim 1, wherein, Calculating a ratio of a modulating signal strength and a blank signal strength of the picked-up audio signal to obtain a signal-to-noise ratio of the audio signal.
8. The echo cancellation method of claim 1, before generating the reference audio signal, the tuning further comprises: adopting an analog audio signal to test to obtain the echo delay; and dynamically tuning digital frequency modulation parameters according to evaluation parameters. comprise:
9. An electronic device, comprising: a processor, a memory, and a program stored on the memory and executable on the processor, the program being executed by the processor to implement steps of the method of any one of claims 1-8. The computer readable storage medium stores a computer program or instructions, the computer program or instructions being executed by the processor to implement steps of the method of any one of claims 1-8.
10. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Echo delay determination method and device, equipment and storage medium
CN113707160A