Method for Filtering Echo, Electronic Device, and Computer-Readable Storage Medium
By applying direct sound filtering and reverse speaker technology in electronic devices, the problem of speaker echo being collected by microphones is solved, the voice assistant wake-up and call quality is improved, and the voice assistant is adapted to a variety of environments.
Patent Information
- Application Number
- CN202010707669.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-07-21
AI Technical Summary
When existing electronic devices play sounds with speakers, the echoes are easily collected by the microphone, which affects the wake-up and call quality of the voice assistant. The conventional acoustic echo cancellation algorithm is not ideal.
Direct sound filtering technology is adopted to obtain the signals of the speaker and microphone and perform direct sound filtering to filter out echoes. Custom direct sound filtering models and reverse speaker technology can be used to cancel echoes.
It improves the wake-up success rate of voice assistant and the quality of voice calls, enhances the echo filtering effect, and adapts to different environmental conditions.
Smart Images

Figure CN113963712B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device, and more particularly to a method for filtering echoes, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the improvement of consumers' requirements for operation experience and voice interaction, there are more and more electronic devices including voice intelligent assistants and call functions, such as smart screens, smart speakers, smart robots, in-vehicle voice assistants, smartphones, tablets, etc. However, the echoes of the sounds played by the speakers of the electronic devices are usually collected by the microphones, which affects the wake-up engine of the intelligent assistant and / or the call quality. The sound played by the speaker can reach the microphone of the electronic device in two ways. One way is through environmental reflection such as a wall, and the formed reflected echo can be collected by the microphone. Another way is that the sounds played by multiple speakers are directly transmitted to the microphone of the electronic device as echoes without any reflection.
[0003] In order to be able to wake up the voice assistant (enhanced voice interaction) or be able to make a voice call when the speaker of the electronic device plays music or a TV program, these electronic devices usually use an acoustic echo cancellation (AEC) algorithm to cancel the signal components associated with the echoes of the sounds played by the speaker in the audio signals collected by the microphone. However, in some cases, the conventional AEC algorithm is still not ideal for canceling the echoes played by the speaker, so further improvement is needed. Summary of the Invention
[0004] In view of the above problems, embodiments of the present disclosure aim to provide a technique for filtering echoes.
[0005] According to a first aspect of the present disclosure, a method for filtering echoes is provided. The method is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The method includes obtaining N speaker signals corresponding to the N speakers; obtaining M microphone signals corresponding to the M microphones; and performing at least direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal, where direct sound filtering indicates filtering the audio components directly output from the N speakers to the M microphones without environmental reflection. By using direct sound filtering, the echo filtering effect can be further improved.
[0006] In some embodiments, the target signal is used to wake up the engine to wake up the intelligent voice assistant or is transmitted to another electronic device for a voice call. In some embodiments, the target signal includes fewer echo components than the M microphone signals, and the echo components are used to characterize the echoes of the sound propagated in space of the N speaker signals collected by the M microphones. By using the target signal filtered by the direct sound, the success rate of waking up the intelligent voice assistant by the wake-up engine and / or the quality of the voice call can be improved.
[0007] In some embodiments, the method further includes causing a display screen of the electronic device to display a customized direct sound filtering interface; receiving user input on the customized direct sound filtering interface; in response to the user input, obtaining N speaker test signals and causing the N speakers to play the N speaker test signals; obtaining M microphone test signals corresponding to the M microphones; and storing a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the method further includes using the customized direct sound filtering model to perform customized direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal. By using customized direct sound filtering, the direct sound filtering can be optimized for the customer's environment and the echo cancellation in the customer's environment can be further improved.
[0008] In some embodiments, the customized direct sound filtering interface displays a prompt item for keeping the environment quiet. The customized direct sound filtering interface may further display a decibel indication item indicating the ambient noise and / or an indication item indicating whether it is suitable for customized echo filtering. By displaying the prompt item for indicating a quiet environment, the decibel indication item, and / or the indication item indicating whether it is suitable for customized echo filtering, the user can be prompted to establish a customized direct sound filtering model in a quiet and suitable environment, thereby laying a good foundation for subsequent use of the customized direct sound filtering specific to this environment and subsequently obtaining a good echo cancellation effect in this environment. In some embodiments, the direct sound filtering includes default direct sound filtering, and the default direct sound filtering indicates filtering based at least on the model relationship between the N speaker signals played by the N speakers in an anechoic environment and the M microphone signals directly collected by the M microphones.
[0009] In some embodiments, the method further includes generating a reverse speaker signal based on the N speaker signals; and causing a reverse speaker close to at least one of the M microphones to play reverse audio based on the reverse speaker signal to cancel the echo of the audio output corresponding to the N speaker signals played by the N speakers. The reverse speaker is different from the N speakers. By causing the reverse speaker to play reverse audio, a part of the echo components can be filtered before the echo is collected by the microphone, thereby providing an echo cancellation effect.
[0010] In some embodiments, the method further includes generating an echo estimation signal based on N speaker signals; filtering the echo estimation signal from the M microphone signals to generate a residual signal; and obtaining a target signal including performing direct-path filtering on the residual signal to obtain the target signal. By preprocessing the microphone signals to filter the echo estimation signal before direct-path filtering, the echo filtering effect can be further improved.
[0011] In some embodiments, generating the echo estimation signal includes performing non-interleaved preprocessing on the N speaker signals to generate at least one preprocessing signal; and performing adaptive filtering on the at least one preprocessing signal to generate the echo estimation signal. By performing non-interleaved preprocessing on the N speaker signals, a preprocessing signal that continuously represents the echo in time can be obtained, and by using an adaptive filtering signal based on the non-interleaved preprocessing signal to estimate the echo component in the M microphone signals, a better echo filtering effect can be obtained.
[0012] In some embodiments, generating at least one preprocessing signal includes linearly summing at least two of the N speaker signals to generate a sum signal. By combining at least two of the N speaker signals into a single sum signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and filtering the echo in the full frequency band to obtain a better echo filtering effect.
[0013] In some embodiments, generating at least one preprocessing signal further includes linearly subtracting at least two of the N speaker signals to generate a difference signal. In some cases, echo filtering focuses on the difference between a certain frequency band or the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo filtering effect.
[0014] In some embodiments, generating at least one preprocessing signal further includes sorting the sum signal and the difference signal; generating the echo estimation signal further includes performing adaptive filtering on the sorted sum signal and difference signal in sequence to generate the echo estimation signal corresponding to the sorting; and generating the residual signal includes filtering the corresponding sorted echo estimation signal from the M microphone signals in sequence to generate the residual signal. By sorting the sum signal and the difference signal and generating the residual signal accordingly, a better echo filtering effect can be obtained for different cases.
[0015] In some embodiments, generating at least one preprocessing signal includes generating N sorted preprocessing signals by sorting N speaker signals. Generating an echo estimation signal includes adaptively filtering the N sorted preprocessing signals in sequence to generate N echo estimation signals corresponding to the sorting. Generating a residual signal includes sequentially filtering the N echo estimation signals corresponding to the sorting from M microphone signals to generate a residual signal. By sorting the N speaker signals and filtering them in sequence, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0016] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components of each of the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo filtering effect can be improved.
[0017] In some embodiments, generating at least one preprocessing signal includes performing non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessing signal. By using the M microphone signals as auxiliary reference signals, the frequency bands with greater echo can be restricted to assist in improving echo filtering.
[0018] In some embodiments, the method further includes adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessing signal so that the gain of the echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0019] According to a second aspect of the present disclosure, a method for filtering echo is provided. The method is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The method includes obtaining N speaker signals corresponding to the N speakers; generating reverse speaker signals based on the N speaker signals; causing a reverse speaker near at least one of the M microphones to play reverse audio based on the reverse speaker signals to cancel the audio output corresponding to the N speaker signals played by the N speakers. The reverse speaker is different from the N speakers. By causing the reverse speaker to play reverse audio, a portion of the echo components can be filtered before the echo is collected by the microphone, thereby providing an echo filtering effect.
[0020] In some embodiments, the method further includes obtaining M microphone signals corresponding to M microphones; performing at least direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal, where direct sound filtering indicates filtering of the audio components that are directly output from the N speakers to the M microphones without environmental reflection. By using direct sound filtering, the effect of echo filtering can be further improved.
[0021] In some embodiments, the target signal is used to wake up the engine to wake up the intelligent voice assistant or is transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, where the echo components are used to characterize the echoes of the audio outputs of the N speaker signals propagating in space as collected by the M microphones. By using the target signal filtered by direct sound, the success rate of waking up the intelligent voice assistant by the wake-up engine and / or the quality of the voice call can be improved.
[0022] In some embodiments, the direct sound filtering includes default direct sound filtering, where the default direct sound filtering indicates filtering based at least on the model relationship between the N speaker signals played by the N speakers and the M microphone signals directly collected by the M microphones in an anechoic environment.
[0023] In some embodiments, the method further includes causing a display screen of the electronic device to display a customized direct sound filtering interface; receiving user input on the customized direct sound filtering interface; in response to the user input, obtaining N speaker test signals and causing the N speakers to play the N speaker test signals; obtaining M microphone test signals corresponding to the M microphones; and storing a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the method further includes performing customized direct sound filtering on the N speaker signals and the M microphone signals using the customized direct sound filtering model to obtain a target signal. By using customized direct sound filtering, the direct sound filtering can be optimized for the customer's environment and the echo cancellation in the customer's environment can be further improved.
[0024] In some embodiments, the customized direct sound filtering interface displays a prompt item for keeping the environment quiet. The customized direct sound filtering interface may also display a decibel indication item indicating the ambient noise and / or an indication item indicating whether it is suitable for customized echo filtering. By displaying the prompt item for keeping the environment quiet, the decibel indication item, and / or the indication item indicating whether it is suitable for customized echo filtering, the user can be prompted to establish a customized direct sound filtering model in a quiet and suitable environment, thereby laying a good foundation for subsequent use of the customized direct sound filtering specific to this environment and obtaining a good echo cancellation effect in this environment subsequently.
[0025] In some embodiments, the method further includes generating an echo estimation signal based on N speaker signals; filtering the echo estimation signal from the M microphone signals to generate a residual signal; and obtaining a target signal including performing direct sound filtering on the residual signal to obtain the target signal. By preprocessing the microphone signals to filter the echo estimation signal before direct sound filtering, the effect of echo filtering can be further improved.
[0026] In some embodiments, generating the echo estimation signal includes performing non-interleaved preprocessing on the N speaker signals to generate at least one preprocessing signal; and performing adaptive filtering on the at least one preprocessing signal to generate the echo estimation signal. By performing non-interleaved preprocessing on the N speaker signals, a preprocessing signal that continuously represents the echo in time can be obtained, and by using an adaptive filtering signal based on the non-interleaved preprocessing signal to estimate the echo component in the M microphone signals, a better echo filtering effect can be obtained.
[0027] In some embodiments, generating at least one preprocessing signal includes linearly summing at least two of the N speaker signals to generate a sum signal. By combining at least two of the N speaker signals into a single sum signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and filtering the echo in the full frequency band to obtain a better echo filtering effect.
[0028] In some embodiments, generating at least one preprocessing signal further includes linearly differencing at least two of the N speaker signals to generate a difference signal. In some cases, echo filtering focuses on the difference in a certain frequency band or between the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo filtering effect.
[0029] In some embodiments, generating at least one preprocessing signal further includes sorting the sum signal and the difference signal. Generating the echo estimation signal further includes sequentially performing adaptive filtering on the sorted sum signal and difference signal to generate corresponding sorted echo estimation signals. Generating the residual signal includes sequentially filtering the corresponding sorted echo estimation signals from the M microphone signals to generate the residual signal. By sorting the sum signal and the difference signal and generating the residual signal accordingly, a better echo filtering effect can be obtained for different cases.
[0030] In some embodiments, generating at least one preprocessed signal includes generating N sorted preprocessed signals by sorting N speaker signals. Generating an echo estimation signal includes adaptively filtering the N sorted preprocessed signals in sequence to generate N sorted echo estimation signals corresponding to the sorting. Generating a residual signal includes sequentially filtering the N echo estimation signals from M microphone signals to generate a residual signal. By sorting the N speaker signals and filtering them in sequence, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0031] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components of each speaker signal among the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo filtering effect can be improved.
[0032] In some embodiments, generating at least one preprocessed signal includes performing non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessed signal. By using the M microphone signals as auxiliary reference signals, the frequency bands with larger echoes can be restricted to assist in improving echo filtering.
[0033] In some embodiments, the method further includes adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessed signal so that the gain of the echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0034] According to a third aspect of the present disclosure, a method for filtering echoes is provided. The method is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The method includes obtaining N speaker signals corresponding to the N speakers; obtaining M microphone signals corresponding to the M microphones; performing non-interleaved preprocessing on the N speaker signals to generate M groups of preprocessed signals; performing adaptive filtering on the M groups of preprocessed signals to generate M echo estimation signals, and filtering the M echo estimation signals from the M microphone signals to obtain a residual signal. By performing non-interleaved preprocessing on the N speaker signals, preprocessed signals that continuously represent echoes in time can be obtained, and by using adaptive filtering signals based on the non-interleaved preprocessed signals to estimate the echo components in the M microphone signals, a better echo filtering effect can be obtained.
[0035] In some embodiments, the residual signal is the target signal. The target signal is used to wake up the engine to wake up the intelligent voice assistant or to be transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, and the echo components are used to characterize the echoes of the sounds propagated in space by the N speaker signals collected by the M microphones. By using the target signal based on the preprocessed signal after non-interleaved preprocessing, the success rate of the wake-up engine in waking up the intelligent voice assistant and / or the quality of the voice call can be improved.
[0036] In some embodiments, generating at least one preprocessed signal includes linearly summing at least two of the N speaker signals to generate a sum signal. By combining at least two of the N speaker signals into a single sum signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and filtering out echoes in the full frequency band to obtain a better echo filtering effect.
[0037] In some embodiments, generating at least one preprocessed signal includes linearly subtracting at least two of the N speaker signals to generate a difference signal. In some cases, echo filtering focuses on a certain frequency band or the differences between the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo filtering effect.
[0038] In some embodiments, generating at least one preprocessed signal further includes sorting the sum signal and the difference signal. Generating at least one echo estimation signal includes sequentially performing adaptive filtering on the sorted sum signal and difference signal to sequentially generate the corresponding sorted adaptive filtering signals. Generating the residual signal includes generating the residual signal by sequentially filtering out the corresponding sorted adaptive filtering signals from the M microphone signals. By sorting the sum signal and the difference signal and filtering them sequentially, signals that cause greater distortion can be preferentially filtered out to provide an echo filtering effect.
[0039] In some embodiments, generating at least one preprocessed signal includes generating N preprocessed signals in sorted order by sorting the N speaker signals. Generating at least one echo estimation signal includes sequentially performing adaptive filtering on the N preprocessed signals in sorted order to generate the corresponding sorted echo estimation signals. Generating the target signal includes generating the target signal by sequentially filtering out the corresponding sorted echo estimation signals from the M microphone signals. By sorting the N speaker signals and filtering them sequentially, signals that cause greater distortion can be preferentially filtered out to provide an echo filtering effect.
[0040] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components respectively possessed by each of the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo cancellation effect can be improved.
[0041] In some embodiments, generating at least one preprocessed signal includes performing non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessed signal. By using the M microphone signals as auxiliary reference signals, the frequency bands with larger echoes can be restricted to assist in improving echo cancellation.
[0042] In some embodiments, the method further includes adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessed signal such that the gain of at least one echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0043] In some embodiments, generating at least one echo estimation signal includes, in the case of multiple preprocessed signals, performing parallel adaptive filtering on the multiple preprocessed signals to generate at least one echo estimation signal. For multiple preprocessed signals with low similarity, parallel preprocessing can improve the echo cancellation effect.
[0044] In some embodiments, generating at least one echo estimation signal includes converting at least one preprocessed signal from a time-domain signal to a frequency-domain signal, and performing adaptive filtering on the frequency-domain signal to obtain at least one echo estimation signal. By converting the preprocessed signal from the time domain to the frequency domain, the computational overhead of adaptive filtering can be reduced.
[0045] In some embodiments, the method includes performing direct sound filtering on the residual signal to generate a target signal. Direct sound filtering indicates filtering the audio components that are directly output from the N speakers to the M microphones without environmental reflection. More specifically, direct sound filtering indicates filtering based on the model relationship between the sound source signal played by at least one of the first speaker and the second speaker and the audio input signal directly collected by the microphone. By using direct sound filtering, the echo filtering effect can be further improved.
[0046] In some embodiments, generating the target signal includes performing default direct sound filtering on the residual signal to generate the target signal. Default direct sound filtering indicates filtering based on the model relationship between the sound source signal played by the N speakers in an anechoic environment and the audio input signal directly collected by the M microphones.
[0047] In some embodiments, the method further includes causing a display screen of an electronic device to display a customized direct sound filtering interface; receiving user input on the customized direct sound filtering interface; in response to the user input, obtaining N speaker test signals and causing N speakers to play the N speaker test signals; obtaining M microphone test signals corresponding to M microphones; and storing a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the method further includes using the customized direct sound filtering model to perform customized direct sound filtering on the N speaker signals and the M microphone signals to obtain target signals. By using customized direct sound filtering, direct sound filtering can be optimized for the customer's environment and echo cancellation in the customer's environment can be further improved.
[0048] In some embodiments, the method further includes generating a reverse audio signal based on the N speaker signals, and causing a reverse speaker near the microphone to play reverse audio based on the reverse audio signal to cancel the audio output played by at least one of the N speakers. The reverse speaker is different from the first speaker and the second speaker. By causing the reverse speaker to play reverse audio, a part of the echo component can be filtered out before the echo is collected by the microphone, thereby providing an echo cancellation effect.
[0049] According to a fourth aspect of the present disclosure, there is provided an electronic device. The electronic device includes N speakers, M microphones, one or more processors; and a memory storing one or more programs. The one or more programs are configured to be executed by the one or more processors. The one or more programs include instructions for performing the method according to the first aspect.
[0050] According to a fifth aspect of the present disclosure, there is provided an electronic device. The electronic device includes N speakers, M microphones, one or more processors; and a memory storing one or more programs. The one or more programs are configured to be executed by the one or more processors. The one or more programs include instructions for performing the method according to the second aspect.
[0051] According to a sixth aspect of the present disclosure, there is provided an electronic device. The electronic device includes N speakers, M microphone reverse speakers, one or more processors; and a memory storing one or more programs. The one or more programs are configured to be executed by the one or more processors. The one or more programs include instructions for performing the method according to the third aspect.
[0052] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs are configured to be executed by one or more processors of an electronic device. The one or more programs include instructions for performing the method according to the first aspect.
[0053] According to an eighth aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs are configured to be executed by one or more processors of an electronic device. The one or more programs include instructions for performing the method according to the second aspect.
[0054] According to a ninth aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs are configured to be executed by one or more processors of an electronic device. The one or more programs include instructions for performing the method according to the third aspect.
[0055] According to a tenth aspect of the present disclosure, there is provided a device for echo cancellation. The device is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The device includes: a first acquisition module for acquiring N speaker signals corresponding to the N speakers; a second acquisition module for acquiring M microphone signals corresponding to the M microphones; a direct sound filtering module for performing at least direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal, where direct sound filtering refers to filtering the audio components that are directly output from the N speakers to the M microphones without environmental reflection. By using direct sound filtering, the effect of echo cancellation can be further improved.
[0056] In some embodiments, the target signal is used to wake up the engine to wake up the intelligent voice assistant or is transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, and the echo components are used to characterize the echoes of the sounds propagated in space of the N speaker signals collected by the M microphones. By using the target signal filtered by direct sound, the success rate of waking up the intelligent voice assistant by the wake-up engine and / or the quality of the voice call can be improved.
[0057] In some embodiments, the device further includes: a display enabling module for enabling a display screen of an electronic device to display a customized direct sound filtering interface; an input receiving module for receiving user input on the customized direct sound filtering interface; a speaker test module for obtaining N speaker test signals in response to the user input and causing N speakers to play the N speaker test signals; a third obtaining module for obtaining M microphone test signals corresponding to M microphones; and a storage module for storing a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the device further includes a customized direct sound filtering module for performing customized direct sound filtering on the N speaker signals and the M microphone signals using the customized direct sound filtering model to obtain a target signal. By using customized direct sound filtering, the direct sound filtering can be optimized for the customer's environment and the echo cancellation in the customer's environment can be further improved.
[0058] In some embodiments, the customized direct sound filtering interface displays a prompt item for prompting to keep the environment quiet. The customized direct sound filtering interface may further display a decibel indication item indicating the ambient noise and / or an indication item indicating whether it is suitable for customized echo filtering. By displaying the prompt item for prompting the environment to be quiet, the decibel indication item, and / or the indication item indicating whether it is suitable for customized echo filtering, the user can be prompted to establish a customized direct sound filtering model in a quiet and suitable environment, thereby laying a good foundation for subsequent use of the customized direct sound filtering specific to this environment and subsequently obtaining a good echo cancellation effect in this environment. In some embodiments, the direct sound filtering includes default direct sound filtering, and the default direct sound filtering indicates filtering based at least on the model relationship between the N speaker signals played by the N speakers and the M microphone signals directly collected by the M microphones in an anechoic environment.
[0059] In some embodiments, the device further includes a reverse speaker signal generating module for generating a reverse speaker signal based on the N speaker signals; and a playback enabling module for causing a reverse speaker near at least one of the M microphones to play reverse audio based on the reverse speaker signal to cancel the echo of the audio output corresponding to the N speaker signals played by the N speakers. The reverse speaker is different from the N speakers. By causing the reverse speaker to play reverse audio, a part of the echo component can be filtered before the echo is collected by the microphone, thereby providing an echo cancellation effect.
[0060] In some embodiments, the apparatus further includes: an echo estimation module for generating an echo estimation signal based on N speaker signals; a residual signal generation module for filtering the echo estimation signal from M microphone signals to generate a residual signal; and a target signal generation module for performing direct-path filtering on the residual signal to obtain a target signal. By preprocessing the microphone signals to filter the echo estimation signal before direct-path filtering, the echo filtering effect can be further improved.
[0061] In some embodiments, the echo estimation module includes: a preprocessing signal generation module for performing non-interleaved preprocessing on N speaker signals to generate at least one preprocessing signal; and an adaptive filtering module for performing adaptive filtering on the at least one preprocessing signal to generate an echo estimation signal. By performing non-interleaved preprocessing on N speaker signals, a preprocessing signal that continuously represents the echo in time can be obtained, and by using an adaptive filtering signal based on the non-interleaved preprocessing signal to estimate the echo component in M microphone signals, a better echo filtering effect can be obtained.
[0062] In some embodiments, the preprocessing signal generation module includes: a summation module for linearly summing at least two of the N speaker signals to generate a summation signal. By combining at least two of the N speaker signals into a single summation signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and filtering the echo in the full frequency band to obtain a better echo filtering effect.
[0063] In some embodiments, the preprocessing signal generation module further includes: a difference module for linearly differencing at least two of the N speaker signals to generate a difference signal. In some cases, echo filtering focuses on the difference in a certain frequency band or between the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo filtering effect.
[0064] In some embodiments, the preprocessing signal generation module further includes a sorting module for sorting the summation signal and the difference signal. The echo estimation module further includes a sequential adaptive filtering module for performing adaptive filtering on the sorted summation signal and difference signal in sequence to generate echo estimation signals corresponding to the sorting. The residual signal generation module includes a sequential residual signal generation module for sequentially filtering the corresponding sorted echo estimation signals from M microphone signals to generate a residual signal. By sorting the summation signal and the difference signal and generating the residual signal accordingly, a better echo filtering effect can be obtained for different cases.
[0065] In some embodiments, the preprocessing signal generation module includes a speaker signal sorting module for generating N sorted preprocessing signals by sorting N speaker signals. The echo estimation module includes: a sequential adaptive filtering module for sequentially performing adaptive filtering on the N sorted preprocessing signals to generate N echo estimation signals corresponding to the sorting. The residual signal generation module includes a sequential residual signal generation module for sequentially filtering the N echo estimation signals corresponding to the sorting from M microphone signals to generate a residual signal. By sorting the N speaker signals and filtering them sequentially, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0066] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components of each of the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo filtering effect can be improved.
[0067] In some embodiments, the preprocessing signal generation module is further configured to perform non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessing signal. By using the M microphone signals as auxiliary reference signals, the frequency band with greater echo can be restricted to assist in improving echo filtering.
[0068] In some embodiments, the apparatus further includes a gain adjustment module for adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessing signal so that the gain of the echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0069] According to an eleventh aspect of the present disclosure, there is provided an apparatus for filtering echo. The apparatus is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The apparatus includes: a first acquisition module for acquiring N speaker signals corresponding to the N speakers; a reverse speaker signal generation module for generating a reverse speaker signal based on the N speaker signals; a playback enabling module for enabling a reverse speaker close to at least one of the M microphones to play reverse audio based on the reverse speaker signal to cancel the audio output corresponding to the N speaker signals played by the N speakers. The reverse speaker is different from the N speakers. By enabling the reverse speaker to play reverse audio, a part of the echo component can be filtered before the echo is collected by the microphone, thereby providing an echo filtering effect.
[0070] In some embodiments, the device further includes a second acquisition module configured to acquire M microphone signals corresponding to M microphones; and a direct sound filtering module configured to perform at least direct sound filtering on N speaker signals and M microphone signals to obtain a target signal, where direct sound filtering indicates filtering of the audio components that are directly output from the N speakers to the M microphones without environmental reflection. By using direct sound filtering, the effect of echo filtering can be further improved.
[0071] In some embodiments, the target signal is used to wake up the engine to wake up the intelligent voice assistant or is transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, and the echo components are used to characterize the echoes of the audio outputs of the N speaker signals propagating in space collected by the M microphones. By using the target signal filtered by direct sound, the success rate of waking up the intelligent voice assistant by the wake-up engine and / or the quality of the voice call can be improved.
[0072] In some embodiments, the direct sound filtering module includes a default direct sound filtering module, and the default direct sound filtering indicates filtering based at least on the model relationship between the N speaker signals played by the N speakers and the M microphone signals directly collected by the M microphones in an anechoic environment.
[0073] In some embodiments, the device further includes: a display enabling module configured to enable the display screen of the electronic device to display a customized direct sound filtering interface; an input receiving module configured to receive user input on the customized direct sound filtering interface; a speaker test module configured to, in response to the user input, acquire N speaker test signals and cause the N speakers to play the N speaker test signals; a third acquisition module to acquire M microphone test signals corresponding to the M microphones; and a storage module configured to store a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the device further includes using the customized direct sound filtering model to perform customized direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal. By using customized direct sound filtering, the direct sound filtering can be optimized for the customer's environment and the echo cancellation in the customer's environment can be further improved.
[0074] In some embodiments, the customized direct sound filtering interface displays a prompt item for prompting to keep the environment quiet. The customized direct sound filtering interface may also display a decibel indication item indicating the ambient noise and / or an indication item indicating whether it is suitable for customized echo filtering. By displaying the prompt item for prompting the environment to be quiet, the decibel indication item, and / or the indication item indicating whether it is suitable for customized echo filtering, the user can be prompted to establish a customized direct sound filtering model in a quiet and suitable environment, thereby laying a good foundation for subsequent use of the customized direct sound filtering specific to this environment and subsequently obtaining a good echo filtering effect in this environment.
[0075] In some embodiments, the apparatus further includes: an echo estimation module for generating an echo estimation signal based on N speaker signals; a residual signal generation module for filtering the echo estimation signal from M microphone signals to generate a residual signal; and a target signal generation module for performing direct sound filtering on the residual signal to obtain a target signal. By preprocessing the microphone signals to filter the echo estimation signal before direct sound filtering, the echo filtering effect can be further improved.
[0076] In some embodiments, the echo estimation module includes: a preprocessing signal generation module for performing non-interleaved preprocessing on the N speaker signals to generate at least one preprocessing signal; and an adaptive filtering module for performing adaptive filtering on the at least one preprocessing signal to generate an echo estimation signal. By performing non-interleaved preprocessing on the N speaker signals, a preprocessing signal that continuously represents the echo in time can be obtained, and by using the adaptive filtering signal based on the non-interleaved preprocessing signal to estimate the echo component in the M microphone signals, a better echo filtering effect can be obtained.
[0077] In some embodiments, the preprocessing signal generation module includes: a summation module for linearly summing at least two of the N speaker signals to generate a summation signal. By combining at least two of the N speaker signals into a single summation signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and filtering the echo in the full frequency band to obtain a better echo filtering effect.
[0078] In some embodiments, the preprocessing signal generation module further includes: a difference module for linearly differencing at least two of the N speaker signals to generate a difference signal. In some cases, echo filtering focuses on the difference in a certain frequency band or between the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo filtering effect.
[0079] In some embodiments, the preprocessing signal generation module further includes a sorting module for sorting the sum signal and the difference signal. The echo estimation module further includes a sequential adaptive filtering module for sequentially performing adaptive filtering on the sorted sum signal and difference signal to generate echo estimation signals corresponding to the sorting. The residual signal generation module includes a sequential residual signal generation module for sequentially filtering the corresponding sorted echo estimation signals from the M microphone signals to generate a residual signal. By sorting the sum signal and the difference signal and correspondingly generating a residual signal, better echo filtering effects can be obtained for different situations.
[0080] In some embodiments, the preprocessing signal generation module includes a speaker signal sorting module for generating N sorted preprocessing signals by sorting the N speaker signals. The echo estimation module includes: a sequential adaptive filtering module for sequentially performing adaptive filtering on the N sorted preprocessing signals to generate N echo estimation signals corresponding to the sorting. The residual signal generation module includes a sequential residual signal generation module for sequentially filtering the corresponding sorted N echo estimation signals from the M microphone signals to generate a residual signal. By sorting the N speaker signals and sequentially filtering them, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0081] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components of each speaker signal among the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo filtering effect can be improved.
[0082] In some embodiments, the preprocessing signal generation module is further configured to perform non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessing signal. By using the M microphone signals as auxiliary reference signals, the frequency band with larger echoes can be restricted to assist in improving echo filtering.
[0083] In some embodiments, the apparatus further includes a gain adjustment module for adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessing signal so that the gain of the echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0084] According to a twelfth aspect of the present disclosure, a device for echo cancellation is provided. The device is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The device includes a first acquisition module configured to acquire N speaker signals corresponding to the N speakers; a second acquisition module configured to acquire M microphone signals corresponding to the M microphones; a preprocessing signal generation module configured to perform non-interleaved preprocessing on the N speaker signals to generate M groups of preprocessing signals; an echo estimation module configured to perform adaptive filtering on the M groups of preprocessing signals to generate M echo estimation signals, and a residual signal generation module configured to filter the M echo estimation signals from the M microphone signals to obtain a residual signal. By performing non-interleaved preprocessing on the N speaker signals, preprocessing signals that continuously represent the echo in time can be obtained, and by using adaptive filtering signals based on the non-interleaved preprocessing signals to estimate the echo components in the M microphone signals, a better echo cancellation effect can be obtained.
[0085] In some embodiments, the residual signal is a target signal. The target signal is used to wake up the engine to wake up the intelligent voice assistant or to be transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, and the echo components are used to characterize the echo of the sound propagated in space by the N speaker signals collected by the M microphones. By using the target signal based on the preprocessing signal that has been non-interleaved preprocessed, the success rate of waking up the intelligent voice assistant by the wake-up engine and / or the quality of the voice call can be improved.
[0086] In some embodiments, the preprocessing signal generation module includes: a summation module configured to linearly sum at least two of the N speaker signals to generate a summation signal. By combining at least two of the N speaker signals into a single summation signal, the computational overhead of subsequent adaptive filtering can be reduced, thereby reducing the overall computational overhead and enabling echo cancellation in the full frequency band to obtain a better echo cancellation effect.
[0087] In some embodiments, the preprocessing signal generation module further includes: a difference module configured to linearly subtract at least two of the N speaker signals to generate a difference signal. In some cases, echo cancellation focuses on a certain frequency band or the difference between the outputs of different speakers. In this case, the difference signal can be provided to the processor alone or combined with other signals to further improve the echo cancellation effect.
[0088] In some embodiments, the preprocessing signal generation module further includes a sorting module for sorting the sum signal and the difference signal. The echo estimation module further includes a sequential adaptive filtering module for sequentially performing adaptive filtering on the sorted sum signal and difference signal to generate echo estimation signals corresponding to the sorting. The residual signal generation module includes a sequential residual signal generation module for sequentially filtering the corresponding sorted echo estimation signals from the M microphone signals to generate a residual signal. By sorting the sum signal and the difference signal and filtering them sequentially, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0089] In some embodiments, the preprocessing signal generation module includes a speaker signal sorting module for sorting the N speaker signals to generate N sorted preprocessing signals. The echo estimation module includes a sequential adaptive filtering module for sequentially performing adaptive filtering on the N sorted preprocessing signals to generate echo estimation signals corresponding to the sorting. The target signal generation module includes a sequential filtering module for sequentially filtering the corresponding sorted echo estimation signals from the M microphone signals to generate a target signal. By sorting the N speaker signals and filtering them sequentially, signals that cause greater distortion can be preferentially filtered to provide an echo filtering effect.
[0090] In some embodiments, sorting the N speaker signals includes sorting the N speaker signals based on the low-frequency components respectively possessed by each of the N speaker signals. By preferentially filtering the speaker signals with low-frequency components, the echo filtering effect can be improved.
[0091] In some embodiments, the preprocessing signal generation module also performs non-interleaved preprocessing on the N speaker signals and the M microphone signals to generate at least one preprocessing signal. By using the M microphone signals as auxiliary reference signals, the frequency bands with greater echo can be restricted to assist in improving echo filtering.
[0092] In some embodiments, the apparatus further includes a gain adjustment module for adjusting the gain of at least one of the N speaker signals, the M microphone signals, and the at least one preprocessing signal so that the gain of the echo estimation signal matches the gain of the M microphone signals. By adjusting the gain, the gain of the echo estimation signal can be made to match the gain of the M microphone signals, thereby improving the adaptive filtering effect and the echo filtering effect.
[0093] In some embodiments, the echo estimation module includes a parallel adaptive filtering module for, in the case of multiple preprocessing signals, performing parallel adaptive filtering on the multiple preprocessing signals to generate at least one echo estimation signal. For multiple preprocessing signals with low similarity, parallel preprocessing can improve the echo filtering effect.
[0094] In some embodiments, the echo estimation module includes: a conversion module configured to convert at least one preprocessed signal from a time-domain signal to a frequency-domain signal, and a frequency-domain adaptive filtering module configured to adaptively filter the frequency-domain signal to obtain at least one echo estimation signal. By converting the preprocessed signal from the time domain to the frequency domain, the computational overhead of the adaptive filtering can be reduced.
[0095] In some embodiments, the apparatus includes a direct sound filtering module configured to perform direct sound filtering on the residual signal to generate a target signal. Direct sound filtering indicates filtering of the audio components that are directly output from N speakers to M microphones without environmental reflections. More specifically, direct sound filtering indicates filtering based on the model relationship between the sound source signal played by at least one of the first speaker and the second speaker and the audio input signal directly collected by the microphone. By using direct sound filtering, the effect of echo filtering can be further improved.
[0096] In some embodiments, the direct sound filtering module includes performing default direct sound filtering configured to perform default direct sound filtering on the target signal to generate at least one echo estimation signal. Default direct sound filtering indicates filtering based on the model relationship between the sound source signal played by N speakers in an anechoic environment and the audio input signal directly collected by M microphones.
[0097] In some embodiments, the apparatus further includes: a display enabling module configured to enable a display screen of the electronic device to display a customized direct sound filtering interface; an input receiving module configured to receive user input on the customized direct sound filtering interface; a speaker test module configured to, in response to the user input, obtain N speaker test signals and cause the N speakers to play the N speaker test signals; a fourth obtaining module configured to obtain M microphone test signals corresponding to the M microphones; and a storage module configured to store a customized direct sound filtering model. The customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering. In some embodiments, the target signal generation module of the apparatus is further configured to use the customized direct sound filtering model to perform customized direct sound filtering on the N speaker signals and the M microphone signals to obtain a target signal. By using customized direct sound filtering, the direct sound filtering can be optimized for the customer's environment and the echo cancellation in the customer's environment can be further improved.
[0098] In some embodiments, the apparatus further includes: a reverse audio signal generation module configured to generate a reverse audio signal based on N speaker signals, and a playback enabling module configured to enable a reverse speaker close to the microphone to play the reverse audio based on the reverse audio signal so as to cancel the audio output played by at least one of the N speakers. The reverse speaker is different from the first speaker and the second speaker. By enabling the reverse speaker to play the reverse audio, a part of the echo component can be filtered out before the echo is collected by the microphone, thereby providing an echo filtering effect.
[0099] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] With reference to the accompanying drawings and the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals indicate the same or similar elements, where:
[0101] Figure 1 A schematic diagram showing an environment in which the embodiments of the present disclosure can be implemented;
[0102] Figure 2 A schematic block diagram showing an electronic device according to an embodiment of the present disclosure;
[0103] Figure 3 A schematic flowchart showing a method for filtering echo according to an embodiment of the present disclosure;
[0104] Figure 4 A schematic diagram showing a process of direct sound filtering according to an embodiment;
[0105] Figure 5 A schematic flowchart showing a method for filtering echo according to another embodiment of the present disclosure;
[0106] Figure 6 A schematic diagram showing a process of reverse echo cancellation according to an embodiment;
[0107] Figure 7 A schematic flowchart showing a method for filtering echo according to still another embodiment of the present disclosure;
[0108] Figure 8 A schematic diagram showing a process of non-interleaved preprocessing according to an embodiment;
[0109] Figure 9 is Figure 8 A schematic diagram showing audio signal processing of an embodiment of non-interleaved preprocessing in
[0110] Figure 10 is Figure 8 A schematic diagram of audio signal processing for another embodiment of non - interleaved pre - processing in
[0111] Figure 11 is Figure 8 A schematic diagram of audio signal processing for yet another embodiment of non - interleaved pre - processing in
[0112] Figure 12 A schematic diagram of a serial audio signal processing process for echo cancellation according to yet another embodiment of the present disclosure;
[0113] Figure 13 A schematic diagram of a parallel audio signal processing process for echo cancellation according to yet another embodiment of the present disclosure;
[0114] Figure 14 A schematic diagram of an audio signal processing process for echo cancellation according to yet another embodiment of the present disclosure;
[0115] Figure 15 A schematic block diagram of a device for echo cancellation according to an embodiment of the present disclosure; and
[0116] Figure 16 A schematic block diagram of a device for echo cancellation according to yet another embodiment of the present disclosure. Detailed implementation manners
[0117] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0118] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.
[0119] As described above, AEC technology has been widely applied in communication electronic devices. However, the echo cancellation effect of conventional AEC technology is still not ideal in some cases, so further improvement is needed.
[0120] In view of the above problems and other potential problems, embodiments of the present disclosure provide a method, an electronic device, and a computer-readable storage medium for echo cancellation. In an embodiment of the present disclosure, direct sound filtering is performed on multiple microphone signals and multiple speakers to obtain a target signal. This can achieve a better echo cancellation effect. In another embodiment of the present disclosure, by setting a reverse speaker adjacent to at least one of the multiple microphones and causing the reverse speaker to play a reverse audio opposite to the audio output played by the multiple speakers, the echo component of the audio output played by the multiple speakers collected by the at least one microphone can be cancelled. In still another embodiment of the present disclosure, by performing non-interleaved preprocessing on the microphone signals corresponding to multiple sound source signals, the echo components corresponding to the multiple echoes of the multiple sounds played by the multiple speakers in the microphone signals can be better estimated, and a better echo cancellation effect can be obtained by filtering out the echo components. Since the estimated signal can to a certain extent predict the temporally continuous echoes of each speaker, by using the at least one estimated signal, the echoes of the sounds played by each speaker can be effectively cancelled from the microphone signals collected by the microphone. In the present disclosure, the above three embodiments can be used alone or in any combination to obtain a better echo cancellation effect.
[0121] Figure 1 FIG. shows a schematic diagram of an exemplary environment 1 in which embodiments of the present disclosure can be implemented. In one embodiment, the electronic device 100 can be, for example, a smart speaker that can play audio, such as music or language programs. The electronic device 100 can include N speakers located inside, where N is an integer greater than 1. In one embodiment, the electronic device 100 includes 7 speakers, where the first group of speakers (collectively denoted by reference numeral 14) among the 7 speakers can be 6 mid-high frequency speakers arranged in a ring in the middle of the electronic device 100. Among the first group of speakers, 3 speakers located on the front of the electronic device 100 are shown, while the other three speakers in the first group of speakers are located on the back of the electronic device 100 and are thus not shown. The second group of speakers 14-7 among the 7 speakers can be a bass unit provided at the bottom of the electronic device 100. For certain types of audio, the sounds played by the first group of speakers 14 and the second group of speakers 14-7 can be different. For example, the first group of speakers 14 can mainly play mid-high frequency sounds, while the second group of speakers 14-7 mainly play low frequency sounds.
[0122] The electronic device 100 may further include a microphone array 12, and the microphone array 12 may include M microphones, where M is an integer greater than 1. In one embodiment, the electronic device 100 includes, for example, 6 microphones located at the top (collectively shown by reference numeral 12), among which 3 microphones located on the front of the electronic device 100 are shown, while the other three microphones are located on the back of the electronic device 100 and thus not shown. Although Figure 1 illustrates 7 speakers and 6 microphones, this is only illustrative and does not limit the scope of the present disclosure. In some other embodiments, the electronic device 100 may include other numbers of microphones and speakers. In addition, in some embodiments that only use direct sound filtering and / or reverse speakers without using non-interleaved preprocessing, the electronic device 100 may include only one speaker and one microphone.
[0123] Although Figure 1 illustrates the cylindrical configuration of the electronic device 100, the electronic device 100 may also have other configurations. For example, in one embodiment, the electronic device 100 may be a soundbar, where 4 microphones form a linear array at the top of the electronic device 100, and multiple speakers are also arranged in a transverse linear pattern. For example, in another embodiment, the electronic device 100 may be a smart TV, where 6 microphones form a linear array at the top of the smart TV, and multiple speakers are distributed around the bottom edge, left and right side edges of the screen of the smart TV, and the back of the smart TV.
[0124] In Figure 1 , the sound played by the first group of speakers 14 and the second group of speakers 14-7 will generate reflected echoes after propagating to the wall 2. Especially when the electronic device 100 is arranged close to the wall 2, the reflected echoes will be collected by the microphone array 12 of the electronic device 100. On the other hand, the sound played by the first group of speakers 14 and the second group of speakers 14-7 can also be transmitted to the microphone array 12 through the physical continuous surface of the electronic device 100 itself. Therefore, the echo signal collected by the microphone array 12 includes both a reflected echo component that has been reflected by the environment and a direct sound component that indicates the direct path from the first group of speakers 14 and the second group of speakers 14-7 to the microphone array 12. In this document, the term "echo" includes the reflected echo collected by the microphone of the electronic device after the audio played by the speaker of the electronic device is reflected by the environment and the direct sound directly collected by the microphone from the audio played by the speaker of the electronic device. The term "reflected echo" indicates the audio component that the audio signal is reflected back to the microphone by the environment such as a wall after being output from the speaker and received by the microphone. In contrast, the term "direct sound" indicates the audio component that the audio signal is directly output from the speaker to the microphone without being reflected by the environment such as a wall and directly received by the microphone.
[0125] In a situation where the electronic device 100 plays audio and the user 20 speaks, the microphone array 12 can also collect the user's voice. In this situation, each of the M microphone signals generated by the microphone array 12 collecting sound includes the user's voice, the echo of the audio output played by the speaker, and the possible existing noise.
[0126] In some embodiments, the electronic device 100 may have a voice intelligent assistant and a call function. For example, the user 20 can wake up the voice intelligent assistant of the electronic device 100 through the wake-up engine by speaking a wake-up command such as "Hello, intelligent assistant". After waking up the voice intelligent assistant, the user 20 can also make a voice call by speaking a voice such as "Call mom". In a situation where the first set of speakers 14 and the second set of speakers 14-7 of the electronic device 100 play sound, in order for the electronic device 100 to correctly recognize the wake-up command and for the call recipient to clearly hear the voice without being affected by the echo, the electronic device 100 can implement the echo cancellation technology according to the embodiments of the present disclosure to correctly recognize the voice for wake-up and improve the call clarity. Therefore, by using the embodiments according to the present disclosure, the filtering effect of the echo can be further improved. In the present disclosure, the terms "cancellation" and "filtering" can be used interchangeably and both represent removing a part, while the terms "complete cancellation" and "complete filtering" represent removing all parts.
[0127] Although Figure 1 the application environment of the embodiments of the present disclosure is described using a smart speaker, it can be understood that this is only illustrative and does not limit the scope of the present disclosure. The embodiments of the present disclosure can also be implemented in other electronic devices having speakers and microphones. For example, the electronic devices that can implement the embodiments of the present disclosure may include at least one of the following: smart speakers, set-top boxes, entertainment units, navigation devices, communication devices, fixed-position data units, mobile-position data units, mobile phones, cellular phones, tablets, phablets, computers, portable computers, desktop computers, personal digital assistants (PDAs), monitors, computer monitors, televisions, tuners, radios, satellite radios, music players, digital music players, portable music players, digital video players, video players, digital video disc (DVD) players, portable digital video players, etc.
[0128] Figure 2 FIG. shows a schematic block diagram of an electronic device 100 according to an embodiment of the present disclosure. It should be understood that Figure 2The illustrated electronic device 100 is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described in this disclosure. In one embodiment, the electronic device 100 may include a processor 110, a wireless communication module 160, an antenna 1, an audio module 170, a speaker module 170A, a microphone module 170C, a key 190, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, and a power management module 141, which are shown by solid lines and solid boxes. The microphone module 170C may include, for example, M microphones as described above. The speaker module 170A may include N speakers as described above, such as a first group of speakers 14 and a second group of speakers 14-7, etc. In some other embodiments, the speaker module 170A may further include a reverse speaker group having at least one reverse speaker, and the reverse speaker is used to play reverse audio for canceling out echo components to be picked up by the microphone.
[0129] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. In some embodiments, different processing units may be independent devices. In some other embodiments, different processing units may also be integrated in one or more processors. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0130] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may store instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0131] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0132] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 may be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the electronic device 100.
[0133] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 may be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0134] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 may be coupled through the PCM bus interface. In some embodiments, the audio module 170 may also transmit an audio signal to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0135] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through the Bluetooth headset.
[0136] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.
[0137] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0138] The USB interface 130 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.
[0139] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as the echo filtering function, sound playback function, image playback function, etc. in the embodiments of the present disclosure), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0140] The charging management module 140 is used to receive a charging input from a charger. The charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0141] The power management module 141 is used to connect the charging management module 140, the processor 110, and optionally the battery 142. The power management module 141 receives the input from the battery 142 and / or the charging management module 140, and supplies power to components such as the processor 110, the internal memory 121, and the wireless communication module 160. In embodiments with a battery 142, the power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be provided in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.
[0142] The electronic device 100 can implement audio functions through the audio module 170, the speaker module 170A, the microphone module 170C, and the application processor, etc. For example, music playback, recording, etc. The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0143] The speaker module 170A, also known as the "loudspeaker", for example, includes N speakers and an optional reverse speaker (for example, M reverse speakers) and is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker module 170A.
[0144] The microphone module 170C, also known as the "microphone", "transmitter", for example, includes M microphones and is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone module 170C to input the sound signal into the microphone module 170C. In the disclosed embodiments, the electronic device 100 uses an echo cancellation function to reduce the echo component in the sound picked up by the microphone module 170C that is related to the sound played by the speaker module 170A, thereby improving the accuracy of smart assistant wake-up and / or improving the voice call quality. In other embodiments, the electronic device 100 can also identify the sound source and implement a directional recording function, etc.
[0145] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. The antenna 1 is used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.
[0146] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 1, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive the signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 1 for radiation.
[0147] In some other embodiments, the electronic device 100 may further include an antenna 2 and a mobile communication module 150. The antenna 2 is used for transmitting and receiving electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 2 can be multiplexed as a diversity antenna for the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0148] The mobile communication module 150 may provide solutions for wireless communications applied to the electronic device 100, including 2G / 3G / 4G / 5G, etc. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves through the antenna 2, perform filtering, amplification, etc. on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modulation and demodulation processor and convert them into electromagnetic waves through the antenna 2 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.
[0149] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker module 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be disposed in the same device as the mobile communication module 150 or other functional modules.
[0150] In some embodiments, the antenna 2 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 1 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0151] The button 190 includes a power-on button, volume buttons, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100.
[0152] In some other embodiments, in addition to the above components, the electronic device 100 may further include one or more of an external memory interface 120, a battery 142, a receiver 170B, a headphone interface 170D, a sensor module 180, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, which are shown by dotted lines and dotted boxes. The sensor module 180 may include one or more of a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M. The sensor module 180 may also include other types of sensors not listed.
[0153] The electronic device 100 realizes the display function through a GPU, the display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0154] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0155] In some embodiments, the electronic device 100 further includes a receiver 170B and a headphone jack 170D. The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a call or a voice message, the voice can be received by placing the receiver 170B close to the human ear. The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0156] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc. The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.
[0157] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0158] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform a Fourier transform on the frequency point energy, etc.
[0159] Video codecs are used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0160] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0161] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are stored in the external memory card.
[0162] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: When a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0163] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and somatosensory game scenes.
[0164] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude through the air pressure value measured by the air pressure sensor 180C to assist positioning and navigation.
[0165] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover according to the magnetic sensor 180D. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip cover, the flip cover can be automatically unlocked.
[0166] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in all directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0167] The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure the distance to achieve fast focusing.
[0168] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the light emitting diode. The electronic device 100 uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user holds the electronic device 100 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0169] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance during photography. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.
[0170] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering of incoming calls, etc.
[0171] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to prevent abnormal shutdown of the electronic device 100 caused by low temperature. In some other embodiments, when the temperature is lower than yet another threshold, the electronic device 100 boosts the output voltage of the battery 142 to prevent abnormal shutdown caused by low temperature.
[0172] The touch sensor 180K, also known as a "touch control device". The touch sensor 180K can be disposed on the display screen 194, and together with the display screen 194, it forms a touch screen, also known as a "touch control screen". The touch sensor 180K is used to detect touch operations acting on it or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from the display screen 194.
[0173] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in earphones to form a bone conduction earphone. The audio module 170 can analyze the voice signal based on the vibration signals of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can analyze the heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.
[0174] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations applied to different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations applied to different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as: time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0175] The indicator 192 can be an indicator light and can be used to indicate the charging state, battery power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0176] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the electronic device 100. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, that is: an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0177] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.
[0178] In addition, it can be understood that the interface connection relationships between the various modules schematically shown in the embodiments of the present invention are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods as described in the above embodiments, or a combination of multiple interface connection methods.
[0179] Figure 3FIG. 300 is a schematic flow chart of a method for echo cancellation according to an embodiment of the present disclosure. In some cases, multiple microphone signals may be generated by audio effect algorithms such as virtual sound field, upmixing, sound field expansion, etc. Some audio effect algorithms can cause strong correlations between microphone signals, posing a great challenge to adaptive filtering. In one embodiment, direct sound processing can model the acoustic transfer function coefficients from each speaker to each microphone in advance. These pre-estimated models are used for filtering during the filtering process, and as a subsequent process of adaptive filtering, it can improve the filtering effect of adaptive filtering.
[0180] In one embodiment, method 300 may be executed by processor 110 of electronic device 100. Electronic device 100 includes M microphones and N speakers, where both M and N are integers greater than 1. Although method 300 is shown as being executed by processor 110, this is for illustration only and does not limit the scope of the present disclosure. One or more operations in method 300 may be executed by other computing devices such as a digital signal processor (DSP) other than processor 110.
[0181] At 302, processor 110 obtains N speaker signals corresponding to the N speakers. In one embodiment, the N speaker signals are copies of the N audio signals played by the N speakers. In the present disclosure, the term "copy" refers to a replicated signal that has a one-to-one correspondence with the source signal in terms of audio content. In one embodiment, the N speaker signals may be exactly the same as the N audio signals played by the speakers. In another embodiment, the N speaker signals may not be exactly the same as the N audio signals played by the speakers, but can reflect the audio content of the N audio signals. For example, the content of the N speaker signals is substantially the same as the N audio signals, but the gains of the N speaker signals are different from the gains of the N audio signals.
[0182] At 304, processor 110 obtains M microphone signals corresponding to the M microphones. As described above, each of the M microphone signals includes user speech, echoes of the audio output played by the N speakers, and possibly existing noise. In one embodiment, processor 110 may process the M microphone signals in sequence. Alternatively, in the case where processor 110 includes multiple processing cores, the multiple processing cores may process the M microphone signals respectively to improve the processing speed. It can be understood that in the case where the speakers play audio, each of the M microphone signals includes a target audio signal component, N echo signal components of the N speakers, and a noise signal component.
[0183] At 306, the processor 110 performs at least direct sound filtering on N speaker signals and M microphone signals to obtain a target signal. In some embodiments, the target signal is used to wake up the engine to wake up the intelligent voice assistant or is transmitted to another electronic device for a voice call. In some embodiments, the target signal contains fewer echo components than the M microphone signals, and the echo components are used to characterize the echoes of the sounds propagated in space of the N speaker signals collected by the M microphones.
[0184] There is a physical direct sound propagation path from each of the N speakers to each of the M microphones without passing through environmental reflections. The filter coefficients of the direct sound filtering model are fixed and do not need to be updated during filtering because the direct sound propagation path from the speaker to the microphone is usually fixed. Therefore, the direct sound filtering model can be established accordingly. In some embodiments, using the direct sound filtering model for the N×M direct propagation path function from N speakers to M microphones, the direct sound components of the N speakers can be filtered out from the M microphone signals, thereby obtaining a better echo filtering effect, and thus improving the success rate of waking up the engine and / or improving the quality of the voice call.
[0185] In some embodiments, the direct sound filtering includes default direct sound filtering, and the default direct sound filtering indicates filtering based at least on the model relationship between the N speaker signals played by the N speakers in an anechoic environment and the M microphone signals directly collected by the M microphones. In one embodiment, in an anechoic chamber without any reflected echoes, the processor 110 of the electronic device 100 controls the N speakers of the product to play white noise or a music sound source separately, and uses all M microphones to record separately, where M and N represent positive integers greater than 1, and can be the same or different. The electronic device performs filtering on multiple recordings in sequence using adaptive filtering, and stores M×N groups of filter coefficients reflecting the transfer functions from the N speakers to the M microphones as the default direct sound filtering model. The default direct sound filtering model is stored in the electronic device before leaving the factory or is stored in the electronic device by means of firmware update of the electronic device for the user to use when filtering out the direct sound.
[0186] In another embodiment, the direct sound filtering may further include customized direct sound filtering. The user can customize the direct sound cancellation model for the daily-use environment to obtain a better direct sound filtering effect. For example, the electronic device 100 may display a customized direct sound filtering interface on the display 194. The user 20 can click and input customized virtual buttons on the customized direct sound filtering interface, for example, in a touch manner. After receiving the user input, the processor 110 obtains N test signals for N speakers from the internal memory 121, and causes the N speakers to play the N speaker test signals and uses all M microphones to record. The processor 110 also obtains M microphone test signals corresponding to the M microphones, and generates a customized direct sound filtering model for customized direct sound filtering based on the N speaker test signals and the M microphone test signals, and stores it in the internal memory 121. The customized direct sound filtering model is used for the above-mentioned direct sound filtering.
[0187] When the user 20 plays audio subsequently, the user 20 can select the default direct sound filtering or the customized direct sound filtering through the direct filtering selection interface displayed on the display 194. In the case where the user 20 selects the customized direct sound filtering, the electronic device 100 uses the stored customized direct sound filtering model to filter the M microphone signals. It can be understood that in some embodiments, the direct filtering selection interface and the customized direct sound filtering interface may be different options displayed in different display areas of the same interface.
[0188] In some embodiments, the customized direct sound filtering interface displays a prompt item for prompting to keep the environment quiet. The customized direct sound filtering interface may also display a decibel indication item indicating the ambient noise and / or an indication item indicating whether it is suitable for customized echo filtering. For example, in the case where the electronic device 100 includes a sensor for measuring ambient noise, the current ambient noise can be displayed in real time on the customized direct filtering interface, and the above-mentioned direct sound test is only performed when the current ambient noise is lower than a threshold. By displaying the prompt item for prompting the environment to be quiet, the decibel indication item and / or the indication item indicating whether it is suitable for customized echo filtering, the user can be prompted to establish a customized direct sound filtering model in a quiet and suitable environment, thereby laying a good foundation for subsequent use of the customized direct sound filtering specific to this environment, and subsequently obtaining a good echo filtering effect in this environment.
[0189] In some embodiments where the electronic device 100 does not have a display screen, the electronic device 100 can play voice to guide the user to keep quiet in the usage environment and start customization. The processor 110 of the electronic device 100 will cause the N speakers to play white noise or a music sound source separately, and use all M microphones to record.
[0190] In addition, the electronic device 100 may have a plurality of customized direct sound filter options to customize the direct sound filter for different environments. When the electronic device 100 is in different environments, the user 20 may select a direct sound filter model for the environment in the direct sound filter selection interface. In addition, when the electronic device 100 arrives in a new environment, the direct sound filter model for the new environment may be re-customized, and the new direct sound filter model for the new environment may be stored in the internal memory 121 for subsequent use.
[0191] Although in Figure 3 In the description, the electronic device 100 is described as having M microphones and N speakers, but this is only for illustration and does not limit the scope of the present disclosure. The method 300 may also be applied to electronic devices having other numbers. For example, the method 300 may also be applied to electronic devices having multiple speakers and a single microphone, multiple microphones and a single speaker, or a single microphone and a single speaker.
[0192] Figure 4 FIG. 4 is a schematic diagram of a direct sound filtering process 400 according to an embodiment. N are respectively output to N loudspeakers 14-1...14-N, and serve as N sound source signals X1...X N The N copies of the loudspeaker signal X 1C …X NC is provided to the processor 110 for adaptive filtering 440, and the audio source signals X1 . . . X N Another set of copies X 1D …X ND It is used for direct sound filtering 450. The sound source is, for example, audio data stored in a storage device in a server communicating via the Internet, audio data stored in a local storage device, or audio data collected by a microphone from another device.
[0193] Although the N loudspeaker signal X 1C …X NC It is shown as being directly adaptively filtered, but this is only for illustration and does not limit the scope of the present disclosure. It can be understood that the N loudspeaker signals X 1C …X NC Various adjustments and processing may be performed, such as gain scaling and non-interleaved preprocessing described below, and the adjusted and preprocessed signal may be provided for adaptive filtering. In other embodiments, the signal provided to the speaker and the microphone signal used for adaptive filtering may be different (e.g., different gains), but have a correlation so that the microphone signal can indicate the sound played by the speaker.
[0194] The processor 110 processes the N speaker signals X 1C…X NC Perform adaptive filtering 440 to generate M echo estimation signals Y E . In one embodiment, the adaptive filtering 440 includes, for example, least mean square (LMS) filtering, recursive least squares (RLS) filtering, etc. In one embodiment, the adaptive filtering is performed by converting the preprocessed signal from the time domain signal to a frequency signal and then performing adaptive filtering on the frequency signal. For example, for a time domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can be used to convert it to a frequency domain signal. Of course, other overlap lengths and Fourier transforms with other numbers of points can also be used. In the case of using a higher proportion of the overlap length, the continuity of the front and rear audio frames can be improved, but the computational overhead is increased. In the case of using a higher number of points for the Fourier transform, the spectral resolution can be improved to improve the adaptive result, but this also increases the computational overhead. For a time domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can achieve a better balance between the filtering effect and the computational overhead.
[0195] The M microphones 12-1…12-M respectively collect various sounds, which respectively include the speech of the user 20, the audio echo of the audio played by the N speakers 14-1…14-N, and possible noise. In one embodiment, each microphone signal D collected by the M microphones 12-1…12-M includes a speech signal component S, a noise signal component V, and an echo signal component Y, where the echo signal component Y includes a reflected echo signal component and a direct sound signal component. Therefore, the microphone signal D is a composite signal of the speech signal component S, the noise signal component V, and the echo signal component Y. As Figure 4 shown, the first microphone signal can be represented by the formula D1 = S1 + V1 + Y1… The Mth microphone signal can be represented by the formula D M = S M + V M + Y M .
[0196] The M echo estimation signals Y E respectively correspond to the microphone signals D1……D collected by the M microphones M . For example, the first echo estimation signal among the M echo estimation signals Y E corresponds to the first microphone signal D1… The Mth echo estimation signal among the M echo estimation signals Y E corresponds to the first microphone signal D M .
[0197] The processor 110 filters out M echo estimation signals from the M microphone signals respectively to generate M residual signals E. For example, the first echo estimation signal is filtered out from the first microphone signal D1 to obtain the first residual signal... The Mth echo estimation signal is filtered out from the Mth microphone signal D M to obtain the Mth residual signal. The M residual signals E can also be used to update the adaptive filter 440 to improve the accuracy of echo estimation. For example, the first residual signal is used to update the adaptive filter corresponding to the first echo estimation signal... The Mth residual signal is used to update the adaptive filter corresponding to the Mth echo estimation signal.
[0198] The processor 110 can then perform direct sound filtering 450 on the M residual signals E respectively using the default direct sound filtering model or the customized direct sound filtering model described above to generate M target signals T. For example, direct sound filtering 450 is performed on the first residual signal to generate the first target signal... Direct sound filtering 450 is performed on the Mth residual signal to generate the Mth target signal. In one embodiment, the M target signals T are subsequently used to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient via the network. In another embodiment, the M target signals T can then be further filtered through conventional schemes such as non-linear filtering schemes, machine learning schemes, etc. to obtain better echo filtering effects, and then used to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient via the network.
[0199] Figure 5 is a schematic flowchart of a method 500 for filtering echoes according to another embodiment of the present disclosure. The electronic device 100 includes N speakers 14-1...14-N, M microphones 12-1...12-M, and a processor 110. In addition, the electronic device 100 may have at least one reverse speaker disposed near at least one of the microphones in the microphone array 12. For example, for the M microphones, the electronic device 100 may have M reverse speakers corresponding to the M microphones respectively. In another embodiment, when the microphones are close to each other, reverse speakers for adjacent two microphones can be disposed between the microphones to save space and cost. In other words, the electronic device 100 may have fewer reverse speakers than the M microphones.
[0200] In one embodiment, the reverse speaker can be set at a distance of 10 cm, 5 cm, 1 cm or closer to the microphone array 12 to play reverse audio, so as to cancel the echo of the audio played by N speakers. Herein, the term "reverse audio" refers to the audio used to cancel the echo of the audio played by N speakers. The term "reverse speaker" refers to a speaker that is set near the microphone and used to play reverse audio. Compared with methods such as the previously described method 300 and the subsequently described method 700 that filter out echoes through internal echo algorithm estimation, method 500 cancels echoes by externally playing reverse audio. This can reduce the echo component in the microphone signal to obtain an echo filtering effect.
[0201] In one embodiment, method 500 can be executed by the processor 110 of the electronic device 100. The processor 110 of the electronic device 100 can receive the audio signal to be played by the plurality of speakers through the communication unit and transmit it to the plurality of speakers for playback. Alternatively, the processor 110 can receive the audio to be played by the plurality of speakers from the storage unit, ROM or RAM and transmit it to the plurality of speakers for playback. Although method 500 is shown to be executed by the processor 110, this is only illustrative and does not limit the scope of the present disclosure. One or more operations in method 500 can be executed by other computing devices such as a digital signal processor (DSP) other than the processor 110.
[0202] At 502, the processor 110 obtains N speaker signals corresponding to N speakers. In one embodiment, the N speaker signals are copies of the N audio signals played by the N speakers.
[0203] At 504, the processor 110 generates a reverse speaker signal based on the N speaker signals. In one embodiment, in an anechoic environment, the direct sound impulse response model from each speaker to the microphone can be tested, and then a reverse filtering model can be designed based on the direct sound impulse response model. Similar to the default direct sound filtering model, the reverse filtering model can be stored in the electronic device before the electronic device leaves the factory or through a firmware update of the electronic device for subsequent reverse echo filtering. Similarly, the user can also customize the reverse filtering model in their own environment. By controlling the volume and the sound frequency band played by the reverse speaker, the electronic device can make the reverse audio act only at the microphone and in the target frequency band, ultimately enabling the microphone to receive less echo, thereby improving the echo filtering performance.
[0204] At 506, the processor 110 causes at least one reverse speaker near at least one of the M microphones to play reverse audio based on a reverse audio signal to cancel the echo of the audio output of the N speakers. By causing the reverse speaker to play reverse audio, a part of the echo component can be filtered out before the echo is collected by the microphone, thereby providing an echo filtering effect. It can be understood that the method 500 can be used alone to filter out echoes, or can be combined with at least one of the methods 300 and 700 to obtain a better echo filtering effect.
[0205] Although the electronic device 100 is described as having M microphones and N speakers in Figure 5 , this is only illustrative and does not limit the scope of the present disclosure. The method 500 can also be applied to electronic devices with other numbers of microphones and speakers. For example, the method 500 can also be applied to electronic devices having multiple speakers and a single microphone, multiple microphones and a single speaker, and a single microphone and a single speaker.
[0206] Figure 6 is a schematic diagram of a process 600 for reverse echo cancellation according to an embodiment. M anti-phase speakers 15-1... 15-M are respectively arranged near the M microphones 12-1... 12-M to play the echo components in the sound to be collected by the microphones. As N sound source signals X1... X N output from the sound source are output by the processor 110 to the N speakers 14-1... 14-N. The sound source is, for example, audio data stored in a storage device in a server communicating through the Internet, audio data stored in a local storage device, or audio data collected by a microphone of another device.
[0207] As copies of the N sound source signals X1... X N —N speaker signals X 1C ... X NC are reversely processed 620 by the processor 110 to generate M reverse audio signals X 1R ... X MR , and the processor 110 causes the M reverse speakers 15-1... 15-M to play the M reverse audio signals X 1R ... X MR . The first reverse audio signal X 1R cancels the echo of the audio played by the N speakers 14-1... 14-N outside the first microphone 12-1, so that the microphone signal D1 collected by the first microphone 12-1 includes a voice signal S1, noise V1, and a remaining echo signal component Y 1R that is not completely filtered out. And so on, the Mth reverse audio signal X MREcho cancellation is performed outside the Mth microphone 12-M for the echo of the audio played by the N speakers 14-1...14-N, so that the microphone signal D collected by the Mth microphone 12-M M includes the voice signal S M , noise V M and the remaining echo signal component Y that has not been completely filtered out MR . In one embodiment, the microphone signal can be further filtered for echo via at least one of Method 300 and Method 700. In another embodiment, the microphone signal can be used as a target signal to wake up the engine to wake up the intelligent assistant or be transmitted to the call recipient via the network, or be further filtered via conventional schemes such as non-linear filtering schemes, machine learning schemes, etc. to obtain a better echo filtering effect and then be used to wake up the engine to wake up the intelligent assistant or be transmitted to the call recipient via the network.
[0208] Figure 7 is a schematic flowchart of Method 700 for filtering echo according to another embodiment of the present disclosure. The electronic device 100 includes M microphones and N speakers, where both M and N are integers greater than 1. Although it is shown that Method 700 is executed by the processor 110, this is only illustrative and does not limit the scope of the present disclosure. One or more operations in Method 700 can be executed by other computing devices such as a digital signal processor (DSP) other than the processor 110.
[0209] At 702, the processor 110 obtains N speaker signals corresponding to the N speakers. In one embodiment, the N speaker signals are copies of the N audio signals played by the N speakers.
[0210] At 704, the processor 110 obtains M microphone signals corresponding to the M microphones. As described above, each of the M microphone signals includes user voice, the echo of the audio output played by the N speakers, and possibly existing noise. In one embodiment, the processor 110 can process the M microphone signals sequentially. Alternatively, in the case where the processor 110 includes multiple processing cores, the multiple processing cores can process the M microphone signals respectively to improve the processing speed. It can be understood that in the case where the speaker plays audio, each of the M microphone signals includes a target audio signal component, N echo signal components of the N speakers, and a noise signal component.
[0211] At 706, the processor 110 performs non-interleaved preprocessing on N speaker signals to generate M groups of preprocessed signals. Each group of preprocessed signals includes at least one non-interleaved preprocessed signal. In this document, "non-interleaved preprocessing" indicates that multiple audio frames corresponding to the same time slot of multiple speaker signals are preprocessed simultaneously in a combined manner or multiple speaker signals are preprocessed sequentially and the adaptive filter is updated using a target signal corresponding to a single speaker signal, such that the preprocessed signals corresponding to the time slot can reflect at least a portion of the audio characteristics of each speaker signal among the multiple speaker signals. In other words, "non-interleaved preprocessing" includes processing multiple speaker signals in a manner other than interleaving multiple speaker signals by a unit time period or by frame as described above, to continuously indicate the status of multiple speaker signals at the same time in time, without interleaving or splicing multiple audio frames from multiple speaker signals alternately. In any time slot, the preprocessed signals after non-interleaved preprocessing are related to each speaker signal among the multiple speaker signals at any time in the preprocessed audio stream. Non-interleaved preprocessing includes, for example, linear summation, linear subtraction, reference audio signal sorting, serial-parallel filter mode adjustment, gain adjustment, filtering based on microphone signals, etc., as described below.
[0212] In one embodiment, the processor 110 may perform linear summation on at least two of the N speaker signals. For example, the processor performs linear summation on the N speaker signals over time to combine the N speaker signals into a single audio signal. In this way, subsequent filtering by the processor only needs to be performed on the single combined audio signal, thus reducing the subsequent computational operation overhead. In contrast, the conventional interleaved method sequentially selects audio segments of a unit time length from different reference audio signals according to a predetermined unit time length and interleaves them into a single reference audio signal. For example, a single reference audio signal includes a first audio segment of a first unit time period selected from a first reference audio signal, a second audio segment of a first unit time period selected from a second reference audio signal, a third audio segment of a second unit time period selected from the first reference audio signal, and so on. Therefore, the computational overhead of the conventional interleaved method can be N times that of the linear summation operation method. By combining the N speaker signals into a single audio signal, the computational amount of audio signal processing can be significantly reduced.
[0213] In another embodiment, the processor 110 may linearly subtract at least two of the two speaker signals from among the N speaker signals. For example, the processor 110 linearly subtracts the first speaker signal and the second speaker signal over time, thereby synthesizing the two audio signals into a single audio signal. In this way, subsequent filtering by the processor 110 only needs to be performed on the single combined audio signal, rather than on these two speaker signals, thus reducing the subsequent computational operation overhead. In addition, in some cases, echo cancellation focuses on a certain frequency band or the difference between the outputs of different speakers. In this case, the difference signal can be separately provided to the processor 110 or combined with other speaker signals to further improve the echo cancellation effect.
[0214] In yet another embodiment, the processor 110 may sort the N speaker signals to adjust the order of echo cancellation. For example, the processor may first cancel the echo component for the second speaker signal from the microphone signal, and then cancel the echo component for the first speaker signal from the filtered audio signal. In some cases, it is beneficial to sort different speaker signals and perform echo cancellation according to the sorting result. For example, in the case where the second speaker signal is a low-frequency audio indicating the output of a subwoofer speaker, first canceling the low-frequency audio component significantly improves the echo cancellation effect.
[0215] In another embodiment, when at least one of the preprocessing signals includes multiple signals among the sum signal, the difference signal, and the speaker signals, these multiple signals may be sorted according to the echo cancellation effect. For example, a signal that may cause relatively large distortion may be used as a signal with a higher priority for subsequent adaptive filtering. In one embodiment, a signal that may cause relatively large distortion includes a signal with a relatively high low-frequency content. Research shows that by preferentially filtering the low-frequency signal, the filtering effect can be significantly improved. Therefore, by changing the filtering order, the degree of distortion during speech recognition and calls can be reduced.
[0216] In addition to the above-mentioned multiple pre-processed signals being adaptively processed in sequence in series, multiple pre-processed signals can also be combined into matrix signals to perform adaptive processing in parallel. For example, L pre-processed signals can be combined into an L-dimensional matrix to perform adaptive filtering in parallel, where L is an integer greater than 1. It can be selected based on the similarity between multiple speaker signals or pre-processed signals to perform serial adaptive filtering or parallel adaptive filtering on multiple pre-processed signals. In one embodiment, the similarity between multiple speaker signals is high, such as the electronic device 100 plays mono audio. That is, N speakers play the same audio. In this case, serial adaptive filtering can be performed based on N speaker signals associated with N speakers to obtain better filtering effects and better protect the voice input of microphone group 12. In another embodiment, the similarity between multiple speaker signals is low, such as the electronic device 100 plays stereo or 5.1 surround sound audio. In this case, N speaker signals indicating each channel can be combined into a matrix to perform adaptive filtering in parallel to obtain better filtering effects.
[0217] In another embodiment, the processor 110 can adjust the gains of the N speaker signals so that the gain of the echo estimation signal matches the gain of the microphone signal to better filter out the echo. The microphone signal actually received by the microphone is affected by factors including the analog-to-digital conversion gain of the microphone 12. Therefore, if the gain of the microphone signal collected by the microphone 12 does not match the gain of the echo estimation signal, for example, the gain of the echo estimation signal is much lower than the gain of the microphone signal, it may cause the filter to fail to converge to a better state, resulting in the target signal still containing a higher echo signal. Accordingly, the smart assistant of the electronic device may not be activated, and / or the other party may still hear more echoes.
[0218] In one embodiment, the gain may be adjusted based on the acoustic characteristics of the electronic device 100. For example, during the design and manufacture of the electronic device 100, the gain of the echo picked up or received by the test microphone is adjusted based on the final echo filtering effect in the case of only echo, so as to obtain a default gain adjustment setting. Therefore, by adjusting the gain of the N speaker signals or preprocessed signals, the gain of the echo estimation signal can be matched or equivalent to the gain of the microphone signal, thereby more effectively filtering the echo from the microphone signal.
[0219] Alternatively, the gain of the combined signal of the N speaker signals can also be adjusted so that the gain of the echo estimation signal matches the gain of the microphone signal. It can be understood that the gain adjustment can be performed at any operation before the last operation of echo cancellation to achieve gain matching, thereby obtaining a better echo cancellation effect. In addition, it can also be understood that the amplitude of the gain adjustment performed at each stage can be related to the specific operation and does not have to be adjusted by the same amplitude.
[0220] In yet another embodiment, the M microphone signals collected by the microphone array 12 can be used as reference signals for non-interleaved preprocessing. For example, each of the M microphone signals can be subjected to operations such as summation, subtraction, sorting, and gain adjustment with the N speaker signals. In addition, the microphone signals can also be band-pass filtered (such as low-pass filtered) to filter out the frequency bands with larger residual echoes and improve the echo cancellation effect.
[0221] It can be understood that in some embodiments of the present disclosure, different non-interleaved preprocessing can be performed for different speaker signals. For example, in the case where N is 7, summation can be performed for the first and second speaker signals, subtraction can be performed for the third and fourth speaker signals, and gain adjustment can be performed for the fifth and sixth speaker signals, and then the summation signal, subtraction signal, gain-adjusted fifth and sixth speaker signals, and the seventh speaker signal are sorted to generate 5 sorted preprocessing signals.
[0222] At 708, the processor 110 performs adaptive filtering on the M groups of preprocessing signals to generate M echo estimation signals. The M echo estimation signals represent the echo signal components in the M microphone signals estimated based on the speaker signals. Adaptive filtering includes, for example, LMS filtering, RLS filtering, etc. In one embodiment, the adaptive filtering is performed by converting the preprocessing signal from a time-domain signal to a frequency signal and then performing adaptive filtering on the frequency signal. For example, for a time-domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can be used to convert it into a frequency-domain signal. Of course, other overlap lengths and other numbers of points of the Fourier transform can also be used. In the case of using a higher proportion of the overlap length, the continuity of the front and rear audio frames can be improved, but the computational overhead is increased. In the case of using a higher number of points of the Fourier transform, the spectral resolution can be improved to improve the adaptive result, but this also increases the computational overhead. For a time-domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can achieve a better balance between the filtering effect and the computational overhead.
[0223] At 710, the processor 110 filters out M echo estimation signals from the M microphone signals respectively to generate M target signals. In this embodiment, the target signal is a residual signal, which mainly includes a speech signal component, and may also include an echo component that has not been completely filtered out and a noise component. In one embodiment, the target signal is subsequently used by the wake-up engine to wake up the intelligent assistant or transmitted to the call recipient via the network. In another embodiment, the target signal can subsequently be further filtered by conventional schemes such as a non-linear filtering scheme, a machine learning scheme, etc. to obtain a better echo filtering effect, and then used by the wake-up engine to wake up the intelligent assistant or transmitted to the call recipient via the network. In addition, the target signal can also be used to update the adaptive filtering so that the echo estimation signal is closer to the echo component in the microphone signal.
[0224] Although the operations of method 700 are shown in the flowchart of Figure 7 , this is only exemplary and does not limit the scope of the present disclosure. Method 700 may have other additional or optional operations. For example, after 710, the residual signal can be subjected to direct sound filtering as specifically described above to generate the target signal. In other words, method 700 can be used in combination with method 300 to obtain a better echo filtering effect. In addition, method 700 can also be used in combination with method 500, or in combination with method 300 and method 500, as described below with reference to Figure 14 .
[0225] Figure 8 is a schematic diagram of the process 800 of non-interleaved preprocessing according to an embodiment. N source signals X1…X N from the sound source are respectively output by the processor 110 to N speakers. The sound source is, for example, audio data stored in a storage device in a server communicating via the Internet, audio data stored in a local storage device, or audio data collected by a microphone of another device. In addition, N speaker signals X N which are copies of the N source signals X1…X 1C …X NC are used by the processor 110 for non-interleaved preprocessing as described above to estimate the echo components Y1…Y M in the M microphone signals D1…D M . It can be understood that the N source signals X1…X N can be the same as the N speaker signals X 1C …X NC . Alternatively, the N speaker signals X 1C …X NC can be different from the N source signals X1…X N but can reflect the N source signals X1…XN of the audio content so that the N speaker signals X 1C … X NC can indicate the sounds played by the speakers. For example, the N source signals X1… X N from the sound source can be subjected to various adjustments and processes, such as gain scaling, and the adjusted and processed signals are used as the N speaker signals X 1C … X NC respectively for non-interleaved preprocessing.
[0226] The processor 110 performs non-interleaved preprocessing 810 on the obtained N speaker signals X 1C … X NC to generate M groups of preprocessed signals X P . The non-interleaved preprocessing 810 may include linear summation, linear subtraction, reference audio signal sorting, serial-parallel filtering mode adjustment, gain adjustment, filtering based on microphone signals, etc. as described above. In the case where at least one preprocessed signal X P includes multiple preprocessed signals, based on the nature of the preprocessed signals, it can be selected to perform serial adaptive filtering or parallel adaptive filtering on them. For example, in the case where multiple preprocessed signals are derived from speaker signals indicating monaural, they can be adaptively processed in a serial manner. In the case where multiple preprocessed signals are derived from speaker signals indicating 5.1 surround sound, they can be adaptively processed in a parallel manner. Therefore, the processor can also judge the correlation of multiple speaker signals before non-interleaved preprocessing and perform corresponding non-interleaved preprocessing based on the judgment result. The M groups of preprocessed signals can be generated using different or the same speaker signals in different or the same non-interleaved preprocessing manners. In other words, the M groups of preprocessed signals can be generated independently of each other, and each group of preprocessed signals can include at least one preprocessed signal, and the at least one preprocessed signal is related to the speaker signals and non-interleaved preprocessing manner selected for this group of preprocessed signals.
[0227] The processor 110 then performs adaptive filtering on the M groups of preprocessed signals X P to generate M echo estimation signals Y E . In one embodiment, the adaptive filtering can convert the M groups of preprocessed signals X P from time-domain signals to frequency signals and use the most LMS or RLS for filtering. For example, for a time-domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can be used to convert it to a frequency-domain signal to achieve a better balance between filtering effect and computational overhead. Of course, other overlap lengths and Fourier transforms with other numbers of points can also be used.
[0228] M microphones 12-1…12-M respectively collect various sounds including audio echo, speech, and noise of the audio played by the speaker, and convert them into M microphone signals and provide them to the processor 110. In one embodiment, the microphone signal D1 collected by the microphone 12-1 includes a speech signal component S1, a noise signal component V1, and an echo signal component Y1, where the echo signal component Y1 includes a reflected echo signal component and a direct sound signal component… The microphone signal D collected by the microphone 12-M M includes a speech signal component S M , a noise signal component V M , and an echo signal component Y M , where the echo signal component Y M includes a reflected echo signal component and a direct sound signal component. Therefore, the microphone signal D is a composite signal of the speech signal component S, the noise signal component V, and the echo signal component Y. The M microphone signals D1……D M correspond to the M echo estimation signals Y E . For example, the first echo estimation signal in the M echo estimation signals Y E corresponds to the first microphone signal D1… The Mth echo estimation signal in the M echo estimation signals Y E corresponds to the first microphone signal D M .
[0229] The processor 110 respectively filters out the M echo estimation signals from the M microphone signals to generate M target signals T. For example, the first echo estimation signal is filtered out from the first microphone signal D1 to obtain the first target signal… The Mth echo estimation signal is filtered out from the Mth microphone signal D M to obtain the Mth target signal. The M target signals T can also be used to update the adaptive filter 440 to improve the accuracy of echo estimation. For example, the first target signal is used to update the adaptive filter corresponding to the first echo estimation signal… The Mth target signal is used to update the adaptive filter corresponding to the Mth echo estimation signal.
[0230] In one embodiment, the target signal T is subsequently used to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient through the network. In another embodiment, the target signal T can subsequently be further filtered through conventional schemes such as non-linear filtering schemes, machine learning schemes, etc. to obtain a better echo filtering effect, and then used to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient through the network. In addition, the target signal T can also be used to update the adaptive filter to improve the accuracy of echo estimation.
[0231] Figure 9 is Figure 8Schematic diagram of the summation process 810-1 of an embodiment of non-interleaved preprocessing. The processor 110 processes the first speaker signal X among the N speaker signals 1C and the second speaker signal X 2C to perform a summation process 812 to generate a summation signal X 12 . In one embodiment, the summation signal X 12 can be used as a preprocessing signal for adaptive filtering. By linear summation, two audio signals can be combined into a single audio signal. In this way, subsequent adaptive filtering of the processor 110 only needs to be performed on the single combined audio signal, thus reducing the subsequent computational operation overhead.
[0232] The summation signal X 12 can undergo other preprocessing before being adaptively filtered. In another embodiment, the processor 110 can adjust the gain 814 of the summation signal X 12 to generate a gain-adjusted preprocessing signal X 12A . By the gain adjustment 814, the gain of the final echo estimation signal can be made to match or be comparable to the gain of the echo component Y in the microphone signal, thereby more effectively filtering out the echo component. Alternatively, the gains of the first speaker signal X 1C and the second speaker signal X 2C can be adjusted before the summation process 812. In addition, the processor 110 can also average the first speaker signal X 1C and the second speaker signal X 2C to generate an averaged signal. It can be understood that averaging is equivalent to halving the gain after summation, so averaging can be an alternative specific implementation of summation.
[0233] Although only two speaker signals among the N speaker signals are described herein, it can be understood that the scope of the present disclosure is not limited thereto. In other embodiments, more speaker signals can be used for summation, or signals subjected to non-interleaved preprocessing can be used for summation.
[0234] Figure 10 is Figure 8 Schematic diagram of the difference process 810-2 of another embodiment of non-interleaved preprocessing. The processor 110 performs a difference process 816 on the first speaker signal X among the N speaker signals 1C and the second speaker signal X 2C to generate a difference signal X 21 . In one embodiment, the difference signal X 21 can be used as a preprocessing signal for adaptive filtering. By subtracting the first speaker signal X 1C from the second speaker signal X 2CCombined into a single difference signal X 21 , subsequent filtering by the processor 110 only needs to be performed on the difference signal X 21 , thus reducing the subsequent computational operation overhead. In addition, in some cases, echo cancellation focuses on a certain frequency band or the difference between the outputs of different speakers. In this case, the difference signal X 21 can be separately provided to the processor 110 or combined with other speaker signals to further improve the echo cancellation effect.
[0235] Difference signal X 21 can undergo other preprocessing before being adaptively filtered. In another embodiment, the processor 110 can adjust the gain 814 of the difference signal X 21 to generate a gain-adjusted preprocessing signal X 21A . By adjusting the gain, the gain of the final echo estimation signal can be made to match or be comparable to the gain of the microphone signal, thereby more effectively filtering out the echo component. Alternatively, the gains of the first speaker signal X 1C and the second speaker signal X 2C can be adjusted before the difference processing 816. Although only two speaker signals among the N speaker signals are described herein, it can be understood that the scope of the present disclosure is not limited thereto. In other embodiments, more speaker signals can be used for difference calculation. For example, the first speaker signal is subtracted from the second speaker signal, and the third speaker signal and the fourth speaker signal are subtracted. In another embodiment, the third speaker signal can also be subtracted from the difference signal X 21 again.
[0236] Figure 11 is Figure 8 a schematic diagram of the sorting process 810-3 of another embodiment of the non-interleaved preprocessing in Figure 9 . In one embodiment, the processor 110 performs a sorting process 818 on the gain-adjusted sum signal X 12A in Figure 10 , the gain-adjusted difference signal X 21A in and the third speaker signal X 3C among the N speaker signals. The processor 110 sorts the difference signal X 21A as the first preprocessing signal, sorts the sum signal X 12A as the second preprocessing signal, and sorts the third speaker signal X 3C as the third preprocessing signal. The first preprocessing signal, the second preprocessing signal, and the third preprocessing signal are adaptively filtered in sequence.
[0237] In some cases, it is beneficial to sort different speaker signals and / or preprocessed signals and perform echo cancellation according to the sorting result. For example, signals that may generate relatively large distortion include signals with a relatively high low-frequency content. In the case where the difference signal X 21A indicates low-frequency audio, filtering the low-frequency audio component first can significantly improve the echo cancellation effect. By sorting the signals that may generate relatively large distortion as the priority sorting signals to be adaptively filtered subsequently, the filtering effect can be significantly improved. It can be understood that Figure 11 the sorting shown is only illustrative, and there can be other combinations and sortings of speaker signals and preprocessed signals. For example, N speaker signals can be directly sorted, for example, based on the low-frequency content in each speaker signal, and the sorted speaker signals can be serially used for adaptive filtering.
[0238] Figure 12 is a schematic diagram of a serial filtering process 1200 for echo cancellation according to another embodiment of the present disclosure. The processor 110 duplicates N sound source signals X1…X N --N speaker signals X 1C …X NC for non-interleaved preprocessing 810 to generate N preprocessed signals X P1 …X PN . The processor 110 then performs adaptive processing 440-1 on the first preprocessed signal X P1 …X PN among the N preprocessed signals to generate a first echo estimation signal Y P1 . The processor 110 then filters 470-1 the first echo estimation signal Y E1 from the microphone signal D (e.g., the first microphone signal) to generate a first residual signal E1. E1
[0239] The processor 110 performs adaptive processing on the second preprocessed signal among the N preprocessed signals X P1 …X PN to generate a second echo estimation signal. The processor 110 then filters the second echo estimation signal from the first residual signal E1 to generate a second residual signal. And so on, until the Nth residual signal E N is generated as the final residual signal. In one embodiment, the residual signal E N is then used as the target signal T to wake up the engine to wake up the intelligent assistant or is transmitted to the call recipient via the network. In another embodiment, the residual signal E NSubsequently, the target signal T can be further filtered through conventional schemes such as non - linear filtering schemes, machine learning schemes, etc. to obtain a better echo filtering effect, and then be used to wake up the engine to wake up the intelligent assistant or be transmitted to the call recipient through the network. The N residual signals E1…E N can be respectively used to update the corresponding adaptive filters 440 - 1…440 - N to make the echo estimation signal closer to the echo component in the microphone signal.
[0240] In the case where the similarity between multiple speaker signals is high, serial adaptive filtering and estimation signal filtering can obtain a better filtering effect and better protect the voice input of the microphone 112. Although the principle of serial filtering is described with a single microphone signal D in Figure 12 , this is only illustrative and does not limit the scope of the present disclosure. For example, M microphone signals can be respectively subjected to the above - mentioned serial filtering, and the corresponding target signals are finally synthesized into a voice signal.
[0241] Figure 13 is a schematic diagram of a parallel filtering process 1300 for echo filtering according to another embodiment of the present disclosure. The processor 110 performs non - interleaving pre - processing 810 on N copies of the sound source signals X1…X N -- N speaker signals X 1C …X NC to generate N pre - processed signals. The N pre - processed signals are combined into a matrix signal X PM for parallel adaptive processing. For example, the N pre - processed signals can be combined into an N - dimensional matrix for parallel adaptive filtering. Whether to perform serial adaptive filtering or parallel adaptive filtering on the multiple pre - processed signals can be selected based on the similarity between multiple speaker signals or pre - processed signals. In one embodiment, the similarity between multiple speaker signals is high. For example, the electronic device 100 plays a mono audio. That is, the N speakers 14 - 1…14 - N play the same audio. In this case, the N speaker signals X N associated with the N sound source signals X1…X 1C …X NC can be serially adaptively filtered to obtain a better filtering effect and better protect the voice input of the microphone group 12. In another embodiment, the similarity between multiple speaker signals is low. For example, the electronic device 100 plays stereo or 5.1 surround sound audio. In this case, the N pre - processed signals indicating each channel can be combined into a matrix for parallel adaptive filtering to obtain a better filtering effect.
[0242] The processor 110 then performs parallel adaptive processing 440 on the matrix signal X PM to generate an echo estimation signal Y E。The processor 110 then filters the echo estimation signal Y from the microphone signal D E to generate a residual signal E. In one embodiment, the residual signal E is then used as a target signal T to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient via a network. In another embodiment, the residual signal E as the target signal T can then be further filtered by conventional schemes such as non-linear filtering schemes, machine learning schemes, etc. to obtain a better echo filtering effect, and then used to wake up the engine to wake up the intelligent assistant or transmitted to the call recipient via a network. The residual signal E can be used to update the adaptive filter 440 to make the echo estimation signal closer to the echo component in the microphone signal.
[0243] In the case where the similarity between the N speaker signals is low, parallel adaptive filtering and estimated signal filtering can obtain a better filtering effect. In one embodiment, the processor 110 can determine the similarity between the N speaker signals before performing non-interleaved preprocessing on the N speaker signals, and select serial filtering or parallel filtering based on the determination result. Although the principle of serial filtering is described with a single microphone signal D in Figure 13 , this is only illustrative and does not limit the scope of the present disclosure. For example, the M microphone signals can be respectively subjected to the above parallel filtering, and the corresponding target signals are finally synthesized into a voice signal.
[0244] Figure 14 is a schematic diagram of an audio signal processing process 1400 for echo filtering according to another embodiment of the present disclosure. The electronic device 100 has a processor 110, M microphones 12-1...12-M, M reverse speakers 15-1...15-M, and N speakers 14-1...14-N, where the M reverse speakers are respectively arranged near the M microphones, and M and N are integers greater than 1. The N sound source signals X1...X from the sound source N are respectively output by the processor 110 to the N speakers 14-1...14-N, and the processor 110 causes the N speakers 14-1...14-N to respectively play the corresponding audio. The sound source is, for example, audio data stored in a storage device in a server communicating via the Internet, audio data stored in a local storage device, or audio data collected by a microphone of another device.
[0245] The M microphones 12-1...12-M collect various sounds including audio echoes, voices, and noises of the audio played by the speakers. In one embodiment, the microphone signals collected by the microphones include a voice signal component, a noise signal component, and an echo signal component remaining after the reverse audio fails to be completely canceled, where the echo signal component Y RIt includes a reflected echo signal component and a direct sound signal component. For example, the microphone signal D1 collected by the first microphone 12-1 includes a voice signal component S1, a noise signal component V1, and an echo signal component Y remaining after the reverse audio cannot be completely cancelled out. 1R … The microphone signal D collected by the Mth microphone 12-M M includes a voice signal component S M , a noise signal component V M and an echo signal component Y remaining after the reverse audio cannot be completely cancelled out. MR . Therefore, the microphone signal is a composite signal of a voice signal component, a noise signal component, and an echo signal component.
[0246] As N source signals X1…X N copies, N speaker signals X 1C …X NC are non-interleaved preprocessed 1410 by the processor 110 to generate M groups of preprocessed signals. The M groups of preprocessed signals are then adaptively filtered 1420-1…1420-M respectively to generate M reverse audio signals. The processor 110 causes the reverse speakers 15-1…15-M to play the reverse audio signals. In one embodiment, the adaptive filtering can convert the M groups of preprocessed signals from time-domain signals to frequency signals and perform filtering using LMS or RLS. For example, for a time-domain signal with a sampling rate of 16 kHz, a 75% overlap length and a 1024-point Fourier transform can be used to convert it to a frequency-domain signal to achieve a better balance between filtering effect and computational overhead. Of course, other overlap lengths and Fourier transforms with other numbers of points can also be used. In addition, copies D C1 ……D CM of the microphone signals D collected by the microphones 12-1……12-M can be used to guide the update of the adaptive filtering 1420-1…1420-M.
[0247] The reverse audio cancels out the echo of the audio played by the N speakers 14-1…14-N outside the M microphones 12-1…12-M, so that the microphone signals collected by the M microphones 12-1…12-M include voice signals, noise, and the remaining echo signal components that are not completely filtered out. For example, the microphone signal D1 includes a voice signal S1, noise V1, and the remaining echo signal component Y that is not completely filtered out. 1R。。。 The microphone signal D M includes a voice signal S M , noise V M and the remaining echo signal component Y that is not completely filtered out. MR .
[0248] On the other hand, the N speaker signals X 1C …XNC is used by the processor 110 to perform the non-interleaved preprocessing 810 as described above for estimating the echo components in the microphone signals collected by the M microphones 12-1...12-M. Although the N sound source signals X1...X from the sound source are shown as being directly provided to the N speakers 14-1...14-N and the N speaker signals X N are shown as being directly used for non-interleaved preprocessing, this is merely illustrative and does not limit the scope of the present disclosure. It can be understood that the N sound source signals X1...X from the sound source 1C …X NC and the N speaker signals X N …X 1C …X NC can be subjected to various adjustments and processes, such as gain scaling, and the adjusted and processed signals are provided to the N speakers 14-1...14-N and used for non-interleaved preprocessing. In other embodiments, the sound source signals X1...X provided to the speakers N and the speaker signals X used for non-interleaved preprocessing 1C …X NC can be different (e.g., different gains), but are correlated such that the speaker signals X 1C …X NC can indicate the sounds played by the speakers 14-1...14-N.
[0249] The processor 110 performs non-interleaved preprocessing 810 on the received N speaker signals X 1C …X NC to generate M groups of preprocessed signals X P . The non-interleaved preprocessing 810 can include linear summation, linear subtraction, reference audio signal sorting, serial / parallel filtering mode adjustment, gain adjustment, and replica signals D based on the microphone signals D1...D M …D C1 …D CMFiltering, etc. In the case where a set of preprocessed signals includes multiple preprocessed signals, serial adaptive filtering or parallel adaptive filtering can be selected based on the nature of the preprocessed signals. For example, in the case where multiple preprocessed signals are derived from a reference signal indicating monaural sound, they can be adaptively processed in a serial manner. In the case where multiple preprocessed signals are derived from a reference signal indicating 5.1 surround sound, they can be adaptively processed in a parallel manner. The M sets of preprocessed signals can be generated using different or the same speaker signals in different or the same non-interleaved preprocessing manners. In other words, the M sets of preprocessed signals can be generated independently of each other, and each set of preprocessed signals can include at least one preprocessed signal that is related to the speaker signal and non-interleaved preprocessing manner selected for that set of preprocessed signals. The processor 110 then performs adaptive filtering 440 on the M preprocessed signals to generate M echo estimation signals. The processor 110 filters out the M echo estimation signals Y from the microphone signals E to generate M residual signals. For example, the processor 110 filters out the M echo estimation signals from the M microphone signals to generate M residual signals E. For example, the first echo estimation signal is filtered out from the first microphone signal D1 to obtain the first residual signal... The Mth echo estimation signal is filtered out from the Mth microphone signal D M to obtain the Mth residual signal.
[0250] In addition, the M residual signals E can also be used to update the adaptive filtering 440 respectively to improve the accuracy of echo estimation. For example, the first residual signal is used to update the adaptive filtering corresponding to the first echo estimation signal... The Mth residual signal is used to update the adaptive filtering corresponding to the Mth echo estimation signal. On the other hand, copies D C1 …D CM of the microphone signals can be used to update the adaptive filtering 1420-1…1420-M to improve the accuracy of echo estimation.
[0251] The processor 110 then can use the N speaker signals X 1C …X NCPerform direct-path filtering on M residual signals E to generate M target signals T. For example, perform direct-path filtering 450 on the first residual signal to generate the first target signal... perform direct-path filtering 450 on the Mth residual signal to generate the Mth target signal. In one embodiment, the M target signals T are then used by the wake-up engine to wake up the intelligent assistant or transmitted to the call recipient via a network. In another embodiment, the M target signals T can then be further filtered by conventional schemes such as non-linear filtering schemes, machine learning schemes, etc. to obtain a better echo cancellation effect, and then used by the wake-up engine to wake up the intelligent assistant or transmitted to the call recipient via a network.
[0252] Figure 15 FIG. is a schematic block diagram of an echo cancellation device 1500 according to an embodiment of the present disclosure. The device 1500 is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The device 1500 includes a first acquisition module 1502 for acquiring N speaker signals corresponding to the N speakers. The device 1500 further includes a second acquisition module 1504 for acquiring M microphone signals corresponding to the M microphones. The device 1500 further includes a direct-path filtering module for performing at least direct-path filtering on the N speaker signals and the M microphone signals to obtain target signals. By using direct-path filtering, the echo cancellation effect can be further improved.
[0253] Although only three modules are shown in Figure 15 it can be understood that this is only illustrative and does not limit the scope of the present disclosure. The device 1500 may further include corresponding modules for performing each step of the above-described method 300, method 500, and / or method 700, such as at least one of the modules described in the tenth aspect of the present disclosure.
[0254] Figure 16 FIG. is a schematic block diagram of an echo cancellation device 1600 according to another embodiment of the present disclosure. The device 1600 is applied to an electronic device. The electronic device includes M microphones and N speakers, where both M and N are integers greater than 1. The device 1600 includes: an acquisition module 1602 for acquiring N speaker signals corresponding to the N speakers; a reverse speaker signal generation module 1604 for generating a reverse speaker signal based on the N speaker signals; a playback enable module 1606 for causing a reverse speaker near at least one of the M microphones to play reverse audio based on the reverse speaker signal to cancel the audio output corresponding to the N speaker signals played by the N speakers. The reverse speaker is different from the N speakers. By causing the reverse speaker to play reverse audio, a part of the echo component can be filtered out before the echo is collected by the microphone, thereby providing an echo cancellation effect.
[0255] Although only three modules are shown in Figure 16 it is to be understood that this is for illustration only and not to be construed as a limitation on the scope of the present disclosure. The apparatus 1600 may also include corresponding modules for performing the respective steps in the above-described method 300, method 500, and / or method 700, such as at least one of the modules described above with respect to the eleventh aspect of the present disclosure.
[0256] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for echo cancellation, characterized in that, The method is applied to an electronic device (100), the electronic device (100) includes M microphones (12-1... 12-M) and N speakers (14-1... 14-N), both M and N are integers greater than 1, and the method includes: Obtain N speaker signals (X 1C …X NC ) corresponding to N speakers (14-1…14-N); Obtain M microphone signals (D1... D M ) corresponding to M microphones (12-1…12-M); and Perform non-interleaved preprocessing (810) on the N speaker signals (X 1C … X NC ) to generate at least one preprocessed signal (X P ), where the non-interleaved preprocessing (810) indicates that multiple audio frames corresponding to the same time slot of the N speaker signals (X 1C … X NC ) are preprocessed simultaneously in a combined manner or the N speaker signals (X 1C … X NC ) are preprocessed sequentially and the adaptive filter is updated using a target signal corresponding to a single speaker signal; Adaptive filtering (440) is performed on the at least one preprocessed signal (X P ) to generate an echo estimation signal (Y E ); Filter the echo estimation signal (Y M ) from the M microphone signals (D1…D E ) to generate a residual signal (E); perform direct sound filtering (450) on the residual signal (E) to obtain a target signal (T), where the direct sound filtering indicates filtering of an audio component that is directly output from the N speakers to the M microphones without environmental reflection.
2. The method according to claim 1, wherein wherein the target signal (T) is used to wake up the intelligent voice assistant through the wake-up engine or to be transmitted to another electronic device for a voice call.
3. The method according to claim 1, wherein wherein the target signal (T) contains fewer echo components than the M microphone signals (D1... D M ), and the echo components are used to characterize the echoes of the sound propagated in space by the N speaker signals (X 1C ... X NC ) collected by the M microphones (12-1... 12-M).
4. The method according to claim 1, characterized in that, It further includes: causing a display screen (194) of the electronic device (100) to display a customized direct sound filtering interface; receiving user input on the customized direct sound filtering interface; in response to the user input, obtaining N speaker test signals and causing the N speakers (14-1... 14-N) to play the N speaker test signals; obtaining M microphone test signals corresponding to the M microphones (12-1... 12-M); and storing a customized direct sound filtering model, which is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering.
5. The method according to claim 4, wherein wherein the customized direct sound filtering interface displays a prompt item for prompting to keep the environment quiet.
6. The method according to claim 1, characterized in that, wherein the direct sound filtering (450) includes default direct sound filtering, and the default direct sound filtering indicates filtering based at least on a model relationship between N speaker signals played by the N speakers (14-1... 14-N) and M microphone signals directly collected by the M microphones (12-1... 12-M) in an anechoic environment.
7. The method according to claim 1, wherein It further includes: Generate an inverse speaker signal (X 1C … X NC ) based on the N speaker signals (X 1R … X MR ); and Cause the reverse speakers (15-1…15-M) near the M microphones (12-1…12-M) to play reverse audio based on the reverse speaker signals (X 1R …X MR ) to cancel the echo of the audio output corresponding to the N speaker signals (X 1C …X NC ) played by the N speakers (14-1…14-N), where the reverse speakers (15-1…15-M) are different from the N speakers (14-1…14-N).
8. The method according to claim 1, wherein wherein generating the at least one preprocessed signal (X P ) comprises: The N loudspeaker signals (X 1C …X NC ) of at least two loudspeaker signals (X 1C ,X 2C ) is linearly summed to generate a sum signal (X 12 );as well as At least two of the N speaker signals (X 1C , X 2C ) are linearly differenced to generate a difference signal (X 21 ).
9. The method according to claim 8, wherein wherein generating the at least one preprocessed signal (X P ) further includes sorting the summation signal (X 12 ) and the difference signal (X 21 ); Generating the echo estimation signal (Y E ) further includes adaptively filtering the sorted summation signal (X 12 ) and the difference signal (X 21 ) in sequence to generate echo estimation signals corresponding to the sorting; and Generating the residual signal (E) includes sequentially filtering the echo estimation signals in the corresponding order from the M microphone signals (D1…D M ) to generate the residual signal (E).
10. The method according to claim 1, wherein where generating the at least one preprocessed signal (X P ) includes generating a sorted set of N preprocessed signals (X 1C … X NC ) by sorting the N speaker signals (X P1 … X PN ); Generating the echo estimation signal (Y E ) includes performing adaptive filtering on the sorted N preprocessed signals (X P1 …X PN ) in sequence to generate the N echo estimation signals (Y E1 …Y EN ) corresponding to the sorting; and Generating the residual signal (E) includes sequentially filtering the N echo estimation signals in the corresponding order from the M microphone signals (D1…D M ) to generate the residual signal (E).
11. The method according to claim 10, characterized in that, Among them, sorting the N speaker signals (X 1C … X NC ) includes: Based on the low-frequency components of each of the N speaker signals (X 1C … X NC ), sort the N speaker signals (X 1C … X NC ).
12. The method according to claim 1, wherein where generating the at least one preprocessed signal (X P ) includes performing non-interleaved preprocessing (810) on the N speaker signals (X 1C …X NC ) and the M microphone signals (D1…D M ) to generate the at least one preprocessed signal (X P ).
13. The method according to claim 12, characterized in that, It further includes: Adjust the gain of at least one of the N speaker signals (X 1C … X NC ), the M microphone signals (D1 … D M ), and the at least one preprocessed signal (X P ) such that the gain of the echo estimation signal (Y E ) matches the gain of the M microphone signals (D1 … D M ).
14. An electronic device (100), characterized in that, It includes: N speakers (14-1... 14-N), N is an integer greater than 1; M microphones (12-1... 12-M), M is an integer greater than 1; one or more processors (110); and a memory (121) storing one or more programs, the one or more programs being configured to be executed by the one or more processors (110), and the one or more programs include instructions for executing the method according to any one of claims 1-13.
15. A computer-readable storage medium (121), characterized in that, Storing one or more programs, the one or more programs being configured to be executed by one or more processors (110) of the electronic device (100), and the one or more programs include instructions for executing the method according to any one of claims 1-13.
16. An apparatus for filtering echo, characterized in that, The device is applied to an electronic device (100), the electronic device (100) includes M microphones (12-1... 12-M) and N speakers (14-1... 14-N), both M and N are integers greater than 1, and the device includes: The first acquisition module is configured to acquire N speaker signals (X 1C …X NC ) corresponding to N speakers (14-1…14-N); A second acquisition module, configured to acquire M microphone signals (D1... D M ) corresponding to M microphones (12-1... 12-M); An echo estimation module for performing non-interleaved preprocessing (810) on the N speaker signals (X 1C … X NC ) to generate at least one preprocessed signal (X P ), where the non-interleaved preprocessing (810) indicates that multiple audio frames corresponding to the same time slot of the N speaker signals (X 1C … X NC ) are preprocessed simultaneously in a combined manner or the N speaker signals (X 1C … X NC ) are preprocessed sequentially and the adaptive filter is updated using a target signal corresponding to a single speaker signal; performing adaptive filtering (440) on the at least one preprocessed signal (X P ) to generate an echo estimation signal (Y E ); Residual signal generation module, configured to filter out the echo estimation signal (Y M ) from the M microphone signals (D1…D E ) to generate a residual signal (E); a target signal generation module for performing direct sound filtering (450) on the residual signal (E) to obtain a target signal (T), and the direct sound filtering indicates filtering of an audio component that is directly output from the N speakers to the M microphones without environmental reflection.
17. The device according to claim 16, wherein wherein the target signal (T) is used to wake up the intelligent voice assistant through the wake-up engine or to be transmitted to another electronic device for a voice call.
18. The device according to claim 16, characterized in that, wherein the target signal (T) contains fewer echo components than the M microphone signals (D1... D M ), and the echo components are used to characterize the echoes of the sound propagated in space by the N speaker signals (X 1C ... X NC ) collected by the M microphones (12-1... 12-M).
19. The device according to claim 16, characterized in that, It further includes: A display enabling module, configured to enable a display screen (194) of the electronic device (100) to display a customized direct sound filtering interface; An input receiving module, configured to receive user input on the customized direct sound filtering interface; A speaker test enabling module, configured to obtain N speaker test signals and cause the N speakers (14-1…14-N) to play the N speaker test signals in response to the user input; A third obtaining module, configured to obtain M microphone test signals corresponding to the M microphones (12-1…12-M); And A storage module, configured to store a customized direct sound filtering model, where the customized direct sound filtering model is obtained based on the N speaker test signals and the M microphone test signals, and the customized direct sound filtering model is used for direct sound filtering.
20. The device according to claim 19, wherein Wherein the customized direct sound filtering interface displays a prompt item for prompting to keep the environment quiet.
21. The device according to claim 16, characterized in that, Wherein the direct sound filtering (450) includes a default direct sound filtering module, and the default direct sound filtering indicates filtering based at least on a model relationship between N speaker signals played by the N speakers (14-1…14-N) and M microphone signals directly collected by the M microphones (12-1…12-M) in an anechoic environment.
22. The device according to claim 16, wherein The apparatus further includes: An inverse loudspeaker signal generation module for generating an inverse loudspeaker signal (X 1C … X NC ) based on the N loudspeaker signals (X 1R … X MR ); and The playback enabling module is configured to cause the reverse speakers (15-1...15-M) near at least one of the M microphones (12-1...12-M) to play reverse audio based on the reverse speaker signals (X 1R …X MR ) so as to cancel the echo of the audio output corresponding to the N speaker signals (X 1C …X NC ) played by the N speakers (14-1...14-N), wherein the reverse speakers (15-1...15-M) are different from the N speakers (14-1...14-N).
Citation Information
Patent Citations
Echo cancellation
CN101040512A
Speech enhancement method used for loudspeaking communication system
CN106448691A
Noise reduction earphone configuration method and device, intelligent terminal and noise reduction earphone
CN111107461A