A voice signal processing method, device and equipment
By performing frequency division compression and echo cancellation on the audio source signal of the audio device, the problem of low recognition rate and wake-up rate of voice wake-up function in smart audio devices is solved, improving the accuracy and efficiency of voice recognition, especially when playing at high volume.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
- Filing Date
- 2024-07-02
- Publication Date
- 2026-05-08
AI Technical Summary
In the existing technology, the voice wake-up function of smart audio devices has poor echo cancellation, resulting in low voice recognition rate and wake-up rate. In particular, the signal-to-noise ratio is low when playing at high volume, making it difficult to wake up accurately or causing incorrect wake-up.
By performing frequency-division compression on the audio source signal to be played, selecting a portion of the frequency band for signal strength compression, generating a reference signal for echo cancellation, and increasing the recognition weight of the compressed frequency band during speech recognition, the signal-to-noise ratio is improved.
It improves the voice recognition rate and wake-up rate, especially significantly improving the accuracy and efficiency of voice wake-up under high volume playback, balancing sound quality and recognition rate.
Smart Images

Figure CN118782064B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, specifically to a speech signal processing method, apparatus, and electronic device. This application also relates to an audio device and a voice wake-up system. Background Technology
[0002] With the development of speech signal processing technology, voice wake-up has gradually become an important voice interaction method. Smart audio devices that provide voice wake-up functionality (such as smart speakers, smartphones, and smart TVs) have both speakers and microphones to collect and play sound signals. Smart audio devices often integrate other audio playback functions, so simultaneous speaker playback and voice wake-up is a common application scenario. The microphone picks up the speaker signal and other audio signals, such as the user's voice wake-up signal. The speaker signal picked up by the microphone is called the echo. Echo cancellation is a crucial factor affecting the recognition rate of voice wake-up signals.
[0003] In existing technologies, smart audio devices generally use acoustic echo cancellers (AECs) for voice recognition of audio signals such as voice wake-up signals. The reference signal of the acoustic echo canceller (AEC) is the signal before the speaker, such as the signal after power amplification or the signal before digital-to-analog conversion. However, this reference signal differs significantly from the actual signal after nonlinear distortion caused by the speaker path and acoustic cavity structure. Therefore, there is a problem of poor echo cancellation effect, resulting in low voice recognition rate of voice wake-up signals. Especially when playing at high volumes, the speaker signal picked up by the microphone is much larger than the voice wake-up signal, resulting in a very low signal-to-noise ratio, which further reduces the voice recognition rate, making it difficult to wake up the device or causing it to wake up incorrectly.
[0004] Therefore, how to solve the problems of low speech recognition rate and low wake-up rate is a problem that needs to be addressed.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The speech signal processing method provided in this application solves the problems of low speech recognition rate and low speech wake-up rate, and improves the efficiency of human-computer interaction.
[0007] This application provides a speech signal processing method applied to an audio device, the audio device including a speaker and a microphone, comprising: acquiring a first audio source signal to be played; performing frequency division compression processing on the first audio source signal to compress the signal strength of a portion of the frequency band to obtain a second audio source signal; obtaining a reference signal for echo cancellation based on the second audio source signal; and continuing to transmit the second audio source signal and playing it through the speaker; acquiring a speech acquisition signal to be recognized, the speech acquisition signal being an audio signal acquired by the microphone during the playback of the second audio source signal by the speaker, including a target audio signal to be recognized; performing echo cancellation on the speech acquisition signal based on the reference signal; and performing speech recognition based on the echo-cancelled speech signal to obtain a speech recognition result corresponding to the target audio signal to be recognized.
[0008] Optionally, the speech recognition based on the echo-cancelled speech signal includes: increasing the recognition weight of the speech signal corresponding to the portion of the frequency band in the echo-cancelled speech signal in the speech recognition algorithm used for speech recognition.
[0009] Optionally, the step of acquiring the first audio source signal to be played and performing frequency division compression processing on the first audio source signal to compress the signal strength of a portion of the frequency bands to obtain the second audio source signal includes: dividing the total bandwidth of the first audio source signal into a series of non-overlapping frequency bands, and selecting a portion of the frequency bands for signal strength compression; wherein, the step of selecting a portion of the frequency bands for signal strength compression includes: uniformly selecting or non-uniformly selecting the frequency bands to be compressed.
[0010] Optionally, the step of dividing the total bandwidth of the first sound source signal into a series of non-overlapping frequency bands and selecting a portion of these frequency bands for signal strength compression includes: performing a Fourier transform on the first sound source signal to obtain a corresponding first frequency signal; dividing the first frequency signal into multiple non-overlapping frequency bands and identifying each frequency band as an index representing the information of that frequency band, with each index being a frequency point of the first frequency signal; determining a compression frequency point from the frequency points of the first frequency signal and compressing the frequency band signal identified by the compression frequency point, using the first frequency signal after partial signal compression as a second frequency signal; performing an inverse Fourier transform on the second frequency signal and using the time-domain signal obtained from the inverse Fourier transform as the second sound source signal; the step of obtaining a reference signal for echo cancellation based on the second sound source signal includes: obtaining the reference signal for echo cancellation based on the second sound source signal.
[0011] Optionally, dividing the first frequency signal into multiple non-overlapping frequency bands includes: determining the frequency bandwidth for dividing the frequency bands based on the sampling frequency of the speech signal and the number of points in the Fourier transform, and dividing the first frequency signal into multiple non-overlapping frequency bands according to the frequency bandwidth; the frequency bandwidth is the frequency range of the frequency band; determining the compression frequency point from the frequency points of the first frequency signal includes: uniformly or non-uniformly selecting a portion of the frequency points as the compression frequency point.
[0012] Optionally, it may also include: dynamically setting the compression ratio at the compression frequency point based on the voice signal strength at the compression frequency point, or setting the compression ratio according to a linear relationship with the voice signal strength; wherein the compression ratio represents the signal compression amplitude of the frequency band identified by the compression frequency point; and / or, omitting some compression frequency points.
[0013] Optionally, it also includes: controlling the proportion of the compressed frequency points in the total number of frequency points, so as to reduce the impact of compression processing on the playback sound quality of the first audio source signal.
[0014] Optionally, the step of performing echo cancellation on the voice acquisition signal based on the reference signal and performing speech recognition based on the echo-cancelled voice signal includes: the voice acquisition signal includes the speaker signal and the target audio signal to be recognized; wherein, the target audio signal to be recognized is a voice wake-up signal; the speaker signal is generated by playing a second sound source signal compressed at the compressed frequency point through the speaker; the voice acquisition signal is converted from analog to digital to obtain a digital microphone signal; the reference signal is used as the echo digital estimation signal corresponding to the speaker signal, and the reference signal is removed from the digital microphone signal to obtain a cleaner signal at the compressed frequency point, which is used as the echo-cancelled voice signal; in the voice wake-up algorithm for performing speech recognition, the recognition weight of the voice signal corresponding to the compressed frequency point is increased to identify the voice wake-up word corresponding to the voice wake-up signal.
[0015] Optionally, obtaining the reference signal based on the second audio source signal includes: using the signal obtained after digital-to-analog conversion and power amplification of the second audio source signal, and before it enters the speaker, as a first reference signal; and using the digital reference signal obtained by analog-to-digital conversion of the first reference signal as the reference signal for echo cancellation; or, using the signal of the second audio source signal before digital-to-analog conversion as a second reference signal, and using the second reference signal as the reference signal for echo cancellation.
[0016] This application also provides a speech signal processing device applied to an audio device, the audio device including: a speaker and a microphone, comprising: a frequency division compression unit, configured to acquire a first audio source signal to be played, perform frequency division compression processing on the first audio source signal, so that the signal strength of a portion of the frequency band is compressed to obtain a second audio source signal; a reference signal acquisition unit, configured to obtain a reference signal for echo cancellation based on the second audio source signal, and continue to transmit the second audio source signal and play it through the speaker; a pickup unit, configured to acquire a speech acquisition signal to be recognized, the speech acquisition signal being an audio signal acquired by the microphone during the playback of the second audio source signal by the speaker, including a target audio signal to be recognized; and a recognition unit, configured to perform echo cancellation on the speech acquisition signal based on the reference signal, perform speech recognition based on the echo-cancelled speech signal, and obtain a speech recognition result corresponding to the target audio signal to be recognized.
[0017] This application also provides an audio device, including: a speaker, a microphone, and a speech signal processing device as described above.
[0018] This application also provides a voice wake-up system, including: a speaker, multiple microphones, a frequency division compression module, a reference signal acquisition module, a digital-to-analog converter, a power amplifier, an analog-to-digital converter, an acoustic echo canceller, and a wake-up processing module; the frequency division compression module is used to acquire a first audio source signal to be played, and perform frequency division compression processing on the first audio source signal, so that the signal strength of a portion of the frequency band is compressed to obtain a second audio source signal; the second audio source signal passes through the digital-to-analog converter and enters the power amplifier, and is then output by the power amplifier to the speaker for playback; the reference signal acquisition module is used to obtain a reference signal for echo cancellation based on the second audio source signal; the reference signal is transmitted to the acoustic echo canceller. The system includes: a voice canceller; a plurality of microphones for acquiring a voice acquisition signal to be recognized, the voice acquisition signal being an audio signal acquired by the microphones during the playback of the second sound source signal by the speaker, including a voice wake-up signal; the voice acquisition signal being converted into a digital microphone signal by the analog-to-digital converter; an echo canceller for performing echo cancellation on the digital microphone signal according to the reference signal to obtain an echo-cancelled voice signal, the voice signal being a cleaner voice wake-up signal; and a wake-up processing module for performing voice recognition based on the echo-cancelled voice signal to obtain a voice recognition result corresponding to the voice wake-up signal, the voice recognition result including a wake-up word for triggering the voice wake-up function.
[0019] This application also provides an electronic device, including: a memory and a processor; the memory is used to store a computer program, which, when run by the processor, executes the method provided in this application.
[0020] Compared with the prior art, the advantages of this application are as follows:
[0021] This application provides a speech signal processing method, apparatus, electronic device, and storage medium. The method is applied to an audio device, which includes a speaker and a microphone. The method involves acquiring a first audio source signal to be played, performing frequency-division compression processing on the first audio source signal to compress the signal strength of a portion of the frequency band to obtain a second audio source signal; obtaining a reference signal for echo cancellation based on the second audio source signal; and continuing to transmit the second audio source signal and playing it through the speaker. The method also involves acquiring a speech acquisition signal to be recognized, which is an audio signal acquired by the microphone during the playback of the second audio source signal by the speaker, including a target audio signal to be recognized; performing echo cancellation on the speech acquisition signal based on the reference signal; and performing speech recognition based on the echo-cancelled speech signal to obtain a speech recognition result corresponding to the target audio signal to be recognized. By performing frequency division and partial signal compression on the audio speech signal, and then applying echo cancellation based on the compressed signal, the output power of the compressed portion of the signal is reduced. However, audio signals within the frequency band of the compressed portion, such as the voice wake-up signal, remain unchanged. Therefore, the Sentence Error Rate (SER) during speech recognition is reduced, thereby improving both speech recognition and wake-up rates. Furthermore, by controlling the proportion of the compressed signal to the total signal (bandwidth), the balance between the speaker's playback quality and the echo cancellation effect can be controlled. This ensures a clean audio signal while maintaining playback quality, further improving speech recognition and wake-up efficiency. This approach is particularly suitable for addressing the problem of low voice wake-up / recognition rates under high-volume speaker playback.
[0022] In a preferred approach, during the frequency division and compression of the speech signal, a certain number of frequency points are selected as compression frequency points without affecting the sound quality of the speaker playback. The signal strength of some of these compression frequency points is compressed. During echo cancellation, since the output power of these compression frequency points is compressed, while the voice wake-up signal at these compression frequency points remains unchanged in the acquired speech signal, the signal-to-noise ratio (SNR) at these compression frequency points can be improved. The noise in the SNR refers to the audio signal picked up by the microphone from the speaker playback. Furthermore, by controlling the proportion of compression frequency points in the total frequency points, the impact of signal strength compression on the sound quality of the speaker playback is kept within a certain range. At the same time, the recognition weight at the compression frequency points is increased. This reduces the sentence error rate (SER) when the voice wake-up signal at these compression frequency points is used for speech recognition, thus balancing sound quality and recognition rate. By controlling the compression ratio of the compression frequency points, the SER at the compression frequency points can be significantly reduced, thereby improving the speech signal recognition rate, voice wake-up efficiency, and human-computer interaction efficiency, especially under high-power audio equipment, where the improvement is more pronounced. Attached Figure Description
[0023] Figure 1 This is a flowchart of a speech signal processing method provided in the first embodiment of this application;
[0024] Figure 2 This is a schematic diagram of frequency division compression provided in the first embodiment of this application;
[0025] Figure 3 This is a schematic diagram illustrating the working principle of the audio device provided in the first embodiment of this application;
[0026] Figure 4 This is a schematic diagram of a speech signal processing device provided in the second embodiment of this application;
[0027] Figure 5 This is a schematic diagram of an audio device provided in the third embodiment of this application;
[0028] Figure 6 This is a schematic diagram of a voice wake-up system provided in the fourth embodiment of this application;
[0029] Figure 7 This is a schematic diagram of the electronic device provided in this application. Detailed Implementation
[0030] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0031] This application provides a voice signal processing method, apparatus, electronic device, and storage medium. This application also provides an audio device and a voice wake-up system. These will be described in detail in the following embodiments.
[0032] To facilitate understanding, an application scenario and related concepts of a speech signal processing method are first presented. This scenario is merely one embodiment of the use of the speech signal processing method, and its purpose in providing this scenario embodiment is to facilitate understanding of the application of the speech signal processing method provided in this application, and is not intended to limit the method.
[0033] The speech signal processing method provided in this application is applied to intelligent audio devices with voice wake-up functionality. Intelligent audio devices (hereinafter referred to as audio devices) include, but are not limited to, smart speakers, smartphones, smart TVs, computers, smart wearable devices, and other devices capable of human-computer interaction via voice. The audio device has a speaker and multiple microphones. While the speaker plays audio, the microphones can pick up sound; that is, the microphones pick up the speaker signal played by the speaker and the user's voice signal, such as a voice wake-up signal. The speaker signal picked up by the microphone is called an echo. Echoes interfere with the recognition of the user's voice signal. Before speech recognition processing, echoes are eliminated to obtain a clean speech signal, thereby obtaining an accurate speech recognition result. Applying the speech signal processing method provided in this application, frequency-division compression can be performed on the audio source signal to be played. That is, signals in a portion of the frequency band (i.e., frequency range) are selected according to different frequencies for compression. The strength of audio signals within the frequency band range to which the compressed signal belongs, such as the voice wake-up signal, remains unchanged. Therefore, the signal-to-noise ratio of the audio signal recognition within that frequency band range can be improved, thereby increasing the speech recognition rate and voice wake-up rate.
[0034] Speech wakeup is a speech recognition technology used to wake up smart devices. Users can activate the device, launch its voice assistant, or perform other operations by issuing specific voice commands (usually a word or phrase). The core of speech wakeup is the real-time detection of a specific segment of the speaker within a continuous stream of speech.
[0035] Wake-up rate: This is a core metric related to voice wake-up technology, referring to the probability that a device will be successfully woken up by a wake-up word. A higher wake-up rate indicates better performance, and it is usually expressed as a percentage. Conversely, false wake-up rate refers to the probability that a device will be woken up by a word other than a wake-up word. A higher false wake-up rate indicates poorer performance.
[0036] Acoustic Echo Cancellation (AEC) is an algorithm for eliminating nonlinear echoes. It establishes a speech model of the reference signal based on the correlation between the speaker signal and the echo it produces. The echo is estimated using this model, and the filter coefficients are continuously modified to make the echo estimate closer to the real echo. The echo estimate is then removed from the acquired speech signal, thereby achieving the purpose of eliminating the echo.
[0037] A digital-to-analog converter (DAC) is used to convert analog signals into digital signals.
[0038] An analog-to-digital converter (ADC) is used to convert analog signals into digital signals.
[0039] The Fourier Transform (FT) is used to convert a time-domain signal into a frequency-domain signal. Generally, a sampled analog signal is converted into a discrete frequency-domain signal by performing a Discrete Fourier Transform (DFT) or a Fast Fourier Transform (FFT).
[0040] Inverse Fourier Transform (IFFT): The inverse Fourier transform is used to convert a frequency domain signal into a time domain signal.
[0041] A power amplifier (PA) is used to amplify signals.
[0042] It should be noted that the information disclosed above is only for the purpose of helping to understand this application and does not constitute prior art known to those skilled in the art.
[0043] The following combination Figures 1 to 3 The speech signal processing method provided in the first embodiment of this application will be described. Figure 1 The speech signal processing method shown is applied to an audio device, which includes a speaker and a microphone. The method includes steps S101 to S104.
[0044] Step S101: Obtain the first audio source signal to be played, and perform frequency division compression processing on the first audio source signal so that the signal strength of some frequency bands is compressed to obtain the second audio source signal.
[0045] Step S102: Obtain a reference signal for echo cancellation based on the second sound source signal, and continue to transmit the second sound source signal and play it through the speaker;
[0046] Step S103: Obtain the voice acquisition signal to be recognized, wherein the voice acquisition signal is the audio signal acquired by the microphone during the playback of the second sound source signal by the speaker, including the target audio signal to be recognized;
[0047] Step S104: Perform echo cancellation on the voice acquisition signal based on the reference signal, and perform voice recognition on the voice signal after echo cancellation to obtain a voice recognition result corresponding to the target audio signal to be recognized.
[0048] In this embodiment, an audio frequency division scheme is used to divide the audio speech signal into frequencies, and a portion of the frequency band is selected for signal compression. The frequency division compression process must maintain the same sound quality as the speaker playback. A reference signal is determined based on the frequency division compression. This reference signal is used for echo cancellation of the speech acquisition signal and to adjust the recognition weight of the speech signal within the frequency range of the reference signal's frequency band during speech recognition, thereby achieving a balance between sound quality and speech recognition rate. The frequency range of the compressed portion of the signal is called the compression band. Controlling the compression ratio of the compressed portion of the signal can significantly improve the speech recognition signal-to-noise ratio within the compression band of the speech acquisition signal, thereby improving the speech recognition rate.
[0049] As described in step S101, the first audio source signal to be played undergoes frequency-division compression processing, and a portion of the signal strength is compressed to obtain the second audio source signal. Specifically, the first audio source signal is the audio source signal that enters the speaker path and is played out by the speaker. In this scheme, the audio source signal is processed by frequency-division compression to obtain the second audio source signal. The frequency-division compression processing refers to dividing the first audio source signal into a series of non-overlapping frequency bands, and then selecting a portion of the frequency bands for signal strength compression to reduce the signal strength of the selected frequency bands. Therefore, the second audio source signal can be understood as the speech signal obtained by reducing the signal strength of a portion of the frequency range of the first audio source signal. The portion of the frequency band is determined according to the frequency range of the target audio signal to be identified. For example, if the user's voice is the target audio signal to be identified, then the signal strength of a portion of the frequency band corresponding to the human voice frequency range in the spectrum of the first audio source signal is compressed. In subsequent steps, the speech recognition based on the echo-cancelled speech signal includes: increasing the recognition weight of the speech signal corresponding to the specified frequency band in the echo-cancelled speech signal in the speech recognition algorithm used for speech recognition. Speech recognition algorithms are a branch of artificial intelligence that analyze human speech signals to recognize and convert them into text or commands. Preferably, the first audio source signal is an analog speech signal. Acquiring the first audio source signal to be played and performing frequency-division compression processing on the first audio source signal to compress the signal strength of a portion of the frequency band to obtain the second audio source signal includes: dividing the total bandwidth of the first audio source signal into a series of non-overlapping frequency bands, and selecting a portion of these frequency bands for signal strength compression; wherein, selecting a portion of the frequency bands for signal strength compression includes: uniformly selecting or non-uniformly selecting the frequency bands to be compressed. Further, the step of dividing the total bandwidth of the first sound source signal into a series of non-overlapping frequency bands and selecting a portion of these frequency bands for signal strength compression includes: performing a Fourier transform on the first sound source signal to obtain a corresponding first frequency signal; dividing the first frequency signal into multiple non-overlapping frequency bands and identifying each frequency band as an index representing that frequency band information, with each index being a frequency point of the first frequency signal; determining a compression frequency point from the frequency points of the first frequency signal, and compressing the frequency band signal identified by the compression frequency point, using the partially compressed first frequency signal as a second frequency signal; performing an inverse Fourier transform on the second frequency signal, and using the time-domain signal obtained from the inverse Fourier transform as the second sound source signal; in subsequent steps, obtaining the reference signal for echo cancellation based on the second sound source signal. Preferably, in the voice wake-up scenario, the one issuing the voice wake-up signal includes a person; therefore, determining the compression frequency point from the frequency points of the first frequency signal includes: determining the compression frequency point requiring signal strength compression based on the frequency band where the human voice frequency is located.
[0050] In this embodiment, a frequency point can represent a frequency value or a frequency range. For example, a frequency segment is divided into multiple frequency bands according to a certain frequency interval (e.g., 20Hz), each frequency band being a frequency zone, and each frequency zone is numbered; this number is the frequency point. Preferably, dividing the first frequency signal into multiple non-overlapping frequency zones includes: determining the frequency point bandwidth for dividing the frequency zones based on the sampling frequency of the speech signal and the number of points in the Fourier transform, and dividing the first frequency signal into multiple non-overlapping frequency zones according to the frequency point bandwidth; the frequency point bandwidth is the frequency range of the frequency zone; determining the compression frequency point from the frequency points of the first frequency signal includes: uniformly or non-uniformly selecting a portion of the frequency points as the compression frequency point. In practical applications, when digitizing the audio source signal of analog speech, it is necessary to determine a suitable sampling frequency. Assuming the highest frequency of the first audio source signal is 20kHz, according to the Nyquist sampling theorem, the sampling frequency for sampling this first audio source signal (an analog signal) should be greater than or equal to twice the highest frequency in the analog signal's spectrum, i.e., at least greater than or equal to 40kHz. Therefore, a sampling frequency of 48kHz can be chosen. After sampling, the analog signal undergoes a Fourier transform, such as DFT (Discrete Fourier Transform) or FFT (Fast Fourier Transform), to convert it into a discrete frequency domain signal. The interval between discrete frequency points is related to the number of points selected in the Fourier transform. Taking 4096 points as an example, the discrete frequency interval is 10Hz. In the 0-20kHz speech frequency band, human voice frequencies are generally concentrated in the 200-4kHz band. The frequency band corresponding to this human voice frequency is selected for compression processing; of course, a wider band can also be selected. For example, a compression frequency point can be selected every 100Hz within the 200-4kHz band, resulting in 39 frequency points for compression processing. Frequency selection can also be uniform or non-uniform. Please refer to [reference needed]. Figure 2 The frequency division compression diagram shown in the figure uses the 500-1kHz frequency range within the 200-4kHz band as an example, selecting corresponding frequency points every 100Hz for compression by 10dB. In the figure, downward arrows indicate that signals with frequencies of 500Hz, 600Hz, 700Hz, 800Hz, 900Hz, and 1kHz are compressed.
[0051] In this embodiment, the signal at the compression frequency point is compressed according to a compression ratio, where the compression ratio represents the compression magnitude identified by the compression frequency point. Specifically, it further includes: dynamically setting the compression ratio at the compression frequency point based on the speech signal strength at that frequency point, or setting the compression ratio according to a linear relationship with the speech signal strength; and / or, deleting some compression frequency points. For example, when dynamically setting the compression ratio, the compression ratio can be increased for strong signals. Further, it also includes: controlling the proportion of the compression frequency points in the total number of frequency points to reduce the impact of compression processing on the playback sound quality of the first audio source signal. (Continue to refer to...) Figure 2 In the given example, within the 500-1kHz spectrum processing, frequency points are selected at 100Hz intervals for 10dB compression. Since the signal strength at these compressed frequency points is reduced, selecting only a small number of frequencies from the total number of frequencies for compression has minimal impact on the sound quality played by the speaker. Furthermore, in the subsequent echo cancellation step, because the output power of these compressed frequency points is reduced while the corresponding voice wake-up signal strength remains unchanged, the signal-to-noise ratio of the voice wake-up signal at these compressed frequency points is improved by 10dB, effectively increasing the voice wake-up rate. This improvement is particularly noticeable in high-power audio equipment.
[0052] As described in step S102, a reference signal for echo cancellation is obtained based on the second audio source signal. This reference signal can be understood as the echo estimation signal corresponding to the input speech signal of the audio device. This reference signal needs to be removed from the speech acquisition signal during echo cancellation to obtain a cleaner audio signal, such as a voice wake-up signal, which is then used as the speech signal to be recognized for speech recognition. Obtaining the reference signal based on the second audio source signal includes: using the signal obtained after digital-to-analog conversion and power amplification of the second audio source signal, before it enters the speaker, as a first reference signal; and using the digital reference signal obtained by analog-to-digital conversion of the first reference signal as the reference signal for echo cancellation; or, using the signal of the second audio source signal before digital-to-analog conversion as a second reference signal, and using the second reference signal as the reference signal for echo cancellation. Please refer to [link to relevant documentation]. Figure 3 The figure shows a schematic diagram of the working principle of the audio device. Two working principles are shown in the figure. The only difference between working principle one and working principle two is the method of selecting the reference signal. In working principle one, the acoustic echo canceller (AEC) directly uses the signal REF1 after PA before the speaker as the reference signal. In working principle two, the reference signal is selected by using the signal REF2 before DAC as the reference signal.
[0053] As described in step S103, a voice acquisition signal is acquired for voice recognition. Specifically, the microphone continuously picks up the audio signal played by the speaker; simultaneously, the microphone also synchronously picks up the target audio signal to be recognized. For example, when a user utters a voice wake-up word, the microphone also synchronously picks up the voice wake-up signal uttered by the user. Therefore, the voice acquisition signal may include the audio signal acquired by the microphone from the second sound source signal output by the speaker, and the target audio signal to be recognized.
[0054] As described in step S104, echo cancellation is performed based on the reference signal before speech recognition, thereby obtaining a more accurate speech recognition result. When the target audio signal to be recognized is a voice wake-up signal, the speech recognition result includes the wake-up word corresponding to the voice wake-up signal, which is the wake-up command issued by the user. Preferably, the echo cancellation of the voice acquisition signal based on the reference signal and the speech recognition based on the echo-cancelled voice signal include: the voice acquisition signal includes the speaker signal and the target audio signal to be recognized; wherein, the target audio signal to be recognized is a voice wake-up signal, and the speaker signal is generated by playing a second sound source signal compressed at the compressed frequency point through the speaker; the reference signal is used as the echo digital estimation signal corresponding to the speaker signal, and the reference signal is removed from the digital microphone signal to obtain a cleaner signal at the compressed frequency point, which is used as the echo-cancelled voice signal; in the voice wake-up algorithm used for speech recognition, the recognition weight of the voice signal corresponding to the compressed frequency point is increased to identify the voice wake-up word corresponding to the voice wake-up signal. In this embodiment, the voice wake-up algorithm includes voice recognition technology and command matching technology. The voice recognition technology is used to obtain the voice recognition result of the voice signal corresponding to the compressed frequency point, and the voice recognition result is matched with the voice command to identify the wake-up word.
[0055] Please continue to refer to this. Figure 3The figure shows the working principle of the audio device. Fourier transform (such as FFT), inverse Fourier transform (such as IFFT) and frequency division compression module are added before DAC in the speaker path. The working principle includes: (1) The sound source to be played by the speaker (i.e. the first sound source signal to be played) is converted into a frequency signal (such as the first frequency signal mentioned above) through the Fourier transform module; a certain number of frequency points are selected in the main frequency range of voice wake-up (i.e. human voice frequency, such as 200Hz-3kHz range) through the frequency division compression module as compression frequency points; the selected compression frequency point signal is compressed, for example, by 10dB; wherein, the compression frequency point can be a frequency value, or it can have a certain bandwidth, such as 8Hz. The bandwidth of the compression frequency point depends on the sampling frequency of the voice signal and the number of points of the Fourier transform; wherein, the compression ratio can be set to different proportions according to the volume. Generally, the larger the volume, the larger the compression ratio; wherein, the proportion of the selected frequency points in the total number of frequency points can be controlled. When the proportion is small, the impact on the sound quality is also small. (2) The compressed audio information (i.e., frequency signal) is then converted back to a time domain signal by an inverse Fourier transform module. This time domain signal is the second audio source signal. The second audio source signal is then played through a speaker after passing through a digital-to-analog converter (DAC) and a power amplifier (PA). (3) The microphone continuously picks up the audio signal (i.e., echo) played by the speaker. At the same time, when a user utters a voice wake-up word, the microphone also picks up other audio signals, such as the voice wake-up signal. After picking up the speaker signal and the target audio signal to be identified, such as the voice wake-up signal (both of which are the aforementioned voice acquisition signals to be identified), AEC echo cancellation is performed using the voice acquisition signal and the reference signal. After echo cancellation, the signal enters the wake-up processing module for recognition. The recognition of the voice wake-up signal is mainly carried out by sampling and recognizing the compressed frequency points of the speaker. Since the speaker signals at these frequency points are significantly compressed, the signal-to-noise ratio of the voice wake-up signal at these frequency points is significantly improved compared to the traditional scheme. Thus, a cleaner voice wake-up signal can be obtained at the compressed frequency points, thereby reducing the sentence error rate (SER) and effectively improving the voice wake-up rate. SER means that if a sentence contains a word that is incorrectly identified, the sentence is considered to have been incorrectly identified. The percentage of incorrectly identified sentences out of the total number of sentences is the SER.
[0056] In one example, the frequency spectrum of the audio source signal used for speaker playback is 30Hz–10kHz. A 200Hz–3kHz band is selected (based on the main range of human voice frequencies), with a compression frequency point chosen every 200Hz. The speech signal sampling rate is 8kHz, and the Fourier transform number is 1024; therefore, the bandwidth of the compression frequency points is 8Hz, and the compression ratio is set to 10dB. Since the audio source signal is compressed by 10dB at these compression frequency points, the signal-to-noise ratio at these frequencies is improved by approximately 10dB. In voice wake-up algorithms, increasing the wake-up signal recognition weight for these compression frequency points can significantly improve the voice wake-up rate.
[0057] In this embodiment, frequency-division compression processing is performed on the audio source signal before the speaker plays the audio signal. A reference signal is selected based on the compression result to cancel the echo in the voice acquisition signal collected by the microphone. This achieves frequency division between the audio signal played by the speaker and the target audio signal to be recognized collected by the microphone, solving the problems of low voice recognition rate and low voice wake-up rate. It is particularly suitable for voice wake-up scenarios under high-volume audio playback. Furthermore, by adding a Fourier transform module, a signal compression module, and an inverse Fourier transform module before the audio source signal enters the DAC in the speaker path, and selecting a portion of the frequency points of the audio source with a certain bandwidth for compression processing, or even completely deleting them, the signal-to-noise ratio of key frequency points in the target audio signal recognition frequency band can be significantly improved. Furthermore, by increasing the recognition weight of these frequency points in the voice recognition process, the voice wake-up rate and voice recognition rate can be significantly improved.
[0058] It should be noted that, unless otherwise specified, the features given in this embodiment and other embodiments of this application can be combined with each other, and steps S101 and S102 or similar terms do not limit the steps to be performed in a specific order.
[0059] This concludes the description of the method provided in this embodiment. The method involves frequency division and compression of the audio signal, followed by echo cancellation based on the frequency-division and compression-processed signal. The output power of the compressed portion of the signal is reduced, while the target audio signal to be recognized in the audio acquisition signal with a frequency in the compressed portion remains unchanged. Therefore, the signal-to-noise ratio during speech recognition is improved, and the sentence error rate (SER) is reduced, thereby increasing the speech recognition rate and wake-up rate. Furthermore, by controlling the proportion of the compressed signal to the total signal (bandwidth) and / or the compression ratio of the compressed frequency band, the balance between the speaker's playback sound quality and the echo cancellation effect can be controlled, ensuring playback quality while obtaining a cleaner target audio signal to be recognized, thus improving the speech recognition rate. When the target audio signal to be recognized is a voice wake-up signal, the recognition rate and wake-up rate of the voice wake-up command are improved, thereby solving the problems of difficult wake-up and false wake-up. This method is particularly suitable for solving the problem of low voice wake-up rate / speech recognition rate in scenarios with high-volume speaker playback.
[0060] Corresponding to the first embodiment, the fourth embodiment of this application provides a speech signal processing device; for relevant parts, please refer to the description of the corresponding method embodiment. Figure 4 The figure shows a speech signal processing device applied to an audio device, the audio device including a speaker and a microphone, and the device including:
[0061] Frequency division compression unit 401 is used to acquire the first audio source signal to be played, and to perform frequency division compression processing on the first audio source signal so that the signal strength of some frequency bands is compressed to obtain the second audio source signal.
[0062] The reference signal acquisition unit 402 is configured to obtain a reference signal for echo cancellation based on the second sound source signal, and to continue transmitting the second sound source signal and playing it through the speaker;
[0063] The pickup unit 403 is used to acquire the voice acquisition signal to be recognized, which is the audio signal acquired by the microphone during the playback of the second sound source signal by the speaker, including the target audio signal to be recognized;
[0064] The recognition unit 404 is used to perform echo cancellation on the voice acquisition signal based on the reference signal, and to perform voice recognition on the voice signal after echo cancellation, so as to obtain a voice recognition result corresponding to the target audio signal to be recognized.
[0065] Optionally, the recognition unit 404 is specifically used to: increase the recognition weight of the speech signal corresponding to the part of the frequency band in the speech recognition algorithm used for speech recognition.
[0066] Optionally, the frequency division compression unit 401 is specifically used to: divide the total bandwidth of the first audio source signal into a series of non-overlapping frequency bands, and select a portion of the frequency bands for signal strength compression; wherein, selecting a portion of the frequency bands for signal strength compression includes: uniformly selecting or non-uniformly selecting the frequency bands for compression.
[0067] Optionally, the frequency division compression unit 401 is specifically used for: performing a Fourier transform on the first sound source signal to obtain a corresponding first frequency signal; dividing the first frequency signal into multiple non-overlapping frequency bands, and identifying each frequency band as an index representing the information of that frequency band, with each index being a frequency point of the first frequency signal; determining a compression frequency point from the frequency points of the first frequency signal, and compressing the frequency band signal identified by the compression frequency point, using the first frequency signal after partial signal compression as a second frequency signal; performing an inverse Fourier transform on the second frequency signal, and using the time-domain signal obtained by the inverse Fourier transform as the second sound source signal; the reference signal acquisition unit 402 is specifically used for: obtaining the reference signal used for echo cancellation based on the second sound source signal.
[0068] Optionally, the frequency division compression unit 401 is specifically used to: determine the frequency bandwidth for dividing the frequency band according to the sampling frequency of the speech signal and the number of points of the Fourier transform; divide the first frequency signal into multiple non-overlapping frequency bands according to the frequency bandwidth; the frequency bandwidth is the frequency range of the frequency band; and uniformly or non-uniformly select some frequency points from the frequency points as the compression frequency points.
[0069] Optionally, the frequency division compression unit 401 is specifically used to: dynamically set the compression ratio at the compression frequency point according to the speech signal intensity at the compression frequency point, or set the compression ratio according to a linear relationship with the speech signal intensity; wherein the compression ratio represents the signal compression amplitude of the frequency band identified by the compression frequency point; and / or, empty some compression frequency points.
[0070] Optionally, the frequency division compression unit 401 is specifically used to: control the proportion of the compressed frequency points in the total number of frequency points, so as to reduce the impact of compression processing on the playback sound quality of the first audio source signal.
[0071] Optionally, the recognition unit 404 is specifically configured to: the voice acquisition signal includes the speaker signal and the target audio signal to be recognized; wherein, the target audio signal to be recognized is a voice wake-up signal, and the speaker signal is generated by playing a second sound source signal compressed at the compressed frequency point through the speaker; the voice acquisition signal is converted from analog to digital to obtain a digital microphone signal; the reference signal is used as the echo digital estimation signal corresponding to the speaker signal, and the reference signal is removed from the digital microphone signal to obtain a cleaner signal at the compressed frequency point, which is used as the voice signal after echo cancellation; in the voice wake-up algorithm for voice recognition, the recognition weight of the voice signal corresponding to the compressed frequency point is increased to identify the voice wake-up word corresponding to the voice wake-up signal.
[0072] Optionally, the reference signal acquisition unit 402 is specifically used to: take the signal of the second audio source signal after digital-to-analog conversion and power amplification, and before it enters the speaker, as the first reference signal, and take the digital reference signal obtained by analog-to-digital conversion of the first reference signal as the reference signal for echo cancellation; or, take the signal of the second audio source signal before digital-to-analog conversion as the second reference signal, and take the second reference signal as the reference signal for echo cancellation.
[0073] Based on the above embodiments, one embodiment of this application provides an audio device; relevant parts can be found in the corresponding descriptions of the above embodiments. The audio device 501 includes: the voice signal processing device 502 as described above, n microphones 503, and a speaker 504.
[0074] Based on the above embodiments, one embodiment of this application provides a voice wake-up system; for relevant parts, please refer to the corresponding descriptions of the above embodiments. Figure 6The voice wake-up system shown in the figure includes: a speaker 601, multiple microphones 602, a frequency division compression module 603, a reference signal acquisition module 604, a digital-to-analog converter 605, a power amplifier 606, an analog-to-digital converter 607, an acoustic echo canceller 608, and a wake-up processing module 609. The frequency division compression module acquires a first audio source signal to be played, performs frequency division compression on the first audio source signal, compressing the signal strength of a portion of the frequency band to obtain a second audio source signal. The second audio source signal passes through the digital-to-analog converter and enters the power amplifier, which then outputs it to the speaker for playback. The reference signal acquisition module obtains a reference signal for echo cancellation based on the second audio source signal. The reference signal is transmitted to the acoustic echo canceller; the plurality of microphones are used to acquire the voice acquisition signal to be recognized, the voice acquisition signal being the audio signal acquired by the microphones during the playback of the second sound source signal by the speaker, including a voice wake-up signal; the voice acquisition signal is converted into a digital microphone signal by the analog-to-digital converter; the echo canceller is used to perform echo cancellation on the digital microphone signal according to the reference signal to obtain an echo-cancelled voice signal, the voice signal being a cleaner voice wake-up signal; the wake-up processing module is used to perform voice recognition based on the echo-cancelled voice signal to obtain a voice recognition result corresponding to the voice wake-up signal, the voice recognition result including a wake-up word for triggering the voice wake-up function. Specifically, the reference signal acquisition module is used to: use the signal obtained after digital-to-analog conversion and power amplification of the second sound source signal, before it enters the speaker, as a first reference signal; and use the digital reference signal obtained by analog-to-digital conversion of the first reference signal as the reference signal for echo cancellation; or, use the signal of the second sound source signal before digital-to-analog conversion as a second reference signal, and use the second reference signal as the reference signal for echo cancellation.
[0075] Based on the above embodiments, one embodiment of this application provides an electronic device; for relevant parts, please refer to the corresponding descriptions of the above embodiments. Figure 7 A schematic diagram of an electronic device is shown in the figure. The electronic device includes a memory and a processor. The memory is used to store a computer program, which, after being run by the processor, executes the method provided in the embodiments of this application.
[0076] Based on the above embodiments, one embodiment of this application provides a computer storage medium. For relevant details, please refer to the corresponding descriptions in the above embodiments. The schematic diagram of the computer storage medium is similar to that of an electronic device, and the memory in the diagram can be understood as the storage medium. The computer storage medium stores computer execution instructions, which, when executed by a processor, are used to implement the method provided in the embodiments of this application.
[0077] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data should be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country in which it is located (e.g., the user has given explicit consent, the user has been properly notified, etc.).
[0078] In a typical configuration, an electronic device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0079] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0080] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A speech signal processing method, characterized in that, The method is applied to an audio device, which includes a speaker and a microphone, comprising: The first audio source signal to be played is acquired, and the first audio source signal is subjected to frequency division compression processing, so that the signal strength of a portion of the frequency band is compressed to obtain the second audio source signal. The portion of the frequency band is determined according to the frequency range of the target audio signal to be identified. A reference signal for echo cancellation is obtained based on the second sound source signal, and the second sound source signal is continued to be transmitted and played through the speaker; Acquire the voice acquisition signal to be identified, wherein the voice acquisition signal is the audio signal acquired by the microphone during the playback of the second sound source signal by the speaker, including the target audio signal to be identified; Echo cancellation is performed on the voice acquisition signal based on the reference signal, and voice recognition is performed on the voice signal after echo cancellation to obtain a voice recognition result corresponding to the target audio signal to be recognized.
2. The method according to claim 1, characterized in that, The speech recognition based on the echo-cancelled speech signal includes: In a speech recognition algorithm used for speech recognition, the recognition weight of the speech signal corresponding to the specified frequency band in the echo-cancelled speech signal is increased.
3. The method according to claim 1, characterized in that, The process of acquiring the first audio source signal to be played and performing frequency-division compression processing on the first audio source signal to compress the signal strength of a portion of the frequency band to obtain the second audio source signal includes: The total bandwidth of the first audio source signal is divided into a series of non-overlapping frequency bands, and a portion of these frequency bands are selected for signal strength compression. The step of selecting a portion of the frequency band for signal strength compression includes: uniformly selecting or non-uniformly selecting the frequency band for compression.
4. The method according to claim 3, characterized in that, The step of dividing the total bandwidth of the first audio source signal into a series of non-overlapping frequency bands, and selecting a portion of these frequency bands for signal strength compression, includes: Perform a Fourier transform on the first sound source signal to obtain the corresponding first frequency signal; The first frequency signal is divided into multiple non-overlapping frequency bands, and each frequency band is identified by an index representing the information of that frequency band, and each index is a frequency point of the first frequency signal; Determine the compression frequency point from the frequency points of the first frequency signal, and compress the frequency band signal identified by the compression frequency point. Use the first frequency signal after partial signal compression as the second frequency signal. Perform an inverse Fourier transform on the second frequency signal, and use the time-domain signal obtained from the inverse Fourier transform as the second sound source signal.
5. The method according to claim 4, characterized in that, The step of dividing the first frequency signal into multiple non-overlapping frequency bands includes: The frequency bandwidth used to divide the frequency band is determined based on the sampling frequency of the speech signal and the number of points of the Fourier transform. The first frequency signal is divided into multiple non-overlapping frequency bands according to the frequency bandwidth. The frequency bandwidth is the frequency range of the frequency band. Determining the compressed frequency point from the frequency points of the first frequency signal includes: A portion of the frequency points are selected uniformly or non-uniformly from the frequency points to serve as the compressed frequency points.
6. The method according to claim 4, characterized in that, Also includes: The compression ratio at the compression frequency point can be dynamically set according to the voice signal strength at that compression frequency point, or the compression ratio can be set according to a linear relationship with the voice signal strength; wherein, the compression ratio represents the signal compression amplitude of the frequency band identified by the compression frequency point. And / or, Some compressed frequency points are left empty.
7. The method according to claim 4, characterized in that, Also includes: The proportion of the compressed frequency points in the total number of frequency points is controlled to reduce the impact of compression processing on the playback sound quality of the first audio source signal.
8. The method according to claim 4, characterized in that, The process of performing echo cancellation on the voice acquisition signal based on the reference signal, and performing voice recognition based on the echo-cancelled voice signal, includes: The voice acquisition signal includes the speaker signal and the target audio signal to be identified; wherein, the target audio signal to be identified is a voice wake-up signal; the speaker signal is generated by playing a second sound source signal compressed at the compressed frequency point through the speaker. The voice acquisition signal is converted from analog to digital to obtain a digital microphone signal; The reference signal is used as the echo digital estimation signal corresponding to the loudspeaker signal. The reference signal is removed from the digital microphone signal to obtain a cleaner signal at the compressed frequency point, which is used as the echo-cancelled speech signal. In a voice wake-up algorithm used for speech recognition, the recognition weight of the speech signal corresponding to the compressed frequency point is increased in order to identify the voice wake-up word corresponding to the voice wake-up signal.
9. The method according to claim 3, characterized in that, The step of obtaining a reference signal for echo cancellation based on the second sound source signal includes: The signal obtained by performing digital-to-analog conversion and power amplification on the second audio source signal and before it enters the speaker is used as the first reference signal. The digital reference signal obtained by performing analog-to-digital conversion on the first reference signal is used as the reference signal for echo cancellation. or, The signal before the second audio source signal undergoes digital-to-analog conversion is used as the second reference signal, and the second reference signal is used as the reference signal for echo cancellation.
10. A speech signal processing device, characterized in that, Applied to audio devices, the audio devices including: speakers, microphones, including: The frequency division compression unit is used to acquire the first audio source signal to be played, and to perform frequency division compression processing on the first audio source signal so that the signal strength of a portion of the frequency band is compressed to obtain the second audio source signal. The portion of the frequency band is determined according to the frequency range of the target audio signal to be identified. The reference signal acquisition unit is configured to obtain a reference signal for echo cancellation based on the second sound source signal, and to continue transmitting the second sound source signal and playing it through the speaker; A pickup unit is used to acquire a voice acquisition signal to be identified, wherein the voice acquisition signal is an audio signal acquired by the microphone during the playback of the second sound source signal by the speaker, including the target audio signal to be identified; The recognition unit is used to perform echo cancellation on the voice acquisition signal based on the reference signal, and to perform voice recognition on the voice signal after echo cancellation, so as to obtain a voice recognition result corresponding to the target audio signal to be recognized.
11. An audio device, characterized in that, include: A speaker, a microphone, and the device as described in claim 10.
12. A voice wake-up system, characterized in that, include: Speaker, multiple microphones, frequency division compression module, reference signal acquisition module, digital-to-analog converter, power amplifier, analog-to-digital converter, acoustic echo canceller, wake-up processing module; The frequency division compression module is used to acquire a first audio source signal to be played, and to perform frequency division compression processing on the first audio source signal so that the signal strength of a certain frequency band is compressed to obtain a second audio source signal. The certain frequency band is determined according to the frequency range of the voice wake-up signal. The second audio source signal enters the power amplifier after passing through the digital-to-analog converter, and is then output by the power amplifier to the speaker for playback. The reference signal acquisition module is used to obtain a reference signal for echo cancellation based on the second sound source signal; The reference signal is transmitted to the acoustic echo canceller; The plurality of microphones are used to acquire a voice acquisition signal to be identified, the voice acquisition signal being an audio signal acquired by the microphones during the playback of the second sound source signal by the speaker, including the voice wake-up signal; the voice acquisition signal is converted into a digital microphone signal by the analog-to-digital converter; The echo canceller is used to cancel the echo of the digital microphone signal according to the reference signal to obtain an echo-canceled speech signal, which is a cleaner voice wake-up signal. The wake-up processing module is used to perform speech recognition based on the echo-cancelled speech signal to obtain a speech recognition result corresponding to the speech wake-up signal. The speech recognition result includes a wake-up word used to trigger the speech wake-up function.
13. An electronic device, characterized in that, include: A memory and a processor; the memory is used to store a computer program, which, when executed by the processor, performs the method according to any one of claims 1-9.
Citation Information
Patent Citations
Audio signal collection device and audio signal processing method and device
CN109817238A
Voice processing method and device, electronic equipment and storage medium
CN109862200A
Voice recognition circuit, voice interaction device and household appliance
CN111091818A
Echo cancellation method and system, audio equipment and readable storage medium
CN113178203A