Voice wake-up method and devices therefor
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MIDEA GRP (SHANGHAI) CO LTD
- Filing Date
- 2025-08-15
- Publication Date
- 2026-05-28
AI Technical Summary
In existing technologies, the problem of multiple voice wake-up devices responding simultaneously in the same area results in the inability to accurately wake up the device directly in front of the user, thus affecting the user experience.
By acquiring the first and second channel audio signals of each voice wake-up device, calculating the coherent reverberation signal-to-noise ratio, identifying the target voice wake-up device, and controlling it to interact with the user via voice.
It improves the accuracy of direct wake-up and enhances the user experience.
Smart Images

Figure CN2025115069_28052026_PF_FP_ABST
Abstract
Description
Voice wake-up methods and devices
[0001] This application claims priority to Chinese Patent Application No. 2024117055734, filed on November 25, 2024, entitled “Voice Wake-up Method and Device Thereof,” the entirety of which is incorporated herein by reference. [Technical Field]
[0002] This application relates to the field of voice wake-up technology, specifically to a voice wake-up method and device, the device including a voice wake-up device, a voice wake-up device system, a computer program product, a computer device, and a computer-readable storage medium. [Background Technology]
[0003] With the widespread use of voice-activated wake-up devices, multiple devices may exist simultaneously in the same room or area. When a user utters a wake-up phrase, these devices may respond at the same time. Existing technologies employ various logics for selecting the appropriate wake-up device. A relatively simple logic is proximity wake-up, which activates the nearest device. However, the device activated by proximity may not be the one directly in front of the user. From a user experience perspective, the device the user faces is the one they want to wake up, not the one closest to them. Therefore, the methods described above cannot accurately activate the device directly in front of the user, impacting the user experience. [Summary of the Invention]
[0004] To address the aforementioned problems, this application proposes a voice wake-up method and device, which aims to solve the problems described above.
[0005] To solve the above-mentioned technical problems, one technical solution adopted in this application is: to provide a voice wake-up method, which includes: acquiring a first channel audio signal and a second channel audio signal during the wake-up time period of each voice wake-up device; calculating the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device; determining a target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device; and controlling the target voice wake-up device to perform voice interaction with the user.
[0006] The steps of calculating the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device include: performing a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain a first complex spectrum and a second complex spectrum; and calculating the coherent reverberation signal-to-noise ratio based on the first complex spectrum and the second complex spectrum.
[0007] The step of performing a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain a first complex spectrum and a second complex spectrum includes: obtaining a preset window length and a preset window movement corresponding to the short-time Fourier transform; and performing a short-time Fourier transform on the first channel audio signal and the second channel audio signal with the preset window length and the preset window movement to obtain the first complex spectrum and the second complex spectrum respectively.
[0008] The steps for calculating the coherent reverberation signal-to-noise ratio based on the first complex spectrum and the second complex spectrum include: calculating the first self-spectrum of the first complex spectrum, the second self-spectrum of the second complex spectrum, and the cross-spectrum of the first and second complex spectra based on the first complex spectrum and the second complex spectrum, respectively; calculating coherence based on the first self-spectrum, the second self-spectrum, and the cross-spectrum; obtaining the reverberation field correlation function between the first and second channels of the voice wake-up device; and obtaining the coherent reverberation signal-to-noise ratio based on the reverberation field correlation function and the coherence.
[0009] The steps for obtaining the reverberation field correlation function between the first and second channels of the voice wake-up device include: obtaining the spacing and sound velocity between the first and second channels; and obtaining the reverberation field correlation function based on the frequency, spacing, and sound velocity at each frequency point.
[0010] The step of determining the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device includes: calculating the parameter value of each voice wake-up device during the wake-up time period based on the coherent reverberation signal-to-noise ratio of each voice wake-up device; and taking the voice wake-up device corresponding to the maximum value among all parameter values as the target voice wake-up device.
[0011] The steps of acquiring the first channel audio signal and the second channel audio signal within the wake-up time period of each voice wake-up device include: determining whether the voice wake-up device has received a wake-up word; in response to receiving a wake-up word, determining the wake-up start point and wake-up end point of the voice wake-up device based on the time of receiving the wake-up word, and determining the wake-up time period based on the wake-up start point and wake-up end point, so as to acquire the first channel audio signal and the second channel audio signal within the wake-up time period.
[0012] The step of obtaining the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device includes: performing echo cancellation on the first channel audio signal and the second channel audio signal of each voice wake-up device before the step of obtaining the first channel audio signal and the second channel audio signal of each voice wake-up device.
[0013] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a voice wake-up device, which includes a wake-up module, a coherent reverberation signal-to-noise ratio (SNR) calculation module, and an SNR comparison module. The wake-up module is used to acquire a first-channel audio signal and a second-channel audio signal during the wake-up time period of the voice wake-up device; the coherent reverberation SNR calculation module is used to calculate the coherent reverberation SNR of the voice wake-up device based on the first-channel audio signal and the second-channel audio signal; and the SNR comparison module is used to determine a target voice wake-up device based on the coherent reverberation SNR of multiple voice wake-up devices and send a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
[0014] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a voice wake-up device system, which includes multiple voice wake-up devices and a signal-to-noise ratio (SNR) comparison module. Each voice wake-up device includes a wake-up module and a coherent reverberation SNR calculation module. The wake-up module is used to acquire a first-channel audio signal and a second-channel audio signal during the wake-up time period of the voice wake-up device. The coherent reverberation SNR calculation module is used to calculate the coherent reverberation SNR of the voice wake-up device based on the first-channel audio signal and the second-channel audio signal. The SNR comparison module is used to determine a target voice wake-up device based on the coherent reverberation SNR of the multiple voice wake-up devices and send a response command to the target voice wake-up device to control the target voice wake-up device to interact with the user via voice.
[0015] The signal-to-noise ratio comparison module is integrated into at least one of the multiple voice wake-up devices.
[0016] The voice wake-up device system also includes a central control device. The signal-to-noise ratio comparison module is integrated into the central control device, and the central control device is communicatively connected to each voice wake-up device to obtain the coherent reverberation signal-to-noise ratio of each voice wake-up device. Based on the coherent reverberation signal-to-noise ratio of each voice wake-up device, the central control device uses the signal-to-noise ratio comparison module to determine the target voice wake-up device and sends a response command to the target voice wake-up device to control the target voice wake-up device to interact with the user via voice.
[0017] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer program product, which includes instructions that, when executed by an information processing device, cause the information processing device to perform the voice wake-up method of any of the above embodiments.
[0018] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer device, the computer device including a processor and a memory connected to the processor, wherein the memory stores program data, and the processor executes the program data stored in the memory to execute the voice wake-up method that implements any of the above-mentioned methods.
[0019] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium that stores program instructions internally, which are executed by a processor to implement any of the above-mentioned voice wake-up methods.
[0020] The beneficial effects of this application are as follows: Unlike existing technologies, the voice wake-up method of this application includes: acquiring a first-channel audio signal and a second-channel audio signal during the wake-up time period of each voice wake-up device; calculating the coherent reverberation signal-to-noise ratio (SNR) of each voice wake-up device based on the first-channel and second-channel audio signals respectively; determining a target voice wake-up device based on the SNR of each voice wake-up device; and controlling the target voice wake-up device to interact with the user via voice. Through the above method, the voice wake-up method of this application can calculate the SNR of each voice wake-up device using the first-channel and second-channel audio signals, and determine the target voice wake-up device for direct wake-up based on the SNR, thus improving the effectiveness of direct wake-up. [Attached Image Description]
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a flowchart illustrating the first embodiment of the voice wake-up method provided in this application;
[0023] Figure 2 is a flowchart of an embodiment of step S102 in Figure 1;
[0024] Figure 3 is a flowchart of an embodiment of step S201 in Figure 2;
[0025] Figure 4 is a flowchart of an embodiment of step S202 in Figure 2;
[0026] Figure 5 is a flowchart illustrating an embodiment of step S403 in Figure 4;
[0027] Figure 6 is a flowchart of an embodiment of step S103 in Figure 1;
[0028] Figure 7 is a flowchart of an embodiment of step S101 in Figure 1;
[0029] Figure 8 is a flowchart illustrating the second embodiment of the voice wake-up method provided in this application;
[0030] Figure 9 is a structural schematic diagram of an embodiment of the voice wake-up device of this application;
[0031] Figure 10 is a structural schematic diagram of the first embodiment of the voice wake-up device system of this application;
[0032] Figure 11 is a structural schematic diagram of the second embodiment of the voice wake-up device system of this application;
[0033] Figure 12 is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application.
Detailed Implementation Methods
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0035] With the widespread use of voice-activated wake-up devices, multiple devices may exist simultaneously in the same room or area. When a user utters a wake-up phrase, these devices may respond at the same time. Existing technologies employ various logics for selecting the appropriate wake-up device. A relatively simple logic is proximity wake-up, which activates the nearest device. However, the device activated by proximity may not be the one directly in front of the user. From a user experience perspective, the device the user faces is the one they want to wake up, not the one closest to them. Therefore, the methods described above cannot accurately activate the device directly in front of the user, impacting the user experience.
[0036] To address the aforementioned problems, this application proposes a voice wake-up method. Please refer to Figure 1, which is a flowchart illustrating a first embodiment of the voice wake-up method provided in this application. As shown in Figure 1, the voice wake-up method of this embodiment specifically includes steps S101 to S103:
[0037] Step S101: Acquire the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device.
[0038] In this embodiment, the applicant found that after a user utters a wake-up word, the coherent reverberation signal-to-noise ratio of the voice wake-up device directly in front of the user is greater than that of the voice wake-up device not directly in front of the user among multiple voice wake-up devices in the same room or area.
[0039] Therefore, in this embodiment, the target voice wake-up device directly facing the user can be determined based on the coherent reverberation signal-to-noise ratio of each voice wake-up device. Before obtaining the coherent reverberation signal-to-noise ratio of each voice wake-up device, it is first necessary to obtain the first channel audio signal and the second channel audio signal of each voice wake-up device during the wake-up time period.
[0040] In this embodiment, each voice wake-up device needs to acquire the first channel audio signal and the second channel audio signal during the wake-up time period. Therefore, in this embodiment, each voice wake-up device needs to be equipped with at least two audio collection components. In this embodiment, the audio collection components can be microphone components, that is, in this embodiment, each voice wake-up device needs to be equipped with at least two microphone components to acquire the first channel audio signal and the second channel audio signal respectively during the wake-up time period.
[0041] Step S102: Calculate the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device.
[0042] After each voice wake-up device acquires the first and second channel audio signals within its corresponding wake-up time period, each device can calculate its own coherent reverberation signal-to-noise ratio (SNR) based on its respective first and second channel audio signals. The specific method for calculating the coherent reverberation SNR is described below and will not be detailed further here.
[0043] Step S103: Determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device, and control the target voice wake-up device to interact with the user via voice.
[0044] After obtaining the coherent reverberation signal-to-noise ratio (SNR) of each voice wake-up device, the SNR comparison module can then compare the SNR of each device to determine the target voice wake-up device, which corresponds to the maximum value of the SNR. A response command is then sent to the target device to control it to interact with the user via voice upon receiving the user's wake-up word. The specific comparison method of the SNR comparison module is described below and will not be detailed further here.
[0045] Unlike existing technologies, the voice wake-up method of this application includes: acquiring a first-channel audio signal and a second-channel audio signal during the wake-up time period of each voice wake-up device; calculating the coherent reverberation signal-to-noise ratio (RNR) of each voice wake-up device based on the first-channel and second-channel audio signals; determining a target voice wake-up device based on the RNR of each voice wake-up device; and controlling the target voice wake-up device to interact with the user via voice. Through this method, the voice wake-up method of this application can calculate the RNR of each voice wake-up device using the first-channel and second-channel audio signals, and determine the target voice wake-up device for direct wake-up based on the RNR, thus improving the effectiveness of direct wake-up.
[0046] Optionally, the method for calculating the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device is shown in Figure 2. Please refer to Figure 2, which is a flowchart of an embodiment of step S102 in Figure 1. As shown in Figure 2, this embodiment can implement step S102 through the steps shown in Figure 2, specifically including steps S201 to S202:
[0047] Step S201: Perform a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain the first complex spectrum and the second complex spectrum.
[0048] In this embodiment, each voice wake-up device is equipped with a coherent reverberation signal-to-noise ratio (SNR) calculation module. After the voice wake-up device acquires the first channel audio signal and the second channel audio signal using two microphone components, the coherent reverberation SNR calculation module can perform a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f). In this embodiment, the dimensions of the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f) are (t,f), where t is the time dimension and f is the frequency dimension.
[0049] Step S202: Calculate the coherent reverberation signal-to-noise ratio based on the first complex spectrum and the second complex spectrum.
[0050] After acquiring the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f), the coherent reverberation signal-to-noise ratio calculation module can calculate the coherent reverberation signal-to-noise ratio corresponding to the voice wake-up device itself based on the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f). The specific calculation method is described below.
[0051] Optionally, a method for performing a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain the first complex spectrum and the second complex spectrum is shown in Figure 3. Please refer to Figure 3, which is a flowchart illustrating an embodiment of step S201 in Figure 2. As shown in Figure 3, this embodiment can implement step S201 through the steps shown in Figure 3, specifically including steps S301 to S302:
[0052] Step S301: Obtain the preset window length and preset window movement corresponding to the short-time Fourier transform.
[0053] The Short-Time Fourier Transform (SFT) is a mathematical transformation related to the Fourier Transform, used to determine the frequency and phase of a local sinusoidal wave in a time-varying signal. When performing an SFT on an audio signal, it is first necessary to determine the preset window length and preset window shift. In this embodiment, the preset window length can be set to 1024, and the preset window shift can be set to 512. In other embodiments, the preset window length and preset window shift can be set based on actual conditions and are not limited here.
[0054] Step S302: Perform short-time Fourier transform on the first channel audio signal and the second channel audio signal with a preset window length and preset window movement to obtain the first complex spectrum and the second complex spectrum respectively.
[0055] Once the preset window length and preset window movement are obtained, a short-time Fourier transform can be performed on the first channel audio signal and the second channel audio signal using the preset window length and preset window movement to obtain the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f) respectively.
[0056] Optionally, the method for calculating the coherent reverberation signal-to-noise ratio based on the first complex spectrum and the second complex spectrum is shown in Figure 4. Please refer to Figure 4, which is a flowchart illustrating an embodiment of step S202 in Figure 2. As shown in Figure 4, this embodiment can implement step S202 through the steps shown in Figure 4, specifically including steps S401 to S404:
[0057] Step S401: Calculate the first autospectrum of the first complex spectrum, the second autospectrum of the second complex spectrum, and the cross spectrum between the first and second complex spectra based on the first complex spectrum and the second complex spectrum.
[0058] As mentioned above, after obtaining the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f), the coherent reverberation signal-to-noise ratio calculation module can calculate the first self-spectrum Sxx1(t,f) of the first complex spectrum, the second self-spectrum Sxx2(t,f) of the second complex spectrum, and the cross spectrum Sxy(t,f) between the first and second complex spectra based on the first complex spectrum A1(t,f) and the second complex spectrum A2(t,f).
[0059] The calculation formulas for the first self-spectrum Sxx1(t,f), the second self-spectrum Sxx2(t,f), and the cross-spectrum Sxy(t,f) are shown in formulas (1) to (3): Sxx1(t,f)=A1(t,f) * *A1(t,f) (1) Sxx2(t,f)=A2(t,f) * *A2(t,f) (2) Sxy(t,f)=A1(t,f) * *A2(t,f) (3)
[0060] Where Sxx1(t,f) represents the first autospectrum of the first complex spectrum, and A1(t,f) * Let A1(t,f) be the conjugate complex spectrum of the first complex spectrum; let Sxx2(t,f) be the second autospectrum of the second complex spectrum; and let A2(t,f) be the second autospectrum of the second complex spectrum. * A2(t,f) represents the conjugate complex spectrum of the second complex spectrum; Sxy(t,f) represents the second complex spectrum; and Sxy(t,f) represents the cross spectrum of the first and second complex spectra.
[0061] Step S402: Calculate coherence based on the first autospectrum, the second autospectrum, and the cross-spectrum.
[0062] After obtaining the first self-spectrum Sxx1(t,f), the second self-spectrum Sxx2(t,f), and the cross-spectrum Sxy(t,f), the coherent reverberation signal-to-noise ratio calculation module can calculate the coherence Cxx(t,f) based on the first self-spectrum Sxx1(t,f), the second self-spectrum Sxx2(t,f), and the cross-spectrum Sxy(t,f). The calculation formula for the coherence Cxx(t,f) is shown in formula (4):
[0063] Where Cxx(t,f) represents coherence; Sxy(t,f) represents the cross spectrum of the first complex spectrum and the second complex spectrum; Sxx1(t,f) represents the first self-spectrum of the first complex spectrum; and Sxx2(t,f) represents the second self-spectrum of the second complex spectrum.
[0064] Step S403: Obtain the reverberation field correlation function between the first channel and the second channel of the voice wake-up device.
[0065] After calculating the coherence Cxx(t,f), the coherent reverberation signal-to-noise ratio calculation module also needs to obtain the reverberation field correlation function Cnn(t,f) between the first and second channels of the voice wake-up device. The specific calculation method of the reverberation field correlation function Cnn(t,f) is described below and will not be described in detail here.
[0066] Step S404: Obtain the coherent reverberation signal-to-noise ratio based on the reverberation field correlation function and coherence.
[0067] After obtaining the coherence Cxx(t,f) and the reverberation field correlation function Cnn(t,f), the coherent reverberation signal-to-noise ratio calculation module can calculate the coherent reverberation signal-to-noise ratio CRSNR(t,f) based on the coherence Cxx(t,f) and the reverberation field correlation function Cnn(t,f). The calculation formula for CRSNR(t,f) is shown in formula (5):
[0068] Where CRSNR(t,f) represents the coherent reverberation signal-to-noise ratio, Cxx(t,f) represents coherence, and Cnn(t,f) represents the reverberation field correlation function.
[0069] Optionally, the method for obtaining the reverberation field correlation function between the first channel and the second channel of the voice wake-up device is shown in Figure 5. Please refer to Figure 5, which is a flowchart illustrating an embodiment of step S403 in Figure 4. As shown in Figure 5, this embodiment can implement step S403 through the steps shown in Figure 5, specifically including steps S501 to S502:
[0070] Step S501: Obtain the distance and sound velocity between the first channel and the second channel.
[0071] When the coherent reverberation signal-to-noise ratio calculation module obtains the reverberation field correlation function Cnn(t,f), it needs to obtain the distance d between the first channel and the second channel and the speed of sound c. In this embodiment, the distance d between the first channel and the second channel is the distance between the two microphone components set in the voice wake-up device, and the speed of sound refers to the speed at which sound propagates.
[0072] Step S502: Obtain the reverberation field correlation function based on the frequency, spacing and sound velocity at each frequency point.
[0073] After obtaining the spacing d and the sound velocity c, the coherent reverberation signal-to-noise ratio calculation module can obtain the reverberation field correlation function Cnn(t,f) based on the frequency, spacing d, and sound velocity c at each frequency point. The reverberation field correlation function Cnn(t,f) is shown in formula (6):
[0074] Where Cnn(t,f) represents the reverberation field correlation function, f represents the frequency of each frequency point, d represents the distance between the two microphone components set by the voice wake-up device, and c represents the speed of sound.
[0075] Optionally, the method for determining the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device is shown in Figure 6. Please refer to Figure 6, which is a flowchart illustrating an embodiment of step S103 in Figure 1. As shown in Figure 6, this embodiment can implement step S103 through the steps shown in Figure 6, specifically including steps S601 to S602:
[0076] Step S601: Calculate the parameter values of each voice wake-up device during the wake-up time period based on the coherent reverberation signal-to-noise ratio of each voice wake-up device.
[0077] After obtaining the coherent reverberation signal-to-noise ratio (CRSNR)(t,f) of each voice wake-up device, the coherent reverberation signal-to-noise ratio (CRSNR) calculation module of each voice wake-up device can send its corresponding CRSNR(t,f) to the signal-to-noise ratio comparison module. The signal-to-noise ratio comparison module can calculate the parameter value TotalCRSNR corresponding to each voice wake-up device based on the coherent reverberation signal-to-noise ratio (CRSNR)(t,f) of each voice wake-up device. The calculation formula of the parameter value TotalCRSNR is shown in formula (7):
[0078] Where TotalCRSNR represents the parameter value corresponding to the voice wake-up device, CRSNR(t,f) represents the coherent reverberation signal-to-noise ratio CRSNR(t,f) of the voice wake-up device, T represents the maximum value of time in the time dimension of CRSNR(t,f), and F represents the maximum value of frequency in the frequency dimension of CRSNR(t,f).
[0079] Step S602: Select the voice wake-up device corresponding to the maximum value among all parameter values as the target voice wake-up device.
[0080] After the noise ratio comparison module calculates and obtains the parameter value TotalCRSNR corresponding to all voice wake-up devices based on formula (7), it can compare all parameter values TotalCRSNR and take the voice wake-up device corresponding to the maximum value of all parameter values TotalCRSNR as the target voice wake-up device.
[0081] Optionally, the method for acquiring the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device is shown in Figure 7. Please refer to Figure 7, which is a flowchart illustrating an embodiment of step S101 in Figure 1. As shown in Figure 7, this embodiment can implement step S101 through the steps shown in Figure 7, specifically including steps S701 to S702:
[0082] Step S701: Determine whether the voice wake-up device has received the wake-up word.
[0083] In this embodiment, each voice wake-up device is equipped with a wake-up module. The wake-up module of the voice wake-up device needs to receive the corresponding wake-up word before it controls the voice wake-up device to respond and work. Therefore, in this embodiment, when the voice wake-up device acquires the first channel audio signal and the second channel audio signal during the wake-up time period, the wake-up module needs to determine whether the voice wake-up device has received the wake-up word.
[0084] If the voice wake-up device receives a wake-up word, proceed to step S702; if the voice wake-up device does not receive a wake-up word, it enters standby mode and waits to be woken up.
[0085] Step S702: Determine the wake-up start point and wake-up end point of the voice wake-up device based on the time of receiving the wake-up word, and determine the wake-up time period based on the wake-up start point and wake-up end point, so as to obtain the first channel audio signal and the second channel audio signal within the wake-up time period.
[0086] When the voice wake-up device receives a wake-up word, its wake-up module can determine the wake-up start and end points based on the time the wake-up word is received. Specifically, in this embodiment, the wake-up module includes a wake-up start point detection module and a wake-up end point detection module. The wake-up start point detection module determines the start time of the wake-up time period, and the wake-up end point detection module determines the end time of the wake-up time period. The time period between the start and end times is the wake-up time period. The start time of the wake-up time period can be set to the time when the first character of the wake-up word is received, and the end time of the wake-up time period can be set to the time when the last character of the wake-up word is received. After determining the wake-up time period, the wake-up module can extract the first channel audio signal and the second channel audio signal within the wake-up time period for subsequent processing.
[0087] Optionally, this application further proposes a voice wake-up method. Please refer to Figure 8, which is a flowchart illustrating a second embodiment of the voice wake-up method provided in this application. As shown in Figure 8, the voice wake-up method of this embodiment specifically includes steps S801 to S804:
[0088] Step S801: Perform echo cancellation on the first channel audio signal and the second channel audio signal of each voice wake-up device.
[0089] In this embodiment, before each voice wake-up device acquires the first channel audio signal and the second channel audio signal within its corresponding wake-up time period, it is necessary to perform echo cancellation on the first channel audio signal and the second channel audio signal of each voice wake-up device. Therefore, each voice wake-up device can be equipped with an echo cancellation module to eliminate echo.
[0090] In addition, in this embodiment, echo cancellation is also required when the voice wake-up device makes a broadcast. If the voice wake-up device is in dual-talk mode at the same time, the voice wake-up device needs to use a high suppression mode to suppress nonlinear noise in order to reduce the impact of echo on face wake-up.
[0091] Step S802: Obtain the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device.
[0092] Step S802 is the same as step S101, and will not be described again.
[0093] Step S803: Calculate the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device.
[0094] Step S803 is the same as step S102, and will not be described again.
[0095] Step S804: Determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device, and control the target voice wake-up device to interact with the user via voice.
[0096] Step S804 is the same as step S103, and will not be described again.
[0097] Optionally, this application further proposes a voice wake-up device. Please refer to Figure 9, which is a structural schematic diagram of an embodiment of the voice wake-up device of this application. As shown in Figure 9, the voice wake-up device 100 of this embodiment includes a wake-up module 10, a coherent reverberation signal-to-noise ratio calculation module 20, and a signal-to-noise ratio comparison module 30.
[0098] The wake-up module 10 is used to acquire the first channel audio signal and the second channel audio signal during the wake-up time period of the voice wake-up device 100; the coherent reverberation signal-to-noise ratio calculation module 20 is used to calculate the coherent reverberation signal-to-noise ratio of the voice wake-up device 100 based on the first channel audio signal and the second channel audio signal; the signal-to-noise ratio comparison module 30 is used to determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of multiple voice wake-up devices 100, and send a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
[0099] As mentioned above, the wake-up module 10 is equipped with a wake-up start point detection module and a wake-up end point detection module. The wake-up start point detection module is used to determine the start time of the wake-up time period, and the wake-up end point detection module is used to determine the end time of the wake-up time period. The time period between the start time and the end time is the wake-up time period. In addition, the voice wake-up device 100 is also equipped with an echo cancellation module (not shown in the figure) for echo cancellation of the first channel audio signal and the second channel audio signal of the voice wake-up device 100.
[0100] Optionally, this application further proposes a voice wake-up device system. Please refer to Figure 10, which is a structural schematic diagram of the first embodiment of the voice wake-up device system of this application. As shown in Figure 9, the voice wake-up device system 200 of this embodiment includes a plurality of voice wake-up devices 100 and a signal-to-noise ratio comparison module 30.
[0101] Each voice wake-up device 100 includes a wake-up module 10 and a coherent reverberation signal-to-noise ratio (SNR) calculation module 20. The wake-up module 10 is used to acquire the first channel audio signal and the second channel audio signal during the wake-up time period of the voice wake-up device 100. The coherent reverberation SNR calculation module 20 is used to calculate the coherent reverberation SNR of the voice wake-up device 100 based on the first channel audio signal and the second channel audio signal. The SNR comparison module 30 is used to determine the target voice wake-up device based on the coherent reverberation SNR of multiple voice wake-up devices 100 and send a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
[0102] Optionally, in this embodiment, as shown in FIG10, the signal-to-noise ratio comparison module 30 is integrated into at least one of the multiple voice wake-up devices 100.
[0103] In this embodiment, a signal-to-noise ratio (SNR) comparison module 30 can be set in at least one of the multiple voice wake-up devices 100. After the other voice wake-up devices 100 calculate their own coherent reverberation SNR, they send it to the voice wake-up device 100 equipped with the SNR comparison module 30. The voice wake-up device 100 uses the SNR comparison module 30 to determine the target voice wake-up device based on the coherent reverberation SNR of the multiple voice wake-up devices 100, and sends a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
[0104] Optionally, in other embodiments, please refer to FIG11, which is a structural schematic diagram of the second embodiment of the voice wake-up device system of this application. As shown in FIG11, the voice wake-up device system 200 further includes a central control device 210, and a signal-to-noise ratio comparison module 30 is integrated in the central control device 210. The central control device 210 is communicatively connected to each voice wake-up device 100 to obtain the coherent reverberation signal-to-noise ratio of each voice wake-up device 100. The central control device 210 uses the signal-to-noise ratio comparison module 30 to determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each voice wake-up device 100, and sends a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
[0105] Optionally, this application further proposes a computer program product including instructions that, when executed by an information processing device, cause the information processing device to perform the voice wake-up method of any of the above embodiments.
[0106] Furthermore, if the aforementioned functions are implemented as software functions and sold or used as independent products, they can be stored in a mobile terminal-readable storage medium. That is, this application also provides a storage device storing program data, which can be executed to implement the methods of the above embodiments. This storage device can be, for example, a USB flash drive, an optical disc, or a server. In other words, this application can be embodied in the form of a software product, which includes several instructions to cause a smart terminal to execute all or part of the steps of the methods described in the various embodiments.
[0107] Optionally, this application further proposes a computer device including a processor and a memory connected to the processor.
[0108] A processor can also be called a CPU (Central Processing Unit). A processor may be an integrated circuit chip with signal processing capabilities. A processor can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Memory is used to store program data required for processor operation. The processor is also used to execute the program data stored in memory to implement the voice wake-up method of any of the above embodiments.
[0109] Optionally, this application further proposes a computer-readable storage medium. Please refer to FIG12, which is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application.
[0110] The computer-readable storage medium 300 of this application embodiment stores program instructions 310, which are executed by a processor to implement the voice wake-up method of any of the above embodiments.
[0111] Specifically, program instructions 310 can be formed into a program file and stored in the aforementioned storage medium in the form of a software product, so that an electronic device (which may be a personal computer, server, or network device, etc.) or processor can execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0112] In this embodiment, the computer-readable storage medium 300 may be, but is not limited to, a USB flash drive, SD card, PD optical drive, portable hard drive, large-capacity floppy drive, flash memory, multimedia memory card, server, etc.
[0113] Furthermore, if the aforementioned functions are implemented as software functions and sold or used as independent products, they can be stored in a mobile terminal-readable storage medium. That is, this application also provides a storage device storing program data, which can be executed to implement the methods of the above embodiments. This storage device can be, for example, a USB flash drive, an optical disc, or a server. In other words, this application can be embodied in the form of a software product, which includes several instructions to cause a smart terminal to execute all or part of the steps of the methods described in the various embodiments.
[0114] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0115] Any process or method description in the flowchart or otherwise herein can be understood as representing an apparatus, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0116] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (which may be a personal computer, server, network device, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0117] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A voice wake-up method, wherein, include: Acquire the first and second channel audio signals of each voice wake-up device during the wake-up time period; The coherent reverberation signal-to-noise ratio of each voice wake-up device is calculated based on the first channel audio signal and the second channel audio signal of each voice wake-up device, respectively. The target voice wake-up device is determined based on the coherent reverberation signal-to-noise ratio of each of the aforementioned voice wake-up devices, and the target voice wake-up device is controlled to interact with the user via voice.
2. The voice wake-up method according to claim 1, wherein, The step of calculating the coherent reverberation signal-to-noise ratio of each voice wake-up device based on the first channel audio signal and the second channel audio signal of each voice wake-up device includes: Perform a short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain a first complex spectrum and a second complex spectrum; The coherent reverberation signal-to-noise ratio is calculated based on the first complex spectrum and the second complex spectrum.
3. The voice wake-up method according to claim 2, wherein, The step of performing short-time Fourier transform on the first channel audio signal and the second channel audio signal to obtain the first complex spectrum and the second complex spectrum includes: Obtain the preset window length and preset window movement corresponding to the short-time Fourier transform; The first channel audio signal and the second channel audio signal are subjected to the short-time Fourier transform with the preset window length and the preset window movement to obtain the first complex spectrum and the second complex spectrum respectively.
4. The voice wake-up method according to any one of claims 2-3, wherein, The step of calculating the coherent reverberation signal-to-noise ratio based on the first complex spectrum and the second complex spectrum includes: Calculate the first autospectrum of the first complex spectrum, the second autospectrum of the second complex spectrum, and the cross spectrum between the first complex spectrum and the second complex spectrum based on the first complex spectrum and the second complex spectrum, respectively. Coherence is calculated based on the first autospectrum, the second autospectrum, and the cross-spectrum; Obtain the reverberation field correlation function between the first channel and the second channel of the voice wake-up device; The coherent reverberation signal-to-noise ratio is obtained based on the reverberation field correlation function and the coherence.
5. The voice wake-up method according to claim 4, wherein, The step of obtaining the reverberation field correlation function between the first channel and the second channel of the voice wake-up device includes: Obtain the distance and sound velocity between the first channel and the second channel; The reverberation field correlation function is obtained based on the frequency at each frequency point, the spacing, and the sound speed.
6. The voice wake-up method according to any one of claims 1-5, wherein, The step of determining the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of each of the voice wake-up devices includes: Calculate the parameter values of each of the voice wake-up devices during the wake-up time period based on the coherent reverberation signal-to-noise ratio of each of the voice wake-up devices. The voice wake-up device corresponding to the maximum value among all the parameter values is taken as the target voice wake-up device.
7. The voice wake-up method according to any one of claims 1-6, wherein, The steps of acquiring the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device include: Determine whether the voice wake-up device has received a wake-up word; In response to receiving the wake-up word, the wake-up start point and wake-up end point of the voice wake-up device are determined based on the time of receiving the wake-up word, and the wake-up time period is determined based on the wake-up start point and the wake-up end point, so as to obtain the first channel audio signal and the second channel audio signal within the wake-up time period.
8. The voice wake-up method according to any one of claims 1-7, wherein, Before the steps of acquiring the first channel audio signal and the second channel audio signal during the wake-up time period of each voice wake-up device, the method further includes: Echo cancellation is performed on the first channel audio signal and the second channel audio signal of each of the aforementioned voice wake-up devices.
9. A voice wake-up device, wherein, include: The wake-up module is used to acquire the first channel audio signal and the second channel audio signal during the wake-up time period of the voice wake-up device; A coherent reverberation signal-to-noise ratio calculation module is used to calculate the coherent reverberation signal-to-noise ratio of the voice wake-up device based on the first channel audio signal and the second channel audio signal; The signal-to-noise ratio comparison module is used to determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of the multiple voice wake-up devices, and send a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
10. A voice wake-up device system, wherein, include: Multiple voice wake-up devices are provided, each of which includes a wake-up module and a coherent reverberation signal-to-noise ratio (RNR) calculation module. The wake-up module is used to acquire a first channel audio signal and a second channel audio signal during the wake-up time period of the voice wake-up device. The coherent reverberation RNR calculation module is used to calculate the coherent reverberation RNR of the voice wake-up device based on the first channel audio signal and the second channel audio signal. The signal-to-noise ratio comparison module is used to determine the target voice wake-up device based on the coherent reverberation signal-to-noise ratio of the multiple voice wake-up devices, and send a response command to the target voice wake-up device to control the target voice wake-up device to perform voice interaction with the user.
11. The voice wake-up device system according to claim 10, wherein, The signal-to-noise ratio comparison module is integrated into at least one of the multiple voice wake-up devices.
12. The voice wake-up device system according to any one of claims 10-11, wherein, It also includes a central control device, in which the signal-to-noise ratio comparison module is integrated. The central control device is communicatively connected to each of the voice wake-up devices to obtain the coherent reverberation signal-to-noise ratio of each of the voice wake-up devices. Based on the coherent reverberation signal-to-noise ratio of each of the voice wake-up devices, the central control device uses the signal-to-noise ratio comparison module to determine the target voice wake-up device and sends a response command to the target voice wake-up device to control the target voice wake-up device to interact with the user via voice.
13. A computer program product, wherein, The computer program product includes instructions that, when executed by an information processing device, cause the information processing device to perform the voice wake-up method as described in any one of claims 1-8.
14. A computer device, wherein, The computer device includes a processor and a memory connected to the processor, wherein the memory stores program data, and the processor executes the program data stored in the memory to perform the voice wake-up method according to any one of claims 1-8.
15. A computer-readable storage medium, wherein, It stores program instructions that are executed to implement the voice wake-up method according to any one of claims 1-8.
Citation Information
Patent Citations
Method and device for waking up equipment nearby in whole-house intelligent system and related equipment
CN115662425A
Voice wake-up method and device of terminal and storage medium
CN115831113A
Device wakeup method, electronic device and storage medium
CN116453516A
Voice wake-up method and device
CN119580726A
Direct path acoustic signal selection using a soft mask
US10715909B1