Voice wake-up method and device, equipment and storage medium
By enhancing the target sound signal and detecting wake-up words, the problem of inaccurate sound signal recognition is solved, and the accuracy and efficiency of voice wake-up are improved.
Patent Information
- Application Number
- CN202510399573.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the accuracy of identifying whether the sound signal is a wake-up signal of a device function is insufficient, resulting in low accuracy and efficiency of voice wake-up.
By enhancing the target sound signal, the enhanced target wake-up signal is obtained and wake-up word detection is performed to improve the accuracy of wake-up word detection and optimize the voice wake-up method.
提高了唤醒词检测的准确率和语音唤醒的准确度,降低了计算量,提高了语音唤醒的效率。
Smart Images

Figure CN120260585A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to the technical fields of artificial intelligence such as deep learning. Background Art
[0002] With the development of technology, people increasingly prefer to activate relevant device functions by voice wake-up. Among them, the sound signal of the input device may not be the voice wake-up signal of the device function. In this scenario, it is necessary to analyze the sound signal of the input device to identify whether the sound signal is the wake-up signal of the device function.
[0003] Therefore, it is very important to accurately identify whether the sound signal is the wake-up signal of the device function. Summary of the Invention
[0004] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.
[0005] To this end, a first aspect of the present disclosure proposes a voice wake-up method.
[0006] A second aspect of the present disclosure proposes a voice wake-up device.
[0007] A third aspect of the present disclosure proposes an electronic device.
[0008] A fourth aspect of the present disclosure proposes a computer-readable storage medium.
[0009] A fifth aspect of the present disclosure proposes a chip.
[0010] A first aspect of the present disclosure proposes a voice wake-up method, including: in response to the input of a target sound signal, enhancing the input target sound signal to obtain an enhanced target wake-up signal; performing wake-up word detection on the target wake-up signal to perform voice wake-up on a wake-up object.
[0011] A second aspect of the present disclosure proposes a voice wake-up device, including: an enhancement module, configured to enhance the input target sound signal in response to the input of the target sound signal to obtain an enhanced target wake-up signal; a detection and wake-up module, configured to perform wake-up word detection on the target wake-up signal to perform voice wake-up on a wake-up object.
[0012] A third aspect of the present disclosure proposes an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute instructions to implement the voice wake-up method proposed in the first aspect as above.
[0013] A fourth aspect of the present disclosure provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the voice wake-up method provided in the first aspect above.
[0014] A fifth aspect of the present disclosure provides a chip, including one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal and send the signal to the processor, and the signal includes computer instructions stored in a memory. When the processor executes the computer instructions, the chip executes the steps of the voice wake-up method provided in the first aspect.
[0015] For the voice wake-up method provided in the present disclosure, the target sound signal is enhanced to obtain a target wake-up signal, and then the wake-up word detection is performed based on the target wake-up signal obtained by enhancing the target sound signal, which improves the accuracy of the wake-up word detection, and further improves the accuracy and precision of the voice wake-up of the wake-up object. Compared with the voice wake-up method based on sound source localization in the related art, the calculation amount of the voice wake-up of the wake-up object is reduced, the efficiency of the voice wake-up of the wake-up object is improved, and the voice wake-up method is optimized.
[0016] It should be understood that the content described in the present disclosure is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present disclosure will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0018] Figure 1 is a schematic flowchart of a voice wake-up method according to an embodiment of the present disclosure;
[0019] Figure 2 is a schematic flowchart of a voice wake-up method according to another embodiment of the present disclosure;
[0020] Figure 3 is a schematic flowchart of a voice wake-up method according to another embodiment of the present disclosure;
[0021] Figure 4 is a schematic flowchart of a voice wake-up method according to another embodiment of the present disclosure;
[0022] Figure 5 is a schematic structural diagram of a voice wake-up device according to an embodiment of the present disclosure;
[0023] Figure 6 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0025] A voice wake-up method, apparatus, device, and storage medium proposed in an embodiment of the present disclosure will be described below with reference to the accompanying drawings.
[0026] Figure 1 The flowchart of the voice wake-up method according to an embodiment of the present disclosure is shown in Figure 1 As shown, the method includes:
[0027] S101, in response to the input of a target sound signal, enhance the input target sound signal to obtain an enhanced target wake-up signal.
[0028] In an embodiment of the present disclosure, for a device configured with a voice wake-up function, when the device recognizes that there is an input of a sound signal, it can analyze the input sound signal to identify whether the input sound signal is the wake-up sound signal of the device.
[0029] Among them, the sound signal input to the device can be determined as the target sound signal.
[0030] Optionally, for the input target sound signal, it can be enhanced, and the enhanced target sound signal can be analyzed to improve the accuracy of information acquisition. In this scenario, the input target sound signal can be enhanced based on the audio enhancement processing method in the related art to obtain an enhanced target sound signal, and the enhanced target sound signal can be determined as the target wake-up signal.
[0031] S102, perform wake-up word detection on the target wake-up signal to perform voice wake-up on the wake-up object.
[0032] In an embodiment of the present disclosure, the object for receiving the voice wake-up instruction in the device configured with the voice wake-up function can be determined as the wake-up object.
[0033] As a possible implementation, the preset wake-up strategy of the wake-up object can be obtained, and the target wake-up signal can be analyzed based on the wake-up strategy. When it is recognized that the target wake-up signal matches the wake-up strategy, it can be determined that the target wake-up signal is the voice wake-up signal of the wake-up object.
[0034] In this scenario, the wake-up instruction corresponding to the wake-up object can be obtained based on the target wake-up signal, and the wake-up object can be voice-woken up based on the wake-up instruction.
[0035] As another possible implementation, characters that can achieve voice wake-up for the wake-up object can be obtained, and these characters can be determined as the wake-up words of the wake-up object. In this scenario, word detection of the wake-up words is performed on the target wake-up signal based on the word detection method in the related art.
[0036] Among them, when it is recognized from the detection result that the target wake-up signal carries a wake-up word, a voice wake-up instruction for the wake-up object can be obtained based on the wake-up word, thereby realizing the voice wake-up of the wake-up object.
[0037] The voice wake-up method proposed in this disclosure enhances the target sound signal to obtain a target wake-up signal, and then performs wake-up word detection on the target wake-up signal after enhancement based on the target sound signal, improving the accuracy of wake-up word detection, and further improving the accuracy and precision of voice wake-up of the wake-up object. Compared with the method of voice wake-up based on sound source localization in the related art, it reduces the computational complexity of voice wake-up of the wake-up object, improves the efficiency of voice wake-up of the wake-up object, and optimizes the voice wake-up method.
[0038] In the above embodiment, regarding the voice wake-up of the wake-up object, it can be combined with Figure 2 For further understanding, Figure 2 is a schematic flowchart of the voice wake-up method according to another embodiment of the present disclosure. As Figure 2 shown, the method includes:
[0039] S201, obtain the target signal-to-noise ratio corresponding to the target sound signal, and the target signal frame corresponding to the target signal-to-noise ratio.
[0040] Optionally, obtain the first candidate signal-to-noise ratio of the current first candidate signal frame in the target sound signal, and the second candidate signal-to-noise ratios of each previous second candidate signal frame of the first candidate signal frame in the target sound signal.
[0041] In the embodiments of the present disclosure, the signal frame currently calculating the signal-to-noise ratio of the target sound signal can be determined as the first candidate signal frame, and each signal frame at the previous position of the first candidate signal frame in the target sound signal can be determined as each second candidate signal frame. Among them, the signal-to-noise ratio of the first candidate signal frame and each second candidate signal frame can be calculated respectively based on a preset signal-to-noise ratio algorithm, and the signal-to-noise ratio corresponding to the first candidate signal frame is determined as the first candidate signal-to-noise ratio of the first candidate signal frame, and the signal-to-noise ratio of any second candidate signal frame obtained is determined as the second candidate signal-to-noise ratio of the second candidate signal frame.
[0042] Among them, regarding the acquisition of the first candidate signal-to-noise ratio and each second candidate signal-to-noise ratio, it can be understood in combination with the following content:
[0043] Optionally, obtain a third candidate signal frame in the target sound signal, where the third candidate signal frame is any signal frame among the first candidate signal frame and each second candidate signal frame.
[0044] In the embodiments of the present disclosure, any signal frame among the first candidate signal frame and each second candidate signal frame can be determined as the third candidate signal frame, and the signal-to-noise ratio of the third candidate signal frame can be calculated.
[0045] Optionally, obtain a first frame feature vector, a first speech power spectral density, and a first noise power spectral density of the third candidate signal frame.
[0046] In the embodiments of the present disclosure, the acquisition of the first frame feature vector of the third candidate signal frame can be understood in combination with the following content:
[0047] Optionally, a fourth candidate signal frame adjacent to the previous one of the third candidate signal frame in the target sound signal can be obtained, a second speech power spectral density and a second noise power spectral density of the fourth candidate signal frame, and a second frame feature vector corresponding to the fourth candidate signal frame can be obtained.
[0048] In the embodiments of the present disclosure, the signal frame adjacent to the previous one of the third candidate signal frame in the target sound signal can be determined as the fourth candidate signal frame, and the noise part and the speech part included in the fourth candidate signal frame can be respectively processed by an algorithm based on the power spectral density algorithm in the related art. Then, according to the results of the algorithm processing, the power spectral density corresponding to the noise part of the fourth candidate signal frame is determined as the second noise power spectral density, and the power spectral density corresponding to the speech part is determined as the second speech power spectral density.
[0049] In this scenario, the second speech power spectral density and the second noise power spectral density can be processed by an algorithm based on a preset feature vector algorithm, and then according to the results of the algorithm processing, the feature vector of the fourth candidate signal frame is obtained and determined as the second frame feature vector of the fourth candidate signal frame.
[0050] Optionally, determine a candidate frame feature vector of the third candidate signal frame according to the second noise power spectral density, the second speech power spectral density, and the second frame feature vector.
[0051] As an example, it can be understood based on the following algorithm:
[0052] F m =φ NN,M-1 (k) -1 φ XX,m-1 (k)F m-1
[0053] In the above formula, F mThe candidate frame feature vector representing the third candidate signal frame, m represents the number of frames of the third candidate signal frame, m - 1 represents the number of frames of the fourth candidate signal frame, φ NN,m-1 (k) represents the second noise power spectral density of the fourth candidate signal frame, φ XX,m-1 (k) represents the second speech power spectral density of the fourth candidate signal frame, F m-1 represents the second frame feature vector of the fourth signal frame, k represents the frequency point.
[0054] Optionally, normalize the candidate frame feature vector to obtain the first frame feature vector of the third candidate signal frame.
[0055] As an example, the candidate frame feature vector can be normalized based on the following algorithm, and the feature vector obtained after the normalization process is determined as the first frame feature vector of the third candidate signal frame.
[0056]
[0057] In the above formula, F m represents the candidate frame feature vector of the third candidate signal frame, m represents the number of frames of the third candidate signal frame, ||*|| represents the 2-norm.
[0058] In the embodiments of the present disclosure, regarding the acquisition of the first speech power spectral density and the first noise power spectral density of the third candidate signal frame, it can be understood in combination with the following content:
[0059] Optionally, determine the first speech power spectral density of the third candidate frame signal according to the first forgetting factor corresponding to the second speech power spectral density.
[0060] As an example, based on the algorithm shown in the following formula, based on the second speech power spectral density and the corresponding first forgetting factor, obtain the speech power spectral density of the third candidate frame signal as the first speech power spectral density. The algorithm is as follows:
[0061]
[0062] In the above formula, φ XX,m represents the first speech power spectral density, φ XX,m-1 represents the second speech power spectral density of the fourth candidate signal frame, β represents the first forgetting factor, X represents the speech part in the third candidate signal frame, H represents the conjugate transpose, and m represents the number of frames.
[0063] Optionally, determine the first noise power spectral density of the third candidate frame signal according to the second forgetting factor corresponding to the second noise power spectral density.
[0064] As an example, based on the algorithm shown in the following formula, the noise power spectral density of the third candidate frame signal can be obtained based on the second noise power spectral density and the corresponding second forgetting factor, and used as the first noise power spectral density. The algorithm is as follows:
[0065]
[0066] In the above formula, φ NN,m represents the first noise power spectral density, φ NN,m-1 represents the second noise power spectral density of the fourth candidate signal frame, α represents the second forgetting factor, N represents the noise part in the third candidate signal frame, H represents the conjugate transpose, and m represents the number of frames.
[0067] It should be noted that the first forgetting factor and the second forgetting factor proposed above have a preset value range, which can be [0, 1].
[0068] That is to say, for the acquisition of the speech power spectral density and the noise power spectral density, it can be obtained based on the algorithm formula proposed in the above embodiment, based on the noise power spectral density and the speech power spectral density of the previous signal frame corresponding to the signal frame to be calculated currently. In this scenario, the noise power spectral density and the speech power spectral density of the first signal frame in the target sound signal can be determined as preset values, and the noise power spectral density and the speech power spectral density of each subsequent signal frame can be calculated based on this presetting.
[0069] Optionally, according to the first frame feature vector, the first speech power spectral density, and the first noise power spectral density, the third candidate signal-to-noise ratio of the third candidate signal frame is obtained, where the third candidate signal-to-noise ratio is the signal-to-noise ratio corresponding to the third candidate signal frame among the first candidate signal-to-noise ratio and each second candidate signal-to-noise ratio.
[0070] As an example, based on the signal-to-noise ratio algorithm in the related art, algorithm processing can be performed on the first frame feature vector, the first speech power spectral density, and the first noise power spectral density, so as to obtain the signal-to-noise ratio of the third candidate signal frame, as the third candidate signal-to-noise ratio.
[0071] Optionally, the signal-to-noise ratio algorithm can be as follows:
[0072]
[0073] In the above formula, SNR represents the third candidate signal-to-noise ratio, F represents the first frame feature vector, φ NN (k) represents the first speech power spectral density, φ xX (k) represents the first noise power spectral density, H represents the conjugate transpose, and k represents the frequency point.
[0074] Optionally, a target signal-to-noise ratio is obtained based on the first candidate signal-to-noise ratio and each second candidate signal-to-noise ratio.
[0075] Wherein, in response to the first candidate signal-to-noise ratio being greater than each second candidate signal-to-noise ratio, the first candidate signal-to-noise ratio is determined as the target signal-to-noise ratio.
[0076] In the embodiments of the present disclosure, when the first candidate signal-to-noise ratio is greater than each second candidate signal-to-noise ratio, it can be determined that among the signal frames for which the signal ratio has been calculated currently, the first candidate signal-to-noise ratio is the maximum value among them. In this scenario, the first candidate signal-to-noise ratio can be determined as the target signal-to-noise ratio in the target sound signal.
[0077] Correspondingly, in response to the third candidate signal-to-noise ratio, which is the maximum among each second candidate signal-to-noise ratio, being greater than or equal to the first candidate signal-to-noise ratio, the third candidate signal-to-noise ratio is determined as the target signal-to-noise ratio.
[0078] In the embodiments of the present disclosure, for each second candidate signal-to-noise ratio, the maximum value among them can be determined as the third candidate signal-to-noise ratio among each second candidate signal-to-noise ratio. In this scenario, when the third candidate signal-to-noise ratio is greater than or equal to the first candidate signal-to-noise ratio, it can be determined that among the signal frames for which the signal ratio has been calculated currently, both the third candidate signal-to-noise ratio and the first candidate signal-to-noise ratio are the maximum values among them. In this scenario, the third candidate signal-to-noise ratio can be determined as the target signal-to-noise ratio in the target sound signal.
[0079] That is to say, for the input target sound signal, after obtaining the signal-to-noise ratio of each signal frame, the maximum signal-to-noise ratio can be selected therefrom as the target signal-to-noise ratio in the target sound signal. In the scenario of calculating the signal-to-noise ratio of each signal frame in real time, the target signal-to-noise ratio with the maximum current value of the target sound signal can be obtained based on the algorithm proposed in the above embodiments.
[0080] S202. Enhance the target sound signal according to the target signal-to-noise ratio and the corresponding target signal frame to obtain a target wake-up signal.
[0081] Optionally, obtain the target filter coefficient and the target frequency-domain signal corresponding to the target signal frame, and obtain the target wake-up signal based on the target filter coefficient and the target frequency-domain signal.
[0082] In the embodiments of the present disclosure, the signal frame corresponding to the target signal-to-noise ratio in the target sound signal can be determined as the target signal frame in the target sound signal. In this scenario, the filter coefficient corresponding to the target signal frame can be obtained as the target filter coefficient based on the method for obtaining the filter coefficient in the related art, and the frequency-domain signal corresponding to the target signal frame can be obtained as the target frequency-domain signal based on the method for obtaining the frequency-domain signal in the related art.
[0083] In this scenario, based on a preset signal generation algorithm, algorithm processing can be performed on the target filter coefficients and the target frequency-domain signal of the target signal frame, and the signal obtained based on the algorithm is determined as the target wake-up signal obtained after enhancing the target voice signal.
[0084] Among them, the signal generation algorithm can be as follows:
[0085] Y(k) = F(k) H S(k)
[0086] In the above formula, Y(k) represents the target wake-up signal, F(k) H represents the target filter coefficients of the target signal frame, and S(k) represents the target frequency-domain signal of the target signal frame.
[0087] In the embodiments of the present disclosure, the acquisition of the target filter coefficients and the target frequency-domain signal corresponding to the target signal frame can be understood in combination with the following content:
[0088] Optionally, the target voice signal is obtained according to the candidate voice signals input in at least two audio input channels configured for the wake-up object.
[0089] It should be noted that the target voice signal is obtained based on the candidate voice signals respectively input through at least two audio input channels configured for the wake-up object.
[0090] As an example, as Figure 3 shown, in the Figure 3 scenario shown, Microphone 1 and Microphone 2 can be used as two audio input channels configured for the wake-up object. In the Figure 3 shown, the Figure 3 voice signal emitted by the sound source shown can be input through the two audio input channels provided by Microphone 1 and Microphone 2.
[0091] In this scenario, the target voice signals input through each audio input channel in at least two audio input channels can be determined as the candidate voice signals for each channel. Further, the target voice signal is obtained based on each candidate voice signal.
[0092] Among them, the candidate voice signals can be integrated, and the integrated signal can be used as the target voice signal for signal enhancement processing, which is not specifically limited here.
[0093] Optionally, the candidate frequency-domain signals of the target signal frame in each candidate voice signal are obtained, and the target frequency-domain signal of the target signal frame is obtained based on each candidate frequency-domain signal.
[0094] In the embodiments of the present disclosure, based on the frequency-domain signal acquisition algorithm in the related art, algorithm processing can be performed on each candidate sound signal, and then the frequency-domain signal of each candidate sound signal can be obtained based on the result of the algorithm processing as the candidate frequency-domain signal of each candidate sound signal.
[0095] In this scenario, algorithm processing can be performed on each candidate frequency-domain signal based on a preset integration algorithm, and then the target frequency-domain signal of the target signal frame can be obtained.
[0096] As an example, in Figure 3 the scenario shown, the target frequency-domain signal of the target signal frame can be obtained based on the following algorithm for each candidate sound signal:
[0097] S(k) = (S1(k), S2(k)) T
[0098] In the above formula, S(k) represents the target frequency-domain signal, S1(k) represents Figure 3 the candidate frequency-domain signal of the candidate sound signal input through the microphone 1 channel shown, and S2(k) represents Figure 3 the candidate frequency-domain signal of the candidate sound signal input through the microphone 2 channel shown, and T represents transpose.
[0099] Optionally, obtain the candidate filter coefficients of the target signal frame in each candidate sound signal, and based on each candidate filter coefficient, obtain the target filter coefficient of the target signal frame.
[0100] In the embodiments of the present disclosure, based on the method for obtaining filter coefficients in the related art, algorithm processing can be performed on each candidate sound signal to obtain the filter coefficients of each candidate sound signal as the candidate filter coefficients of each candidate sound signal. Further, algorithm processing is performed on each candidate filter coefficient based on a preset algorithm, and the filter coefficient obtained based on the algorithm processing is determined as the target filter coefficient of the target signal frame.
[0101] As an example, in Figure 3 the scenario shown, the acquisition algorithm of the target filter coefficient can be as follows:
[0102] F(k) = (F1(k), F2(k)) T
[0103] In the above embodiment, F(k) represents the target filter coefficient, F1(k) represents Figure 3 the candidate filter coefficient of the candidate sound signal input through the microphone 1 channel shown, and F2(k) represents Figure 3 the candidate filter coefficient of the candidate sound signal input through the microphone 2 channel shown, and T represents transpose.
[0104] S203. Perform wake word detection on the candidate voice signals input based on at least two audio input channels and the target wake-up signal.
[0105] In the embodiments of the present disclosure, wake word detection can be performed on the candidate voice signals input based on at least two audio input channels, and wake word detection can also be performed on the target wake-up signal obtained after enhancing the target voice signal, so as to obtain corresponding wake word detection results.
[0106] As an example, as Figure 3 shown, in the Figure 3 scenario shown, wake word detection can be performed on the candidate voice signals input through the two audio input channels of the Figure 3 shown microphone 1 and microphone 2, and keyword detection can be performed on the enhanced target wake-up signal obtained by the signal enhancement module 51.
[0107] It can be understood that wake word detection is performed on the candidate voice signals input through each of at least two audio input channels and the enhanced target wake-up signal, so as to achieve the purpose of improving the accuracy of wake word detection.
[0108] S204. Based on the detection results, perform voice wake-up on the wake-up object.
[0109] Optionally, perform wake word detection on the target wake-up signal and the candidate voice signals input through at least two audio input channels through a voice wake-up model.
[0110] As an example, as Figure 3 shown, in the Figure 3 scenario shown, the target voice signal can be enhanced by the Figure 3 shown signal enhancement module 51 to obtain an enhanced target wake-up signal, and the target wake-up signal is transmitted to the Figure 3 shown voice wake-up model, and keyword detection is performed on the target wake-up signal through the Figure 3 model capabilities of the shown voice wake-up model.
[0111] In addition, the candidate voice signals input through the two audio input channels of microphone 1 and microphone 2 are input into the Figure 3 shown voice wake-up model, and keyword detection is performed on each candidate voice signal through the Figure 3 model capabilities of the shown voice wake-up model, and then corresponding wake word detection results are obtained.
[0112] Among them, in response to the detection results indicating that the voice wake-up model detects a wake word, the wake-up object is activated.
[0113] In the embodiments of the present disclosure, when the wake word detection result of the voice wake-up model detects a wake word, it can be understood that at least one of the target voice signal and the target wake signal obtained based on the candidate voice signals input through at least two audio input channels carries the wake word required by the wake-up object. In this scenario, it can be determined that the currently detected target wake signal and the target voice signal obtained based on the candidate voice signals input through at least two audio input channels meet the wake-up conditions preset by the wake-up object. In this scenario, a wake-up instruction can be sent to the wake-up object based on the preset voice wake-up process, and the wake-up object can be started based on the wake-up instruction, so as to achieve the purpose of waking up the wake-up object for voice wake-up.
[0114] Correspondingly, in response to the detection result indicating that the voice wake-up model does not detect a wake word, keep the wake-up object closed.
[0115] In the embodiments of the present disclosure, when the wake word detection result of the voice wake-up model does not detect a wake word, it can be understood that neither the target voice signal obtained based on the candidate voice signals input through at least two audio input channels nor the target wake signal carries the wake word required by the wake-up object. In this scenario, it can be determined that the currently detected target wake signal and the target voice signal obtained based on the candidate voice signals input through at least two audio input channels do not meet the wake-up conditions preset by the wake-up object, and it can be determined that the input target voice signal may be a noise signal or a non-wake signal. In this scenario, keep the wake-up object closed.
[0116] It should be noted that there may be only noise signals in the input target voice signal, or there may be both noise signals and voice signals at the same time.
[0117] As an example, as Figure 4 shown, the component detection module shown in Figure 4 can perform component detection on the candidate voice signals input through the two audio input channels of the microphone 1 and the microphone 2 shown in Figure 4 and obtain the component detection result of the input target voice signal based on the component detection results of the respective candidate voice signals.
[0118] Among them, component detection can be understood as detecting the noise part and the voice part included in each signal frame of the input target voice signal.
[0119] It should be noted that Figure 4 the component detection module shown in
[0120] Optionally, in response to recognizing that only a noise signal exists in the target voice signal, update the current noise power spectral density of the target voice signal.
[0121] Optionally, in response to recognizing that a voice signal and a noise signal exist in the target voice signal, update the corresponding voice power spectral density and noise power spectral density of the target voice signal.
[0122] It can be understood that when it is detected that only the noise signal part exists in the current signal frame of the target voice signal, the currently calculated noise power spectral density can be updated, and there is no need to update the voice power spectral density.
[0123] Correspondingly, when it is detected that both the voice signal part and the noise signal part exist in the current signal frame of the target voice signal, the current noise power spectral density and voice power spectral density can be updated respectively based on the algorithms of the noise power spectral density and voice power spectral density proposed in the above embodiments.
[0124] That is to say, based on this detection step, the purpose of reducing the computational amount for updating the noise power spectral density and voice power spectral density can be achieved.
[0125] The voice wake-up method proposed by the present disclosure enhances the target voice signal to obtain a target wake-up signal, and then performs wake-word detection based on the target voice signal and the enhanced target wake-up signal, improving the accuracy of wake-word detection, and further improving the accuracy and precision of voice wake-up of the wake-up object. It detects the noise signal component and voice signal component in the target voice signal, and determines the update strategy of the noise power spectral density and voice power spectral density according to the detection result, reducing the computational amount of updating the noise power spectral density and voice power spectral density, thereby reducing the computational amount of voice wake-up of the wake-up object and improving the efficiency of voice wake-up of the wake-up object.
[0126] Corresponding to the voice wake-up methods proposed in the above several embodiments, an embodiment of the present disclosure also proposes a voice wake-up device. Since the voice wake-up device proposed in the embodiments of the present disclosure corresponds to the voice wake-up methods proposed in the above several embodiments, the implementation manners of the above voice wake-up methods are also applicable to the voice wake-up device proposed in the embodiments of the present disclosure and will not be described in detail in the following embodiments.
[0127] Figure 5 It is a schematic structural diagram of a voice wake-up device according to an embodiment of the present disclosure. As Figure 5 shown, the voice wake-up device 500 includes an enhancement module 51 and a detection and wake-up module 52, where:
[0128] An enhancement module 51, configured to enhance an input target voice signal in response to the input of the target voice signal, so as to obtain an enhanced target wake-up signal;
[0129] A detection and wake-up module 52, configured to perform wake-up word detection on the target wake-up signal, so as to perform voice wake-up on a wake-up object.
[0130] In an embodiment of the present disclosure, the target voice signal is obtained based on candidate voice signals respectively input through at least two audio input channels configured by the wake-up object. The detection and wake-up module is further configured to include: performing wake-up word detection on the candidate voice signals input through the at least two audio input channels and the target wake-up signal; and performing voice wake-up on the wake-up object based on the detection result.
[0131] In an embodiment of the present disclosure, the detection and wake-up module 52 is further configured to: perform wake-up word detection on the target wake-up signal and the candidate voice signals input through the at least two audio input channels through a voice wake-up model; and start the wake-up object in response to the detection result indicating that the voice wake-up model detects a wake-up word.
[0132] In an embodiment of the present disclosure, the detection and wake-up module 52 is further configured to: keep the wake-up object closed in response to the detection result indicating that the voice wake-up model does not detect a wake-up word.
[0133] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a target signal-to-noise ratio corresponding to the target voice signal and a target signal frame corresponding to the target signal-to-noise ratio; and enhance the target voice signal according to the target signal-to-noise ratio and the corresponding target signal frame, so as to obtain the target wake-up signal.
[0134] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a first candidate signal-to-noise ratio of a current first candidate signal frame in the target voice signal and second candidate signal-to-noise ratios of each previous second candidate signal frame of the first candidate signal frame in the target voice signal; and obtain the target signal-to-noise ratio based on the first candidate signal-to-noise ratio and the second candidate signal-to-noise ratios.
[0135] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a third candidate signal frame in the target voice signal, where the third candidate signal frame is any one of the first candidate signal frame and each second candidate signal frame; obtain a first frame feature vector, a first voice power spectral density, and a first noise power spectral density of the third candidate signal frame; and obtain a third candidate signal-to-noise ratio of the third candidate signal frame according to the first frame feature vector, the first voice power spectral density, and the first noise power spectral density, where the third candidate signal-to-noise ratio is the signal-to-noise ratio corresponding to the third candidate signal frame among the first candidate signal-to-noise ratio and the second candidate signal-to-noise ratios.
[0136] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a fourth candidate signal frame adjacent to a third candidate signal frame in a target sound signal; obtain a second speech power spectral density, a second noise power spectral density, and a second frame feature vector corresponding to the fourth candidate signal frame of the fourth candidate signal frame; determine a candidate frame feature vector of the third candidate signal frame according to the second noise power spectral density, the second speech power spectral density, and the second frame feature vector; and normalize the candidate frame feature vector to obtain a first frame feature vector of the third candidate signal frame.
[0137] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: determine a first speech power spectral density of a third candidate frame signal according to a first forgetting factor corresponding to the second speech power spectral density; and determine a first noise power spectral density of the third candidate frame signal according to a second forgetting factor corresponding to the second noise power spectral density.
[0138] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: determine a first candidate signal-to-noise ratio as a target signal-to-noise ratio in response to the first candidate signal-to-noise ratio being greater than each second candidate signal-to-noise ratio; and determine a third candidate signal-to-noise ratio as the target signal-to-noise ratio in response to the largest third candidate signal-to-noise ratio among the second candidate signal-to-noise ratios being greater than or equal to the first candidate signal-to-noise ratio.
[0139] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a target filter coefficient and a target frequency-domain signal corresponding to a target signal frame; and obtain a target wake-up signal based on the target filter coefficient and the target frequency-domain signal.
[0140] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: obtain a target sound signal according to candidate sound signals input through at least two audio input channels configured for a wake-up object; obtain candidate frequency-domain signals of the target signal frame in each candidate sound signal, and obtain a target frequency-domain signal of the target signal frame based on the candidate frequency-domain signals; obtain candidate filter coefficients of the target signal frame in each candidate sound signal, and obtain a target filter coefficient of the target signal frame based on the candidate filter coefficients.
[0141] In an embodiment of the present disclosure, the enhancement module 51 is further configured to: update a current noise power spectral density of the target sound signal in response to identifying that only a noise signal exists in the target sound signal; and update a speech power spectral density and a noise power spectral density corresponding to the target sound signal in response to identifying that both a speech signal and a noise signal exist in the target sound signal.
[0142] The voice wake-up device proposed by the present disclosure enhances the target sound signal to obtain a target wake-up signal, and then performs wake-word detection based on the enhanced target wake-up signal of the target sound signal, improving the accuracy of wake-word detection, thereby improving the accuracy and precision of voice wake-up of the wake-up object. Compared with the method of voice wake-up based on sound source localization in the related art, the computational load of voice wake-up of the wake-up object is reduced, the efficiency of voice wake-up of the wake-up object is improved, and the method of voice wake-up is optimized.
[0143] To achieve the above embodiments, the present disclosure also provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive signals and send signals to the processor, and the signals include computer instructions stored in the memory. When the processor executes the computer instructions, the chip executes the steps of the voice wake-up method and / or the touch point prediction method provided in the above embodiments.
[0144] To achieve the above embodiments, the present disclosure also provides an electronic device, a computer-readable storage medium, and a computer program product.
[0145] Figure 6 For the block diagram of the electronic device 600 according to an embodiment of the present disclosure, as Figure 6 shown, the electronic device 600 includes a memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602. When the processor 602 executes the program instructions, the voice wake-up method provided in the above embodiments is implemented.
[0146] Optionally, the electronic device is a wearable device configured with dual microphones.
[0147] To implement the above embodiments, the present disclosure also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the program is executed by the processor, the voice wake-up method provided in the above embodiments is implemented.
[0148] To implement the above embodiments, the present disclosure also proposes a computer program product, on which a computer program is stored. When the computer program is executed by the processor, the voice wake-up method provided in the above embodiments is implemented.
[0149] To implement the above embodiments, the present disclosure also provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive signals and send the signals to the processor, and the signals include computer instructions stored in the memory. When the processor executes the computer instructions, the chip executes the steps of the voice wake-up method provided in the above embodiments.
[0150] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0151] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0152] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0153] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0154] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0155] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0156] In addition, the functional units in the various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0157] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0158] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.
[0159] The above specific implementation manners do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A voice wake-up method, characterized in that, The method includes: In response to the input of a target voice signal, enhancing the input target voice signal to obtain an enhanced target wake-up signal; Performing wake-word detection on the target wake-up signal to perform voice wake-up on a wake-up object.
2. The method according to claim 1, wherein The target voice signal is obtained based on candidate voice signals respectively input through at least two audio input channels configured by the wake-up object. The performing wake-word detection on the target wake-up signal to perform voice wake-up on a wake-up object includes: Performing wake-word detection on the candidate voice signals input through at least two audio input channels and the target wake-up signal; Based on the detection results, performing voice wake-up on the wake-up object.
3. The method according to claim 2, wherein The performing voice wake-up on the wake-up object based on the detection results includes: Performing wake-word detection on the target wake-up signal and the candidate voice signals input through at least two audio input channels through a voice wake-up model; In response to the detection results indicating that the voice wake-up model detects the wake word, starting the wake-up object.
4. The method according to claim 2, characterized in that The method further includes: In response to the detection results indicating that the voice wake-up model does not detect the wake word, keeping the wake-up object closed.
5. The method according to claim 1, wherein The enhancing the target voice signal to obtain an enhanced target wake-up signal includes: Obtaining a target signal-to-noise ratio corresponding to the target voice signal and a target signal frame corresponding to the target signal-to-noise ratio; Enhancing the target voice signal according to the target signal-to-noise ratio and the corresponding target signal frame to obtain the target wake-up signal.
6. The method according to claim 5, characterized in that The obtaining a target signal-to-noise ratio corresponding to the target voice signal and a target signal frame corresponding to the target signal-to-noise ratio includes: Obtaining a first candidate signal-to-noise ratio of a current first candidate signal frame in the target voice signal and second candidate signal-to-noise ratios of previous second candidate signal frames of the first candidate signal frame in the target voice signal; Based on the first candidate signal-to-noise ratio and the second candidate signal-to-noise ratios, obtaining the target signal-to-noise ratio.
7. The method according to claim 6, wherein The obtaining a first candidate signal-to-noise ratio of a current first candidate signal frame in the target voice signal and second candidate signal-to-noise ratios of previous second candidate signal frames of the first candidate signal frame in the target voice signal includes: Obtaining a third candidate signal frame in the target voice signal, where the third candidate signal frame is any one of the first candidate signal frame and the second candidate signal frames; Obtaining a first frame feature vector, a first speech power spectral density, and a first noise power spectral density of the third candidate signal frame; According to the first frame feature vector, the first speech power spectral density, and the first noise power spectral density, obtaining a third candidate signal-to-noise ratio of the third candidate signal frame, where the third candidate signal-to-noise ratio is the signal-to-noise ratio corresponding to the third candidate signal frame among the first candidate signal-to-noise ratio and the second candidate signal-to-noise ratios.
8. The method according to claim 7, wherein The obtaining a first frame feature vector of the third candidate signal frame includes: Obtaining a previous adjacent fourth candidate signal frame of the third candidate signal frame in the target voice signal; Obtain the second speech power spectral density and the second noise power spectral density of the fourth candidate signal frame, as well as the second frame feature vector corresponding to the fourth candidate signal frame; Determine the candidate frame feature vector of the third candidate signal frame according to the second noise power spectral density, the second speech power spectral density and the second frame feature vector; Normalize the candidate frame feature vector to obtain the first frame feature vector of the third candidate signal frame.
9. The method according to claim 8, characterized in that The obtaining the first speech power spectral density and the first noise power spectral density of the third candidate signal frame includes: Determine the first speech power spectral density of the third candidate frame signal according to the first forgetting factor corresponding to the second speech power spectral density; Determine the first noise power spectral density of the third candidate frame signal according to the second forgetting factor corresponding to the second noise power spectral density.
10. The method according to claim 6, wherein The obtaining the target signal-to-noise ratio based on the first candidate signal-to-noise ratio and each second candidate signal-to-noise ratio includes: In response to the first candidate signal-to-noise ratio being greater than each second candidate signal-to-noise ratio, determine the first candidate signal-to-noise ratio as the target signal-to-noise ratio; In response to the maximum third candidate signal-to-noise ratio among the second candidate signal-to-noise ratios being greater than or equal to the first candidate signal-to-noise ratio, determine the third candidate signal-to-noise ratio as the target signal-to-noise ratio.
11. The method according to claim 5, characterized in that The enhancing the target sound signal according to the target signal-to-noise ratio and the corresponding target signal frame to obtain the target wake-up signal includes: Obtain the target filter coefficients and the target frequency-domain signal corresponding to the target signal frame; Based on the target filter coefficients and the target frequency-domain signal, obtain the target wake-up signal.
12. The method according to claim 11, wherein The obtaining the target filter coefficients and the target frequency-domain signal corresponding to the target signal frame includes: Obtain the target sound signal according to the candidate sound signals input in at least two audio input channels configured by the wake-up object; Obtain the candidate frequency-domain signals of the target signal frame in each candidate sound signal, and based on the candidate frequency-domain signals, obtain the target frequency-domain signal of the target signal frame; Obtain the candidate filter coefficients of the target signal frame in each candidate sound signal, and based on the candidate filter coefficients, obtain the target filter coefficients of the target signal frame.
13. The method according to any one of claims 1 to 12, characterized in that, The method further includes: In response to recognizing that only a noise signal exists in the target sound signal, update the current noise power spectral density of the target sound signal; In response to recognizing that a speech signal and the noise signal exist in the target sound signal, update the speech power spectral density and the noise power spectral density corresponding to the target sound signal.
14. A voice wake-up device, characterized in that, The apparatus includes: An enhancement module, configured to enhance the input target sound signal to obtain an enhanced target wake-up signal in response to the input of the target sound signal; A detection and wake-up module, configured to perform wake-up word detection on the input target sound signal and the target wake-up signal to perform voice wake-up on the wake-up object.
15. An electronic device, characterized in that, Includes: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute instructions to implement the method according to any one of claims 1-13.
16. The electronic device according to claim 15, characterized in that, The electronic device is a wearable device configured with dual microphones.
17. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method according to any one of claims 1-13.
18. A chip, characterized in that, Comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal and send the signal to the processor, the signal includes computer instructions stored in a memory, and when the processor executes the computer instructions, enables the chip to execute the steps of the method according to any one of claims 1-13.