Voice wake-up method, device, electronic device, and storage medium
By combining delay and real-time wake-up detection, the acoustic coding layer is shared, and the real-time and effect of voice wake-up are achieved, and the problems of long response time and high false wake-up rate in the prior art are solved, which improves the accuracy and response speed of voice wake-up.
Patent Information
- Application Number
- CN202111574805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The existing voice wake-up solution cannot take into account the real-time and wake-up effects of voice wake-up, resulting in too long response time or high false wake-up rate.
The method of combining delay wake-up detection and real-time wake-up detection is adopted. The delay wake-up detection model and the real-time wake-up detection model share the acoustic coding layer, respectively, and switch to real-time wake-up detection under pre-wake conditions to ensure the wake-up effect while shortening the response delay.
Without damaging the wake-up effect, the response time of voice wake-up is significantly shortened, the false wake-up rate is reduced, and the accuracy and real-timeness of voice wake-up are improved.
Smart Images

Figure CN114333794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a voice wake-up method, device, electronic device and storage medium. Background Art
[0002] Voice wake-up activates a device from sleep mode to active mode using voice. The more accurate and rapid the detection of wake-up keywords in the real-time voice stream, the better the wake-up effect, the shorter the response time, and the better the user experience.
[0003] Currently, in the real-time task of voice wake-up, a causal model must be selected as the acoustic model. If the selected causal model is a pure causal convolution superposition, there is no delay effect, that is, the response time of voice wake-up is short, but the possibility of false wake-up is high, and the voice wake-up effect is poor; correspondingly, if the selected causal model is a non-causal convolution superposition, the voice wake-up effect can be better guaranteed, but due to the introduction of delay, the response time of voice wake-up will be longer. In summary, the current voice wake-up solution cannot take into account both the real-time performance and wake-up effect of voice wake-up. Summary of the Invention
[0004] The present invention provides a voice wake-up method, device, electronic device and storage medium, which are used to solve the defect that the voice wake-up solution in the prior art cannot take into account both the real-time performance and the wake-up effect of voice wake-up.
[0005] The present invention provides a voice wake-up method, comprising:
[0006] Performing a delay wake-up detection on the real-time voice stream to obtain a delay wake-up detection result, wherein a voice frame corresponding to the delay wake-up detection result differs from a voice frame corresponding to the real-time voice stream by a preset number of frames;
[0007] If the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. The voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0008] According to a voice wake-up method provided by the present invention, performing real-time wake-up detection on the real-time voice stream and performing voice wake-up based on a real-time wake-up detection result obtained by the real-time wake-up detection includes:
[0009] Performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream respectively to obtain the delayed wake-up detection result and the real-time wake-up detection result, and performing voice wake-up based on the delayed wake-up detection result and the real-time wake-up detection result.
[0010] According to a voice wake-up method provided by the present invention, performing voice wake-up based on the delayed wake-up detection result and the real-time wake-up detection result includes:
[0011] If the delayed wake-up detection result and / or the real-time wake-up detection result is wake-up, voice wake-up is performed, and the delayed wake-up detection and the real-time wake-up detection are reset.
[0012] According to a voice wake-up method provided by the present invention, the method further includes performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream, obtaining the delayed wake-up detection result and the real-time wake-up detection result, and then further including:
[0013] If both the delayed wake-up detection result and the real-time wake-up detection result are no wake-up, recording the duration of the real-time wake-up detection result being no wake-up;
[0014] If the duration exceeds a preset time, the real-time wake-up detection is interrupted.
[0015] According to a voice wake-up method provided by the present invention, the recording of the duration of the real-time wake-up detection result being non-awakening further includes:
[0016] If the duration does not exceed the preset time, determining a delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream;
[0017] If the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, real-time wake-up detection is performed on the real-time voice stream until the delayed wake-up detection result and / or the real-time wake-up detection result is awakened.
[0018] According to a voice wake-up method provided by the present invention, performing delay wake-up detection on a real-time voice stream includes:
[0019] Based on the delay wake-up detection model, the real-time voice stream is subjected to delay wake-up detection;
[0020] The performing real-time wake-up detection on the real-time voice stream includes:
[0021] Performing real-time wake-up detection on the real-time voice stream based on a real-time wake-up detection model;
[0022] The delayed wakeup detection model and the real-time wakeup detection model share an acoustic coding layer, and the acoustic coding layer is used to acoustically encode an input real-time speech stream.
[0023] According to a voice wake-up method provided by the present invention, the delayed wake-up detection model is trained based on a sample voice stream and a delayed wake-up detection label, and the delayed wake-up detection label includes at least one of wake-up, non-wake-up and pre-wake-up.
[0024] The present invention also provides a voice wake-up device, comprising:
[0025] A delay wake-up detection unit is used to perform a delay wake-up detection on the real-time voice stream to obtain a delay wake-up detection result, wherein the voice frame corresponding to the delay wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames;
[0026] A real-time wake-up detection unit is configured to perform real-time wake-up detection on the real-time voice stream if the delayed wake-up detection result is pre-wake-up, and perform voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, where the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0027] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described voice wake-up methods are implemented.
[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described voice wake-up methods.
[0029] The voice wake-up method, device, electronic device and storage medium provided by the present invention perform delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result. When the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. This overcomes the defect of traditional solutions that cannot take into account both the real-time performance and the wake-up effect of voice wake-up. It can shorten the response delay without compromising the wake-up effect of voice wake-up, thereby achieving a balance between the wake-up effect and real-time performance of voice wake-up. Moreover, the pre-wake-up judgment is used to switch from delayed wake-up detection to real-time wake-up detection, so that the wake-up detection of the real-time voice stream can be smooth and orderly, thereby reducing the false wake-up rate of voice wake-up and ensuring the accuracy of voice wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 Schematic diagram of the flow of the voice wake-up method provided by the present invention;
[0032] Figure 2 This is a flow chart of delayed wake-up detection and real-time wake-up detection provided by the present invention;
[0033] Figure 3 It is a structural diagram of the delayed wake-up detection model and the real-time wake-up detection model provided by the present invention;
[0034] Figure 4 This is an overall flow chart of the voice wake-up method provided by the present invention;
[0035] Figure 5 This is a structural diagram of the traditional voice wake-up solution provided by the present invention;
[0036] Figure 6 This is a flow chart of the traditional voice wake-up solution provided by the present invention;
[0037] Figure 7 This is an example diagram of the voice wake-up method provided by the present invention;
[0038] Figure 8 It is a structural diagram of the voice wake-up device provided by the present invention;
[0039] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0041] Voice wake-up activates a device from sleep mode to active mode using voice. Current voice wake-up tasks use the following metrics to measure the general performance of the voice wake-up system: wake-up rate, false wake-up rate, response time, and power consumption. Therefore, when performing voice wake-up, the more accurate and rapid the detection of wake-up keywords in the real-time voice stream, the better the wake-up effect, the shorter the response time, and the better the user experience.
[0042] Currently, in real-time voice wake-up tasks, a causal model must be selected as the acoustic model. If the selected causal model is a pure causal convolution superposition, there is no delay effect, that is, the response time of voice wake-up is short. However, the possibility of false wake-up is high, and the voice wake-up effect is poor. Correspondingly, if the selected causal model is a non-causal convolution superposition, the wake-up effect of voice wake-up can be better guaranteed. However, due to the introduction of delay, the response time of voice wake-up will be longer.
[0043] The current voice wake-up response time is the time difference from fully speaking the wake-up keyword to wake-up (the device gives feedback). As users' pursuit of the wake-up effect and response time of voice wake-up continues to increase, the current neural network model that performs voice wake-up tasks has a long response time and is therefore unable to meet users' actual needs. It cannot meet users' dual requirements for wake-up effect and response time.
[0044] In the voice wake-up task, the voice wake-up engine not only needs to correctly detect each word in the wake-up keyword, but also needs to determine the arrangement relationship between words, that is, the collocation between each word. For example, for the wake-up keyword "hello Xiao Ao", when performing wake-up detection, the voice wake-up engine not only needs to correctly detect the four words "ni", "hao", "xiao", and "ao", but also needs to determine the arrangement relationship between these four words, that is, the four words must form "hello Xiao Ao", not "hao ni ao xiao", "hao ni ao xiao", "hao ni Xiao Ao", etc.
[0045] When detecting "Hello Xiao Ao", the future information "Hello Xiao Ao" is used as a reference for "you". To obtain this future information, the traditional model uses non-causal convolution and introduces a fixed window delay. This delay allows the voice wake-up engine to obtain the future information "Hello Xiao Ao", thereby ensuring the voice wake-up effect. Due to the introduction of a fixed window delay, the wake-up detection results obtained by this model must also contain a fixed window delay. As a result, the voice wake-up response time also increases by the delay introduced by the model.
[0046] Accordingly, when performing wake-up detection, if "Hello Xiao Ao" is split, for "you", the future information "Hello Xiao Ao" is required as a reference; and for "Hello", only the future information "Xiao Ao" is required as a reference; further, for "Hello Xiao", only the future information "Ao" is required as a reference, and "Ao" can be detected almost without the need for future information as a reference. It can be seen that as the real-time voice stream advances, the field of view required for word-by-word detection of wake-up keywords is gradually narrowed.
[0047] Based on this, the present invention provides a voice wake-up method, which aims to ensure the wake-up effect while shortening the response delay of the voice wake-up. Figure 1 : is a flow chart of the voice wake-up method provided by the present invention, such as Figure 1 As shown, the method includes:
[0048] Step 110 , performing a delayed wakeup detection on the real-time voice stream to obtain a delayed wakeup detection result. The voice frame corresponding to the delayed wakeup detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames.
[0049] Since the voice wake-up task is to activate the device from the sleep state to the running state through voice, before performing voice wake-up, wake-up detection must first be performed on the real-time voice stream. Considering that the amount of information contained in a single voice frame in the real-time voice stream wake-up detection is small, in order to improve the effectiveness of wake-up detection, it is necessary to assist in the wake-up detection with the information contained in multiple voice frames after the current voice frame. Therefore, when performing wake-up detection on the current voice frame in the real-time voice stream, a delay can be set. That is, delayed wake-up detection is performed on the real-time voice stream to obtain a delayed wake-up detection result. When performing wake-up detection on the current voice frame, in addition to relying on the information contained in the current voice frame, information contained in a preset number of voice frames after the current voice frame can also be referenced. This reduces the false wake-up rate of voice wake-up and ensures the accuracy of voice wake-up.
[0050] Specifically, the process of performing delay wake-up detection on real-time voice stream can be implemented through a delay wake-up detection model. The specific process can be that the real-time voice stream is input into the delay wake-up detection model, and the delay wake-up detection model performs delay wake-up detection on the input real-time voice stream, and finally obtains the delay wake-up detection result output by the delay wake-up detection model.
[0051] It should be noted that, since a delay is set for the input and output of the delay wake-up model, when the delay wake-up detection model performs wake-up detection on the current voice frame in the real-time voice stream, it outputs the delay wake-up detection result of the preset number of voice frames before the current voice frame. It can also be understood that the voice frame corresponding to the delay wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames. For example, when the preset number of frames is 3 frames, and the voice frame corresponding to the real-time voice stream at the current moment is the 4th voice frame, the voice frame corresponding to the delay wake-up detection result output by the delay wake-up detection model is the 1st voice frame. The preset number of frames here is a pre-set frame delay, which can be set according to actual needs. For example, it can be 2 frames, 3 frames, 5 frames, etc.
[0052] Before inputting the real-time voice stream into the delayed wake-up detection model, the delayed wake-up detection model can be pre-trained based on the sample voice stream and the delayed wake-up detection label. The training process of the delayed wake-up detection model includes the following steps: First, a large number of sample voice streams are collected and the delayed wake-up detection labels of the sample voice streams are determined. The delayed wake-up detection labels here include one or more of wake-up, non-wake-up and pre-wake-up; then, based on the sample voice stream and the delayed wake-up detection label, the initial delayed wake-up detection model is trained to obtain a trained delayed wake-up detection model. It should be noted that the initial delayed wake-up detection model here can be constructed on the basis of a traditional delay model.
[0053] Step 120: If the delayed wake-up detection result is pre-wake-up, a real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. The voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0054] Considering that in step 110, the delayed wake-up detection result obtained by performing delayed wake-up detection on the real-time voice stream has already guaranteed the accuracy of voice wake-up to a great extent, based on this, in an embodiment of the present invention, after obtaining the delayed wake-up detection result, real-time wake-up detection can be performed on the real-time voice stream, that is, the delayed wake-up detection of the real-time voice stream is switched to real-time wake-up detection to reduce the response delay in the voice wake-up process, thereby achieving an overall advance in response time and ensuring the real-time nature of voice wake-up.
[0055] Specifically, after obtaining the delayed wake-up detection result in step 110, it is necessary to determine whether the delayed wake-up detection result is a pre-set pre-wake-up. The pre-wake-up here is a pre-set condition for switching from delayed wake-up detection to real-time wake-up detection. It can be set accordingly according to actual needs. For example, for the wake-up keyword "Hello Xiao Ao", "you" can be used as a pre-wake-up, "good" can be used as a pre-wake-up, and "Hello Xiao" can also be used as a pre-wake-up. The embodiment of the present invention does not make specific limitations on this.
[0056] Furthermore, if the delayed wake-up detection result is not pre-wake-up, it indicates that the wake-up detection of the real-time voice stream has not yet reached the pre-set switching condition from delayed wake-up detection to real-time wake-up detection. At this time, it is still necessary to continue the delayed wake-up detection until the delayed wake-up detection result is pre-wake-up.
[0057] Correspondingly, if the delayed wake-up detection result is pre-wake-up, it indicates that the wake-up detection for the real-time voice stream has met the pre-set switching conditions from delayed wake-up detection to real-time wake-up detection. At this time, real-time wake-up detection can be performed on the real-time voice stream to obtain a real-time wake-up detection result. It should be noted that real-time wake-up detection is wake-up detection without delay. It can also be understood that the voice frames corresponding to the real-time wake-up detection result obtained by performing real-time wake-up detection on the real-time voice stream are the same as the voice frames corresponding to the real-time voice stream.
[0058] Here, the process of performing real-time wake-up detection on the real-time voice stream can be implemented through a real-time wake-up detection model. The specific process can be: inputting the real-time voice stream into the real-time wake-up detection model, and the real-time wake-up detection model performs real-time wake-up detection on the input real-time voice stream, and finally obtains the real-time wake-up detection result output by the real-time wake-up detection model.
[0059] Before inputting the real-time voice stream into the real-time wakeup detection model, the real-time wakeup detection model can be pre-trained based on the sample voice stream and the real-time wakeup detection label. The training process of the real-time wakeup detection model includes the following steps: First, a large number of sample voice streams are collected and the real-time wakeup detection labels of the sample voice streams are determined. Here, the delayed wakeup detection labels include wakeup and / or non-wakeup; then, based on the sample voice streams and the real-time wakeup detection labels, the initial real-time wakeup detection model is trained to obtain the trained real-time wakeup detection model. It should be noted that the initial real-time wakeup detection model here can be built on the basis of a pre-trained multitask MLP (multitask multilayer perceptron) and decoder.
[0060] Thereafter, voice wake-up may be performed based on the real-time wake-up detection result obtained by the real-time wake-up detection.
[0061] It should be noted that, in the embodiment of the present invention, when real-time wake-up detection is performed on the real-time voice stream, the delayed wake-up detection of the real-time voice stream is not stopped, that is, after the delayed wake-up detection result is pre-wake-up, for the real-time voice stream, the original delayed wake-up detection is maintained and the real-time wake-up detection is added, that is, the real-time wake-up detection and delayed wake-up detection of the real-time voice stream are parallel, not only can the real-time wake-up detection result be obtained, but also the original delayed wake-up detection result is retained. The wake-up detection result obtained by the dual wake-up detection is used for voice wake-up, which can shorten the response delay while ensuring the wake-up effect of the voice wake-up.
[0062] The voice wake-up method provided by the present invention performs delayed wake-up detection on a real-time voice stream to obtain a delayed wake-up detection result. When the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. This overcomes the defect in traditional solutions that both the real-time performance and the wake-up effect of voice wake-up cannot be taken into account. Without compromising the wake-up effect of voice wake-up, the response delay can be shortened, thereby achieving a balance between the wake-up effect and the real-time performance of voice wake-up. Moreover, the pre-wake-up judgment is used to switch from delayed wake-up detection to real-time wake-up detection, so that the wake-up detection of the real-time voice stream can be smooth and orderly, thereby reducing the false wake-up rate of voice wake-up and ensuring the accuracy of voice wake-up.
[0063] Based on the above embodiment, in step 120, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection, including:
[0064] Performing delayed wakeup detection and real-time wakeup detection on the real-time voice stream respectively to obtain delayed wakeup detection results and real-time wakeup detection results, and performing voice wakeup based on the delayed wakeup detection results and the real-time wakeup detection results.
[0065] Considering that in step 120, voice wake-up is performed based on the real-time wake-up detection result obtained by real-time wake-up detection to shorten the response delay of voice wake-up, since the advance of the response time is based on the removal of delay, and removing the delay will have a very small impact on the accuracy of voice wake-up, based on this, in an embodiment of the present invention, when real-time wake-up detection is performed on the real-time voice stream, the original delayed wake-up detection can be retained. Even if the real-time wake-up detection and the delayed wake-up detection are performed in parallel, and voice wake-up is performed based on the wake-up detection result obtained by the dual wake-up detection, the response delay can be shortened as much as possible without damaging the wake-up effect of the voice wake-up.
[0066] Specifically, in step 120, when performing real-time wake-up detection on the real-time voice stream, the previous delayed wake-up detection can also be continued. This process can be specifically to perform real-time wake-up detection and delayed wake-up detection on the real-time voice stream respectively to obtain real-time wake-up detection results and delayed wake-up detection results. It should be noted that when performing real-time wake-up detection on the real-time voice stream, the various parameters obtained by the delayed wake-up detection can also be assigned to the corresponding parameters in the real-time wake-up detection, so that real-time wake-up detection can be performed on the basis of the delayed wake-up detection, thereby ensuring the continuity and accuracy of voice wake-up.
[0067] It should be noted that the real-time wake-up detection and delayed wake-up detection here can be implemented by the real-time wake-up detection model and the delayed wake-up detection model respectively. The process of performing real-time wake-up detection and delayed wake-up detection by the two models has been described in detail above and will not be repeated here.
[0068] After determining the real-time wake-up detection results and the delayed wake-up detection results, voice wake-up can be performed based on these two. If either of the wake-up detection results is a wake-up, it is considered a successful wake-up, and the device gives feedback to inform the outside world that the wake-up has been successful. Correspondingly, if both the real-time wake-up detection results and the delayed wake-up detection results are no wake-up, it is considered no wake-up. Voice wake-up is performed through dual detection results, which can shorten the response delay while ensuring the wake-up effect of voice wake-up.
[0069] It should be noted that the embodiment of the present invention achieves a reduction in the average response time, that is, the response delay can be shortened only when the real-time wake-up detection result is awakening; when the real-time wake-up detection result is not awakening and the delayed wake-up detection result is awakening, the response delay cannot be reduced.
[0070] Based on the above embodiment, voice wake-up is performed based on the delayed wake-up detection result and the real-time wake-up detection result, including:
[0071] If the delayed wake-up detection result and / or the real-time wake-up detection result is wake-up, voice wake-up is performed and the delayed wake-up detection and real-time wake-up detection are reset.
[0072] Specifically, in step 120, when performing voice wake-up according to the delayed wake-up detection result and the real-time wake-up detection result, it is first necessary to determine the wake-up state indicated by the delayed wake-up detection result and the real-time wake-up detection result. The wake-up state here includes wake-up and not wake-up.
[0073] If either the delayed wake-up detection result or the real-time wake-up detection result is wake-up, or if both of the delayed wake-up detection results are wake-up, voice wake-up is performed.
[0074] After that, the delayed wake-up detection and real-time wake-up detection of the real-time voice stream need to be reset, that is, the voice wake-up engine needs to be restored to the initial state in order to perform the next round of wake-up detection.
[0075] It should be noted that the initial state of the voice wake-up engine here is to perform only delayed wake-up detection on the real-time voice stream.
[0076] Based on the above embodiment, delayed wake-up detection and real-time wake-up detection are performed on the real-time voice stream respectively to obtain delayed wake-up detection results and real-time wake-up detection results, and then further includes:
[0077] If both the delayed wake-up detection result and the real-time wake-up detection result are no wake-up, the duration of the real-time wake-up detection result being no wake-up is recorded;
[0078] If the duration exceeds the preset time, the real-time wake-up detection is interrupted.
[0079] Specifically, in step 120, after performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream respectively and obtaining the delayed wake-up detection result and the real-time wake-up detection result, before performing voice wake-up according to the delayed wake-up detection result and the real-time wake-up detection result, it is necessary to determine the wake-up state indicated by the delayed wake-up detection result and the real-time wake-up detection result. On the premise that either one or both of the delayed wake-up detection result and the real-time wake-up detection result are wake-up, voice wake-up is performed, and after wake-up, the delayed wake-up detection and the real-time wake-up detection of the real-time voice stream are reset.
[0080] Correspondingly, if both the delayed wake-up detection result and the real-time wake-up detection result are no wake-up, it indicates that the wake-up keyword has not been detected, or the wake-up keyword detection error occurs. In this case, a timeout judgment can be performed to determine whether to interrupt the real-time wake-up detection of the real-time voice stream based on the result of the timeout judgment. This process specifically includes the following steps:
[0081] First, if both the delayed wake-up detection result and the real-time wake-up detection result are negative, record the duration of the real-time wake-up detection result being negative. This duration is the duration of the real-time wake-up detection on the real-time voice stream.
[0082] It should be noted that the duration here is measured by the number of voice frames. Recording the real-time wake-up detection result as the duration of not waking up actually records the number of decoded frames for decoding the voice frames in the real-time voice stream.
[0083] Then, it is determined whether the duration exceeds a preset time. The preset time here is a preset time for determining whether the real-time wake-up detection of the real-time voice stream has timed out. That is, under the condition that the real-time wake-up detection result is no wake-up, the real-time wake-up detection duration for the real-time voice stream can be tolerated. It can be set accordingly according to actual conditions. The preset time here is also set by the number of frames.
[0084] If the real-time wake-up detection result shows that the duration of non-wake-up exceeds the preset time, indicating that the real-time wake-up detection of the real-time voice stream has timed out, the real-time wake-up detection of the real-time voice stream will be interrupted and only the delayed wake-up detection will be performed. This can reduce the occupation of computing resources and ensure a smaller amount of calculation.
[0085] Correspondingly, if the real-time wake-up detection result is that the duration of non-awakening does not exceed the preset time, it indicates that the real-time wake-up detection of the real-time voice stream has not timed out. At this time, there is no need to interrupt the real-time wake-up detection of the real-time voice stream, and the delayed wake-up detection and real-time wake-up detection are still maintained in parallel.
[0086] Based on the above embodiment, the real-time wake-up detection result is recorded as the duration of not waking up, and then further includes:
[0087] If the duration does not exceed the preset time, determining the delay wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream;
[0088] If the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, a real-time wake-up detection is performed on the real-time voice stream until the delayed wake-up detection result and / or the real-time wake-up detection result is awakened.
[0089] Since the real-time wake-up detection result is without delay, for the same voice frame in the real-time voice stream, the time of obtaining its corresponding real-time wake-up detection result must be earlier than the time of obtaining the delayed wake-up detection result. That is to say, when the real-time wake-up detection result is obtained and the duration of non-wake-up in the real-time wake-up detection result does not exceed the preset time, the delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream can be determined, and voice wake-up can be performed based on this delayed wake-up detection result to ensure the wake-up effect of voice wake-up.
[0090] Specifically, if the real-time wake-up detection result shows that the duration of non-wake-up does not exceed the preset time, the delayed wake-up detection result of the current window is determined, that is, the delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream is determined.
[0091] Furthermore, it is necessary to determine the wake-up state indicated by the delayed wake-up detection result corresponding to the voice frame at the current moment, so as to perform voice wake-up or other operations based on this wake-up state. Specifically, this process can be: if the delayed wake-up detection result corresponding to the voice frame at the current moment is wake-up, voice wake-up is performed, and the delayed wake-up detection and real-time wake-up detection of the real-time voice stream are reset.
[0092] Correspondingly, if the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, the real-time wake-up detection is continued to be performed on the real-time voice stream to obtain the real-time wake-up detection result, and voice wake-up is performed based on the real-time wake-up detection result and the delayed wake-up detection result obtained by the ongoing delayed wake-up detection. The above process is repeated until either the delayed wake-up detection result or the real-time wake-up detection result is awakening, or both wake-up detection results are awakening. At this time, the wake-up result is thrown out, and the real-time wake-up detection and delayed wake-up detection of the real-time voice stream are reset, that is, the voice wake-up engine is restored to the initial state for the next round of wake-up detection.
[0093] In traditional solutions, a voice wake-up model is constructed by superimposing causal convolution and non-causal convolution, and voice wake-up is performed accordingly to ensure the wake-up effect of voice wake-up while controlling the delay. However, with the iterative update of technology and the improvement of users' requirements for real-time performance, the current voice wake-up solution can no longer meet users' expectations in terms of response time. In response to this situation, based on the above embodiment, the embodiment of the present invention provides a voice wake-up method, which aims to solve the problem of slow response time of voice wake-up during the process of iterative productization of the model. Figure 2 This is a flow chart of delayed wake-up detection and real-time wake-up detection provided by the present invention, such as Figure 2 As shown, in step 110, delay wake-up detection is performed on the real-time voice stream, including:
[0094] Step 111: Perform delay wakeup detection on the real-time voice stream based on the delay wakeup detection model;
[0095] In step 120, real-time wake-up detection is performed on the real-time voice stream, including:
[0096] Step 121: Perform real-time wake-up detection on the real-time voice stream based on the real-time wake-up detection model;
[0097] The delayed wakeup detection model and the real-time wakeup detection model share the acoustic coding layer, which is used to acoustically encode the input real-time speech stream.
[0098] Specifically, in step 110, the process of performing delay wake-up detection on the real-time voice stream can be implemented by a delay wake-up detection model. That is, in step 111, delay wake-up detection is performed on the real-time voice stream according to the delay wake-up detection model to obtain a delay wake-up detection result. The specific process includes the following steps:
[0099] First, the delayed wakeup detection model uses the MLP (Multilayer Perceptron) module to detect the delayed wakeup of the real-time speech stream and obtain the delayed wakeup posterior probability.
[0100] Then, according to the decoding module in the delayed awakening detection model, the delayed awakening posterior probability is decoded to obtain the delayed awakening detection result.
[0101] After obtaining the delayed wake-up detection result, and the delayed wake-up detection result is pre-wake-up, real-time wake-up detection can be performed on the real-time voice stream. In step 120, the process of performing real-time wake-up detection on the real-time voice stream can be implemented by a real-time wake-up detection model, that is, step 121, performing real-time wake-up detection on the real-time voice stream according to the real-time wake-up detection model to obtain a real-time wake-up detection result. The specific process includes the following steps:
[0102] First, the real-time wakeup detection is performed on the real-time speech stream through the MLP module in the real-time wakeup detection model to obtain the real-time wakeup posterior probability;
[0103] Then, according to the decoding module in the real-time wake-up detection model, the real-time wake-up posterior probability is decoded to obtain the real-time wake-up detection result.
[0104] It should be noted that the MLP modules in both the delayed wakeup detection model and the real-time wakeup detection model include an input layer, an acoustic coding layer, and a classification layer. The input layer receives a real-time speech stream, the acoustic coding layer acoustically encodes the input real-time speech stream, and the classification layer classifies the wakeup state indicated by the wakeup detection results.
[0105] Figure 3 Schematic diagram of the structure of the delayed wake-up detection model and the real-time wake-up detection model provided by the present invention, such as Figure 3As shown in the figure, the delayed wakeup detection model and the real-time wakeup detection model share the acoustic coding layer, that is, the parameters of the acoustic coding layer of the MLP module in the delayed wakeup detection model and the real-time wakeup detection model are the same. It can also be understood that a multitask MLP branch is implanted in the MLP module in the delayed wakeup detection model. This branch is inserted before the deconvolution layer and shares the acoustic coding layer with the MLP module in the delayed wakeup detection model. In other words, the structure of this branch is exactly the same as that of the MLP module in the delayed wakeup detection model, both of which are deconv and convout, and only the classification layer parameters are different.
[0106] In the embodiment of the present invention, the two models share parameters except for the classification layer. Such a setting can not only reduce the time and energy consumed in the model training process, but also reduce the memory usage.
[0107] In addition, the decoding modules in the delayed wakeup detection model and the real-time wakeup detection model are used to maintain two decoding instances, namely, posterior probabilities, which correspond to the wakeup state indicated by the delayed wakeup detection result (the first-level state, namely, the t state) and the wakeup state indicated by the real-time wakeup detection result (the second-level state, namely, the t+N state). The two decoding instances are independently decoded and input into the decoding modules in the delayed wakeup detection model and the real-time wakeup detection model, respectively, to obtain the delayed wakeup detection result and the real-time wakeup detection result.
[0108] Based on the above embodiment, the delayed wakeup detection model is trained based on sample voice streams and delayed wakeup detection labels, and the delayed wakeup detection labels include at least one of awakening, non-awakening, and pre-awakening.
[0109] Specifically, in step 111, before performing delayed wakeup detection on the real-time voice stream according to the delayed wakeup detection model, the delayed wakeup detection model may be pre-trained based on the sample voice stream and the delayed wakeup detection label. The training process of the delayed wakeup detection model includes the following steps:
[0110] First, a large number of sample voice streams are collected, and delayed wakeup detection tags of the sample voice streams are determined. The delayed wakeup detection tags here include one or more of wakeup, non-wakeup, and pre-wakeup. The pre-wakeup can be set accordingly according to actual needs. For example, for the wakeup keyword "hello Xiaoao", "you" can be used as a pre-wakeup, "good" can be used as a pre-wakeup, and "hello Xiao" can also be used as a pre-wakeup. This embodiment of the present invention does not specifically limit this.
[0111] Then, the initial delayed wakeup detection model is trained based on the sample voice stream and the delayed wakeup detection label, thereby obtaining a trained delayed wakeup detection model. It should be noted that the initial delayed wakeup detection model here can be constructed based on a traditional delay model.
[0112] Based on the above embodiments, Figure 4 This is the overall flow chart of the voice wake-up method provided by the present invention, such as Figure 4 As shown, the method includes:
[0113] Step 410: Perform delay wakeup detection on the real-time voice stream based on the delay wakeup detection model;
[0114] Step 420, determine whether the delayed wake-up detection result is pre-wake-up; if so, execute step 430; if not, execute step 410;
[0115] Step 430: Perform real-time wake-up detection on the real-time voice stream based on the real-time wake-up detection model;
[0116] Step 440: Determine whether the real-time wake-up detection result and the delayed wake-up detection result are wake-up; if the real-time wake-up detection result and / or the delayed wake-up detection result are wake-up, execute step 490; if the real-time wake-up detection result and the delayed wake-up detection result are both not wake-up, execute step 450;
[0117] Step 450: If both the real-time wake-up detection result and the delayed wake-up detection result are no wake-up, then the duration of the real-time wake-up detection result being no wake-up is recorded;
[0118] Step 460, determining whether the duration exceeds the preset time; if so, interrupting the real-time wake-up detection and returning to step 410; if not, executing step 470;
[0119] Step 470, determining a delayed wakeup detection result corresponding to the current voice frame in the real-time voice stream;
[0120] Step 480: Determine whether the delayed wakeup detection result corresponding to the current voice frame is wakeup; if so, execute step 490; if not, return to step 430 until the delayed wakeup detection result and / or the real-time wakeup detection result is wakeup;
[0121] Step 490: If the real-time wake-up detection result and / or the delayed wake-up detection result is wake-up, voice wake-up is performed, and the delayed wake-up detection and real-time wake-up detection are reset.
[0122] The method provided by the embodiment of the present invention performs delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result. When the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. This overcomes the defect of traditional solutions that cannot take into account both the real-time performance and the wake-up effect of voice wake-up. It can shorten the response delay without compromising the wake-up effect of voice wake-up, thereby achieving a balance between the wake-up effect and real-time performance of voice wake-up. Moreover, the pre-wake-up judgment is used to switch from delayed wake-up detection to real-time wake-up detection, so that the wake-up detection of the real-time voice stream can be smooth and orderly, thereby reducing the false wake-up rate of voice wake-up and ensuring the accuracy of voice wake-up.
[0123] The following uses the wake-up keyword "Hello Xiaoao" as an example to explain the voice wake-up process in detail:
[0124] Figure 5 This is a structural diagram of the traditional voice wake-up solution provided by the present invention, such as Figure 5 As shown in the figure, the traditional solution is to directly perform delay wake-up detection on the real-time voice stream and detect the wake-up keyword ("Hello Xiao Ao") in the real-time voice stream. The specific wake-up detection process includes the following steps: Figure 6 This is a flow chart of the traditional voice wake-up solution provided by the present invention. Figure 6 As shown in the figure: first, feature extraction is performed on the real-time voice stream; then, the features obtained by feature extraction are input into the acoustic model MLP for acoustic encoding, and the acoustic features output by the acoustic model MLP are input into the decoding module for decoding, and finally the delayed wake-up detection result of the delayed wake-up detection is obtained; although this scheme can guarantee the accuracy of voice wake-up to a certain extent, due to the introduction of delay, the real-time performance of voice wake-up is greatly affected by the existence of delay in the wake-up detection result. In order to overcome the defect that the real-time performance of voice wake-up cannot be guaranteed in the traditional scheme, in the embodiment of the present invention, different detection methods are adopted at different stages to perform wake-up detection on the real-time voice stream, and voice wake-up is performed according to the detection results, so as to achieve a balance between the wake-up effect and real-time performance of voice wake-up. This process specifically includes the following steps: Figure 7 This is an example diagram of the voice wake-up method provided by the present invention, such as Figure 7 As shown:
[0125] First, a delayed wakeup detection model is used to detect the wakeup keyword ("Hello Xiaoao") in the real-time voice stream. At the same time, the multitask MLP branch is cached.
[0126] Then, a pre-wake-up judgment is performed on the delayed wake-up detection result obtained by the delayed wake-up detection, that is, whether the delayed wake-up detection result is a pre-wake-up ("Hello, Xiao"); and if the delayed wake-up detection result is a pre-wake-up, the multitask MLP branch cached in the previous step is used to perform real-time wake-up detection on the real-time voice stream, that is, detect "Ao";
[0127] Accordingly, if the delayed wakeup detection result is not pre-wakeup, the delayed wakeup detection is continued through the delayed wakeup detection model until the delayed wakeup detection result is pre-wakeup;
[0128] Subsequently, voice wake-up is performed based on the real-time wake-up detection results and the delayed wake-up detection results. That is, if any one of the real-time wake-up detection and delayed wake-up detection detects "Ao", or if both detect "Ao", the wake-up result is thrown out, and the delayed wake-up detection and real-time wake-up detection of the real-time voice stream are reset, that is, the voice wake-up engine is restored to its initial state for the next round of wake-up detection. At this time, since the multitask mlp branch is a wake-up detection without delay, that is, the wake-up detection of "Ao" does not introduce delay, performing voice wake-up based on this can achieve the purpose of optimizing the response time of voice wake-up, and can also ensure the wake-up effect of voice wake-up;
[0129] Accordingly, if both the real-time wake-up detection result and the delayed wake-up detection result are not awakened, that is, neither the real-time wake-up detection nor the delayed wake-up detection detects "A", a timeout judgment is performed, that is, the duration of the real-time wake-up detection result being not awakened is recorded, and it is determined whether the duration exceeds the preset time; if the duration exceeds the preset time, the real-time wake-up detection of the real-time voice stream is interrupted, and only the delayed wake-up detection is performed on the real-time voice stream to reduce the occupation of computing resources and ensure a smaller amount of calculation; accordingly, if the duration does not exceed the preset time, the delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream is determined, that is, the delayed wake-up detection result corresponding to the voice frame of this window is determined;
[0130] Finally, the system determines whether the delayed wakeup detection result corresponding to the current voice frame indicates a wakeup. If so, the system discards the wakeup result and resets the delayed wakeup detection and real-time wakeup detection for the voice stream, restoring the voice wakeup engine to its initial state for the next round of wakeup detection. While this does not shorten the response delay of voice wakeup, it does ensure the accuracy of voice wakeup.
[0131] Correspondingly, if the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, continue to perform real-time wake-up detection on the real-time voice stream, and repeat the above process until the delayed wake-up detection result and / or the real-time wake-up detection result is awakened.
[0132] The method provided by the embodiment of the present invention performs delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result. When the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. This overcomes the defect of traditional solutions that cannot take into account both the real-time performance and the wake-up effect of voice wake-up. It can shorten the response delay without compromising the wake-up effect of voice wake-up, thereby achieving a balance between the wake-up effect and real-time performance of voice wake-up. Moreover, the pre-wake-up judgment is used to switch from delayed wake-up detection to real-time wake-up detection, so that the wake-up detection of the real-time voice stream can be smooth and orderly, thereby reducing the false wake-up rate of voice wake-up and ensuring the accuracy of voice wake-up.
[0133] The voice wake-up device provided by the present invention is described below. The voice wake-up device described below and the voice wake-up method described above can be referenced to each other.
[0134] Figure 8 This is a schematic diagram of the structure of the voice wake-up device provided by the present invention. Figure 8 As shown, the device includes:
[0135] A delayed wakeup detection unit 810 is configured to perform a delayed wakeup detection on a real-time voice stream and obtain a delayed wakeup detection result, wherein a voice frame corresponding to the delayed wakeup detection result differs from a voice frame corresponding to the real-time voice stream by a preset number of frames;
[0136] The real-time wake-up detection unit 820 is configured to perform real-time wake-up detection on the real-time voice stream if the delayed wake-up detection result is pre-wake-up, and perform voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, where the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0137] The voice wake-up device provided by the present invention performs delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result. When the delayed wake-up detection result is pre-wake-up, real-time wake-up detection is performed on the real-time voice stream, and voice wake-up is performed based on the real-time wake-up detection result obtained by the real-time wake-up detection. This overcomes the defect of traditional solutions that cannot take into account both the real-time performance and the wake-up effect of voice wake-up. It can shorten the response delay without compromising the wake-up effect of voice wake-up, thereby achieving a balance between the wake-up effect and the real-time performance of voice wake-up. Moreover, the pre-wake-up judgment is used to switch from delayed wake-up detection to real-time wake-up detection, so that the wake-up detection of the real-time voice stream can be smooth and orderly, thereby reducing the false wake-up rate of voice wake-up and ensuring the accuracy of voice wake-up.
[0138] Based on the above embodiment, the device further includes a voice wake-up unit, which is configured to:
[0139] Performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream respectively to obtain the delayed wake-up detection result and the real-time wake-up detection result, and performing voice wake-up based on the delayed wake-up detection result and the real-time wake-up detection result.
[0140] Based on the above embodiment, the voice wake-up unit is used to:
[0141] If the delayed wake-up detection result and / or the real-time wake-up detection result is wake-up, voice wake-up is performed, and the delayed wake-up detection and the real-time wake-up detection are reset.
[0142] Based on the above embodiment, the device further includes an interruption unit for
[0143] If both the delayed wake-up detection result and the real-time wake-up detection result are no wake-up, recording the duration of the real-time wake-up detection result being no wake-up;
[0144] If the duration exceeds a preset time, the real-time wake-up detection is interrupted.
[0145] Based on the above embodiment, the real-time wake-up detection unit 820 is configured to:
[0146] If the duration does not exceed the preset time, determining a delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream;
[0147] If the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, real-time wake-up detection is performed on the real-time voice stream until the delayed wake-up detection result and / or the real-time wake-up detection result is awakened.
[0148] Based on the above embodiment, the delayed wake-up detection unit 810 is configured to:
[0149] Based on the delay wake-up detection model, the real-time voice stream is subjected to delay wake-up detection;
[0150] The real-time wake-up detection unit 820 is used to:
[0151] Performing real-time wake-up detection on the real-time voice stream based on a real-time wake-up detection model;
[0152] The delayed wakeup detection model and the real-time wakeup detection model share an acoustic coding layer, and the acoustic coding layer is used to acoustically encode an input real-time speech stream.
[0153] Based on the above embodiment, the delayed wakeup detection model is trained based on a sample voice stream and a delayed wakeup detection label, and the delayed wakeup detection label includes at least one of awakening, non-awakening and pre-awakening.
[0154] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call the logic instructions in the memory 930 to execute the voice wake-up method, which includes: performing a delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result, wherein the voice frame corresponding to the delayed wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames; if the delayed wake-up detection result is a pre-wake-up, performing a real-time wake-up detection on the real-time voice stream, and performing voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, wherein the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0155] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0156] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the voice wake-up method provided by the above methods, the method including: performing delayed wake-up detection on the real-time voice stream to obtain a delayed wake-up detection result, the voice frame corresponding to the delayed wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames; if the delayed wake-up detection result is a pre-wake-up, performing real-time wake-up detection on the real-time voice stream, and performing voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, and the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0157] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the voice wake-up method provided by the above-mentioned methods, the method comprising: performing delayed wake-up detection on a real-time voice stream to obtain a delayed wake-up detection result, wherein the voice frame corresponding to the delayed wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames; if the delayed wake-up detection result is a pre-wake-up, performing real-time wake-up detection on the real-time voice stream, and performing voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, wherein the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream.
[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A voice wake-up method, characterized in that: include: Performing a delay wake-up detection on the real-time voice stream to obtain a delay wake-up detection result, wherein a voice frame corresponding to the delay wake-up detection result differs from a voice frame corresponding to the real-time voice stream by a preset number of frames; If the delayed wake-up detection result is pre-wake-up, performing real-time wake-up detection on the real-time voice stream, and performing voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, where the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream; The delayed wake-up detection result is pre-wake-up, indicating that the wake-up detection for the real-time voice stream has reached a preset condition for switching from delayed wake-up detection to real-time wake-up detection.
2. The voice wake-up method according to claim 1, characterized in that: The performing real-time wake-up detection on the real-time voice stream and performing voice wake-up based on a real-time wake-up detection result obtained by the real-time wake-up detection includes: Performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream respectively to obtain the delayed wake-up detection result and the real-time wake-up detection result, and performing voice wake-up based on the delayed wake-up detection result and the real-time wake-up detection result.
3. The voice wake-up method according to claim 2, characterized in that: The performing voice wake-up based on the delayed wake-up detection result and the real-time wake-up detection result includes: If the delayed wake-up detection result and / or the real-time wake-up detection result is wake-up, voice wake-up is performed, and the delayed wake-up detection and the real-time wake-up detection are reset.
4. The voice wake-up method according to claim 2, characterized in that: The method further includes performing delayed wake-up detection and real-time wake-up detection on the real-time voice stream respectively to obtain the delayed wake-up detection result and the real-time wake-up detection result. If both the delayed wake-up detection result and the real-time wake-up detection result are no wake-up, recording the duration of the real-time wake-up detection result being no wake-up; If the duration exceeds a preset time, the real-time wake-up detection is interrupted.
5. The voice wake-up method according to claim 4, characterized in that: The recording of the duration of the real-time wake-up detection result indicating no wake-up further includes: If the duration does not exceed the preset time, determining a delayed wake-up detection result corresponding to the voice frame at the current moment in the real-time voice stream; If the delayed wake-up detection result corresponding to the voice frame at the current moment is not awakened, real-time wake-up detection is performed on the real-time voice stream until the delayed wake-up detection result and / or the real-time wake-up detection result is awakened.
6. The voice wake-up method according to any one of claims 1 to 5, characterized in that: The delay wake-up detection on the real-time voice stream includes: Based on the delay wake-up detection model, the real-time voice stream is subjected to delay wake-up detection; The performing real-time wake-up detection on the real-time voice stream includes: Performing real-time wake-up detection on the real-time voice stream based on a real-time wake-up detection model; The delayed wakeup detection model and the real-time wakeup detection model share an acoustic coding layer, and the acoustic coding layer is used to acoustically encode an input real-time speech stream.
7. The voice wake-up method according to claim 6, characterized in that: The delayed wakeup detection model is trained based on a sample voice stream and a delayed wakeup detection label, where the delayed wakeup detection label includes at least one of wakeup, non-wakeup, and pre-wakeup.
8. A voice wake-up device, characterized in that: include: A delay wake-up detection unit is used to perform a delay wake-up detection on the real-time voice stream to obtain a delay wake-up detection result, wherein the voice frame corresponding to the delay wake-up detection result differs from the voice frame corresponding to the real-time voice stream by a preset number of frames; a real-time wake-up detection unit, configured to perform real-time wake-up detection on the real-time voice stream if the delayed wake-up detection result is pre-wake-up, and perform voice wake-up based on the real-time wake-up detection result obtained by the real-time wake-up detection, wherein the voice frame corresponding to the real-time wake-up detection result is the same as the voice frame corresponding to the real-time voice stream; The delayed wake-up detection result is pre-wake-up, indicating that the wake-up detection for the real-time voice stream has reached a preset condition for switching from delayed wake-up detection to real-time wake-up detection.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the voice wake-up method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the voice wake-up method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio wake-up method, electronic device and non-transient compute readable storage medium
CN109215647A
Voice awakening method and device and intelligent electronic equipment thereof
CN110364143A