Voice wake-up method, device and electronic device

By using the tail endpoint detection model in the speech recognition system, the accurate tail endpoint of the voice wake-up audio data is determined, and the speech recognition error problem caused by the wake-up audio tail endpoint in the prior art is solved, and the accuracy of the recognition system is improved.

CN115188370BActive Publication Date: 2025-06-10SOUNDAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210782346.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-06-10
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

In the prior art, when the user speaks clearly, the network model may determine the tail endpoint of the wake-up audio in advance, resulting in the subsequent part of the wake-up audio being misunderstood as a control command, resulting in an error in the speech recognition result.

Method used

The voice wake-up model is used to detect the real-time audio data wake-up word. When the matching degree exceeds the threshold, the tail endpoint detection model is used to determine the tail endpoint of the wake-up audio data to ensure the integrity of the wake-up audio data.

Benefits of technology

By accurately determining the tail endpoint of the wake-up audio, voice recognition errors caused by early recognition are avoided, and the accuracy of the voice recognition system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115188370B_ABST
    Figure CN115188370B_ABST
Patent Text Reader

Abstract

The present application discloses a voice wake-up method, device, and electronic device, belonging to the technical field of audio processing. Among them, the method includes: obtaining real-time audio data, and performing wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word; when wake-up audio data with a matching degree exceeding a preset threshold with the preset wake-up word is detected, performing end point detection on the real-time audio data based on an end point detection model and the end point of the preset wake-up word; when the end point corresponding to the wake-up audio data is detected, controlling a voice interaction system to perform a wake-up response. The embodiments of the present application can improve the accuracy of determining the end point of a preset wake-up word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of audio processing, and particularly relates to a voice wake-up method, device, and electronic device. Background Art

[0002] A user can speak a specific voice vocabulary to wake up a voice recognition system in a low-power standby state. In related technologies, a network model is usually used to match the audio input by the user with the specific voice vocabulary. When the matching degree between the two reaches a threshold, it is determined that the wake-up audio has been received, and thus the voice recognition system is woken up to respond to the voice command input by the user next.

[0003] However, in a scenario where the user's speech is relatively clear, it is very likely that the network model can obtain a matching result reaching the threshold before the voice device has received the complete wake-up vocabulary. At this time, the end point of the wake-up audio determined by the network model is earlier than the actual end point. Due to the advance of the end point, the part of the wake-up audio after this end point will be input into the voice recognition system as a control command or part of a control command, which will further cause errors in the voice recognition result. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a voice wake-up method, device, and electronic device, which can determine a more accurate end point of the wake-up audio.

[0005] In a first aspect, the embodiments of this application provide a voice wake-up method, which includes:

[0006] Obtain real-time audio data, and perform wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word;

[0007] When wake-up audio data with a matching degree exceeding a preset threshold with the preset wake-up word is detected, perform end point detection on the real-time audio data based on an end point detection model and the end point of the preset wake-up word;

[0008] When the end point corresponding to the wake-up audio data is detected, control the voice interaction system to perform a wake-up response.

[0009] In a second aspect, the embodiments of this application provide a voice wake-up device, including:

[0010] An obtaining module, configured to obtain real-time audio data, and perform wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word;

[0011] A detection module, configured to perform endpoint detection on the real-time audio data based on an endpoint detection model and the endpoint of the preset wake-up word when wake-up audio data with a matching degree exceeding a preset threshold with the preset wake-up word is detected;

[0012] A first control module, configured to control the voice interaction system to perform a wake-up response when the endpoint corresponding to the wake-up audio data is detected.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0017] In the embodiment of the present application, a voice wake-up model is used to detect a preset wake-up word in real-time audio data. When it is determined that real-time audio data of the preset wake-up word is received, an endpoint detection model can be used to detect the endpoint of the wake-up audio data corresponding to the preset wake-up word, and the real-time audio data after the endpoint of the preset wake-up word is input to the control voice interaction system for a wake-up response. Among them, using the endpoint detection model to perform endpoint detection on the real-time audio data can make the determined endpoint of the wake-up audio data be the endpoint of the last phoneme in the preset wake-up word, and can overcome the problem of incorrect speech recognition results caused by the premature endpoint. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of a voice wake-up method provided by an embodiment of the present application;

[0019] Figure 2 is a schematic structural diagram of a voice wake-up device provided by an embodiment of the present application;

[0020] Figure 3It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0023] Next, in conjunction with the accompanying drawings, the voice wake-up method, voice wake-up device, and electronic device provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0024] Please refer to Figure 1 , a voice wake-up method provided by an embodiment of the present application may include the following steps:

[0025] Step 101: Obtain real-time audio data, and perform wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word.

[0026] In implementation, the above real-time audio data may be audio data obtained in real time, and the above preset wake-up word may be a vocabulary preset by the user, for example: "Xiaomi AI Assistant".

[0027] Step 102: When wake-up audio data whose matching degree with the preset wake-up word exceeds a preset threshold is detected, perform end point detection on the real-time audio data based on an end point detection model and the end point of the preset wake-up word.

[0028] Step 103: When the end point corresponding to the wake-up audio data is detected, control the voice interaction system to perform a wake-up response.

[0029] In implementation, the voice wake-up model can perform frame-by-frame matching between the audio in the real-time audio data and the preset wake-up word. When the matching degree reaches the preset matching threshold (such as 80% or 90%, etc.), it can indicate that the voice wake-up model has detected the preset wake-up word.

[0030] It should be noted that in some scenarios, there may be a moment when the voice wake-up model detects the preset wake-up word at a certain phoneme in the middle of the preset wake-up word, rather than the actual end point of the real-time audio data corresponding to the preset wake-up word.

[0031] For example: Suppose the preset wake-up word is "Xiaoi Assistant", and the preset matching threshold is 80%. There may be a situation where when the user has only spoken half of the word "Assistant", the voice wake-up model detects that the matching degree between the audio data and the preset wake-up word has reached 80%. In the related art, the moment when the matching degree between the audio data and the preset wake-up word reaches 80% is determined as the end point of the real-time audio data corresponding to the preset wake-up word, and after this end point, the voice interaction system is controlled to perform a wake-up response. This will result in taking the second half of the half word "Assistant" as the voice command input to the voice interaction system, thus causing incorrect recognition of the voice command.

[0032] In the embodiments of the present application, when the preset wake-up word is detected, the end point detection model is further used to detect the end point of the real-time audio data corresponding to the preset wake-up word, which can improve the accuracy of the end point detection result, so that the detected end point of the real-time audio data can completely divide the last word of the preset wake-up word within the real-time audio data corresponding to the preset wake-up word, and reduce the probability of incorrect recognition of subsequent voice commands caused by inaccurate recognition of the end point of the preset wake-up word.

[0033] Optionally, the obtaining of the real-time audio data and the performing of the preset wake-up word detection on the real-time audio data based on the voice wake-up model and the preset wake-up word include:

[0034] Obtain real-time audio data;

[0035] Based on the voice wake-up model, perform frame-by-frame matching between the real-time audio data and the preset wake-up word;

[0036] When the matching degree reaches the preset matching threshold, it is determined that the preset wake-up word has been detected.

[0037] In implementation, the above voice wake-up model can be similar to the network model used to detect wake-up words in the related art, and will not be elaborated here.

[0038] The above end point detection model can be any network model such as a neural network model or a machine learning model. This end point detection model can be trained based on a large number of voice training samples with labeled conversion points of adjacent words.

[0039] Optionally, before obtaining the real-time audio data and performing a preset wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word, the voice wake-up method further includes:

[0040] Labeling the time points between two adjacent characters in the voice sample as a first value, and labeling other time points as a second value to obtain a training sample;

[0041] Inputting the training sample into a character segmentation model to be trained for model training to obtain the end point detection model.

[0042] In implementation, methods such as manual labeling or artificial intelligence detection can be adopted. According to the changes in the tone, audio amplitude, etc. of the voice sample, the time points (i.e., the conversion points between two adjacent characters) between two adjacent characters in the voice sample are distinguished from other non-conversion points, and the conversion points and non-conversion points are labeled differently. For example, the conversion points are labeled as 1 (i.e., the first value can be equal to 1), and other non-conversion points are labeled as 0 (i.e., the second value can be equal to 1).

[0043] After labeling, the conversion points and non-conversion points can be distinguished by the first value and the second number. In this way, the end point detection model trained based on this training sample can identify the conversion points between two adjacent characters in the voice audio data. During the process of detecting the end point of the wake-up audio data corresponding to the preset wake-up word based on the end point detection model, the end point of the wake-up audio data can be determined from the conversion points, so that the wake-up audio data corresponding to the preset wake-up word before the end point is complete.

[0044] In this embodiment, after the real-time audio data is input into the end point detection model, the end point detection model can determine the conversion points between two adjacent characters in the real-time audio data. In this way, when the user says the complete preset wake-up word, the end point detection model processes the voice spoken by the user to obtain the conversion point between the last character in the preset wake-up word and other subsequent words, so that the conversion point can be determined as the end point of the wake-up audio data corresponding to the preset wake-up word, ensuring that the determined end point will not be advanced.

[0045] In implementation, if the end point of the wake-up audio data corresponding to the estimated preset wake-up word is advanced, it will cause speech recognition errors. For example: Assume the preset wake-up word is "Xiaomi AI Assistant". When the user says "Xiaomi AI Assistant, play music", if the user's pronunciation is relatively clear, then in the process of determining the matching degree between the user's speech and "Xiaomi AI Assistant" based on the method in the related technology, it is possible that when the user only utters a part of the word "Assistant", it is determined that the matching degree between the user's speech and "Xiaomi AI Assistant" reaches the threshold, and based on this, the end point of the preset wake-up word is determined as the time point when the matching degree reaches the threshold. Thus, it can be seen that the end point of the preset wake-up word estimated in the related technology is advanced relative to the end point of the word "Assistant", which will cause the remaining part of the word "Assistant" to be input into the speech recognition system, that is, the speech command received by the speech recognition system is a part of the word "Assistant" + "play music", and the speech recognition system may misjudge the received speech command as "Don't play music" accordingly.

[0046] It should be noted that in the embodiments of the present application, the above step 101 can be a continuous process, which can continue until after step 102 or even step 103 is executed. For example: If the user's speech is "Xiaomi AI Assistant, play music", then the real-time audio data obtained is the speech data stream of "Xiaomi AI Assistant, play music", and when the speech data of "Xiaomi AI Assistant" is obtained, step 102 can be executed, and after step 102 is executed, the speech data of "play music" can continue to be obtained, so as to, after step 3, use the speech interaction system to respond to the speech command of "play music".

[0047] In an alternative implementation manner, the method further includes:

[0048] Obtain at least one preset wake-up word;

[0049] Determine the termination point of the last phoneme in the preset wake-up word as the end point of the preset wake-up word;

[0050] Determine the end point of the preset wake-up word as the detection object of the end point detection model.

[0051] Wherein, a phoneme can be a character, a letter, a note, etc., which is not specifically limited herein.

[0052] The above process of updating the end point detection model according to the at least one preset wake-up word can be a process of adjusting the model parameters of the end point detection model with the goal that the end point detection model can determine the termination point of the last phoneme corresponding to the preset wake-up word as the end point of the wake-up audio data corresponding to the preset wake-up word.

[0053] In this embodiment, during the model training process, the model parameters of the end point detection model obtained by training can be adjusted according to the preset wake-up words set by the user, so that the end point detection model can detect the audio data including the preset wake-up words and determine the end point of the last phoneme in the preset wake-up words as the end point of the wake-up audio data corresponding to the preset wake-up words.

[0054] In another embodiment, the end point detection of the real-time audio data based on the end point detection model and the end point of the preset wake-up word includes:

[0055] Detecting at least one phoneme in the preset wake-up word based on the end point detection model;

[0056] Determining the end point of the last phoneme in the preset wake-up word as the end point of the real-time audio data.

[0057] In practice, the end point detection model can be used to detect the real-time audio data received after detecting the preset wake-up word to obtain the end point of the last phoneme in the preset wake-up word. In practice, when the preset wake-up word is detected, at least N-1 phoneme audio data in the preset wake-up word has been obtained. At this time, the end point detection model only needs to detect the end point of the first phoneme in the audio data obtained after that, and then it can determine that this end point is the end point of the last phoneme in the preset wake-up word.

[0058] Of course, the end point detection model can also detect the cached audio data (i.e., the audio data before detecting the preset wake-up word) and the real-time audio data received after detecting the preset wake-up word to determine the end point of the last phoneme in the preset wake-up word according to the mutual relationship between the phonemes in the preset wake-up word, where the cached audio data includes the real-time audio data received at historical times.

[0059] For example: Suppose the preset wake-up word is "Xiaomi AI". The user says "Xiaomi AI" between 0ms and 500ms, and the reception time of the audio data corresponding to the word "AI" is between 350ms and 500ms. And based on the voice wake-up model, at 400ms, it is determined that the matching degree between the audio data received between 0ms and 400ms and the preset wake-up word reaches the preset matching threshold (such as 85%). Then the end point detection model can detect the audio data after 400ms, or detect the cached audio data received between 0ms and 400ms and the audio data after 400ms to determine that the end point of the audio of "Xiaomi AI" is 500ms.

[0060] In this embodiment, a voice wake-up model is used to determine the matching degree between the real-time audio data and the preset wake-up word. When the matching degree reaches the preset matching threshold, it can indicate that the electronic device has indeed received the preset wake-up word, thereby allowing the voice interaction system to be started. However, the end point of the wake-up audio data corresponding to the preset wake-up word is determined based on an end point detection model. In this way, the starting point of the voice command input to the voice interaction system can be the first sampling point after the end point of the wake-up audio data, that is, the starting point of the voice command input to the voice interaction system can be 501 ms.

[0061] In practice, when the electronic device is not working, the voice interaction system is in a low-power standby state. When a wake-up word is received, the voice interaction system is awakened to enter the working state, and at least one of the processes such as recognizing, executing, and responding to the voice command is performed.

[0062] In one embodiment, the voice wake-up model can be started first, and when the matching degree between the real-time voice data received by the electronic device and the preset wake-up word determined based on the voice wake-up model reaches the preset matching threshold, the end point detection model is started to determine the end point of the wake-up audio data corresponding to the preset wake-up word based on the end point detection model.

[0063] In this way, it can be determined whether to wake up the voice interaction system based on the voice wake-up model, and when it is determined to wake up the voice interaction system based on the voice wake-up model, the end point of the wake-up audio data corresponding to the preset wake-up word is determined based on the end point detection model. In this embodiment, before it is determined to wake up the voice interaction system based on the voice wake-up model, the end point detection model is not used for calculation temporarily. In this way, the computing amount of the electronic device can be reduced, and the standby power consumption can be reduced.

[0064] In addition, after it is determined to wake up the voice interaction system based on the voice wake-up model, the calculation based on the voice wake-up model can be stopped, and only the end point of the wake-up audio data corresponding to the preset wake-up word needs to be determined based on the end point detection model.

[0065] In another embodiment, the end point detection model can be started before the matching degree between the voice received by the electronic device and the preset wake-up word determined based on the voice wake-up model reaches the preset matching threshold, so that the end point detection model determines the end point of each phoneme in the real-time audio data.

[0066] In this way, when the matching degree between the voice received by the electronic device and the preset wake-up word reaches the preset matching threshold based on the voice wake-up model, the end points of each phoneme in the preset wake-up word except the last phoneme have been determined based on the end point detection model. At this time, the voice wake-up model can comprehensively determine the end point of the last phoneme in the preset wake-up word according to the end points of the previous several phonemes, which can improve the accuracy of the end point of the last phoneme in the determined preset wake-up word.

[0067] Optionally, when the preset wake-up word includes N phonemes and N is an integer greater than 1, detecting at least one phoneme in the preset wake-up word based on the end point detection model includes:

[0068] Based on the end point detection model, determine N end points in the real-time audio data that correspond one-to-one to the N phonemes.

[0069] In this embodiment, the end point of each phoneme in the preset wake-up word can be determined based on the end point detection model. Thus, when determining the end points of the N phonemes in sequence, it can be determined that the end point of the Nth phoneme determined last is the end point of the wake-up audio data corresponding to the preset wake-up word.

[0070] It is worth noting that in implementation, due to the influence of environmental noise interference, low user pronunciation clarity, etc., the end points determined by the end point detection model may be inaccurate. At this time, after detecting the end point of the wake-up audio data corresponding to the preset wake-up word based on the end point detection model, if the end point of the wake-up audio data corresponding to the preset wake-up word has not been determined based on the end point detection model after a certain time interval, the voice interaction system can be directly awakened, and the subsequent voice commands can be processed by awakening the voice interaction system.

[0071] Optionally, after performing end point detection on the real-time audio data based on the end point detection model and the end point of the preset wake-up word, the method further includes:

[0072] When the detection time for the end point of the real-time audio data exceeds the preset time threshold, control the voice interaction system to perform a wake-up response.

[0073] In implementation, the above preset time threshold can be a preset time length, for example: 100 ms, 200 ms, or 300 ms, etc.

[0074] In this embodiment, when the end point detection model fails to detect the end point of the wake-up audio data corresponding to the preset wake-up word due to the influence of environmental noise interference, low user pronunciation clarity, etc., when the detection time exceeds the preset time threshold, the voice interaction system can be directly controlled to perform a wake-up response.

[0075] For example: when the user's pronunciation is unclear, causing the end point detection model to recognize two phonemes in the preset wake-up word as one phoneme, then based on the end point detection model, only the end point of (N - 1) phonemes in the preset wake-up word is recognized; or, due to relatively high environmental noise, the end point detection model cannot recognize the end point of the Nth phoneme in the preset wake-up word, etc. In such cases, when the detection time for the end point of the wake-up audio data corresponding to the preset wake-up word exceeds the preset time threshold, the voice interaction system can be controlled to perform a wake-up response.

[0076] As an optional embodiment, the voice wake-up method further includes:

[0077] When wake-up audio data with a matching degree exceeding a preset threshold with the preset wake-up word is detected, the voice interaction system is controlled to enter a pre-wake-up state. In this pre-wake-up state, the voice interaction system is started, and no wake-up response is made to the real-time audio data.

[0078] In this embodiment, when the preset wake-up word is detected, the voice interaction system is pre-woken up, and when the end point of the wake-up audio data is determined based on the end point detection model, the received audio data is divided based on the time of this end point, so that the audio data received after this end point is input into the already woken-up voice interaction system for processing. In this way, compared with the method of waking up the voice interaction system after the end point of the wake-up audio data is determined based on the end point detection model, the waiting time between when the user speaks a voice command and when the voice interaction system is woken up can be reduced.

[0079] For example: assuming the preset wake-up word is "Xiaoi", the electronic device receives the audio data stream of "Xiaoi, play music" spoken by the user between 0ms and 1000ms, and the reception time of the audio data corresponding to the word "xue" is between 350ms and 500ms. And based on the voice wake-up model, it is determined at 400ms that the matching degree between the audio data received between 0ms and 400ms and the preset wake-up word reaches the preset matching threshold (such as 85%). Then the voice interaction system will be pre-woken up at 400ms, and the real-time audio data after 500ms will be input into the voice interaction system for processing.

[0080] In an embodiment of the present application, a voice wake-up model is used to detect a preset wake-up word in real-time audio data. When it is determined that real-time audio data containing the preset wake-up word is received, a tail-end point detection model can be used to detect the tail-end point of the wake-up audio data corresponding to the preset wake-up word, and the real-time audio data after the tail-end point of the preset wake-up word is input to a control voice interaction system for wake-up response. Among them, using the tail-end point detection model to perform tail-end point detection on the real-time audio data can make the determined tail-end point of the wake-up audio data be the tail-end point of the last phoneme in the preset wake-up word, and can overcome the problem of incorrect speech recognition results caused by the premature tail-end point.

[0081] For the voice wake-up method provided by an embodiment of the present application, the execution subject may be a voice wake-up device. In an embodiment of the present application, taking the voice wake-up device executing the voice wake-up method as an example, the voice wake-up device provided by the embodiment of the present application is described.

[0082] Please refer to Figure 2 , the voice wake-up device 200 provided by an embodiment of the present application may include the following modules:

[0083] A first acquisition module 201, configured to acquire real-time audio data, and perform wake-up word detection on the real-time audio data based on a voice wake-up model and a preset wake-up word;

[0084] A detection module 202, configured to, when detecting wake-up audio data whose matching degree with the preset wake-up word exceeds a preset threshold, perform tail-end point detection on the real-time audio data based on the tail-end point detection model and the tail-end point of the preset wake-up word;

[0085] A first control module 203, configured to, when detecting the tail-end point corresponding to the wake-up audio data, control the voice interaction system to perform a wake-up response.

[0086] Optionally, the voice wake-up device 200 further includes:

[0087] A marking module, configured to mark the time points between adjacent two characters in a voice sample as a first value, and mark other time points as a second value, to obtain a training sample;

[0088] A training module, configured to input the training sample into a to-be-trained character segmentation model for model training, to obtain the tail-end point detection model.

[0089] Optionally, the voice wake-up device 200 further includes:

[0090] A second acquisition module, configured to acquire at least one preset wake-up word;

[0091] A first determination module, configured to determine the termination point of the last phoneme in the preset wake-up word as the tail-end point of the preset wake-up word;

[0092] A second determination module, configured to determine the end point of the preset wake-up word as the detection object of the end point detection model.

[0093] Optionally, the detection module 202 includes:

[0094] A detection unit, configured to detect at least one phoneme in the preset wake-up word based on the end point detection model;

[0095] A first determination unit, configured to determine the termination point of the last phoneme in the preset wake-up word as the end point of the real-time audio data.

[0096] Optionally, the voice wake-up device 200 further includes:

[0097] A second control module, configured to control the voice interaction system to perform a wake-up response when the detection time for the end point of the real-time audio data exceeds a preset time threshold.

[0098] Optionally, the first acquisition module 201 includes:

[0099] An acquisition unit, configured to acquire real-time audio data;

[0100] A matching unit, configured to perform frame-by-frame matching between the real-time audio data and the preset wake-up word based on the voice wake-up model;

[0101] A second determination unit, configured to determine that the preset wake-up word is detected when the matching degree reaches a preset matching threshold.

[0102] Optionally, the voice wake-up device 200 further includes:

[0103] A third control module, configured to control the voice interaction system to enter a pre-wake-up state when wake-up audio data with a matching degree exceeding a preset threshold with respect to the preset wake-up word is detected, where in the pre-wake-up state, the voice interaction system is started and does not perform a wake-up response to the real-time audio data.

[0104] The voice wake-up device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0105] The voice wake-up device provided in the embodiments of the present application can implement each process implemented by the method embodiments as Figure 1 shown, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.

[0106] Optionally, as Figure 3 shown, the embodiments of the present application further provide an electronic device 600, including a processor 601 and a memory 602. A program or instruction that can run on the processor 601 is stored on the memory 602. When the program or instruction is executed by the processor 601, it implements each step of the above-mentioned voice wake-up method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0107] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0108] The embodiments of the present application further provide a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements each process of the above-mentioned voice wake-up method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0109] Wherein, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0110] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-described embodiment of the voice wake-up method and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0111] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-a-chip, etc.

[0112] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-described embodiment of the method for determining the end point of a voice tail and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0113] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0115] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A voice wake-up method, characterized in that, it includes: while collecting audio data in real time, performing the following steps based on the currently collected real-time audio data: Based on the voice wake-up model, frame-by-frame matching the real-time audio data with a preset wake-up word; When the matching degree reaches a preset matching threshold, it is determined that the preset wake-up word is detected, and based on the end-point detection model and the end point of the preset wake-up word, end-point detection is performed on the audio data after the first moment in the real-time audio data, and the voice interaction system is controlled to enter the pre-wake-up state. Wherein, in the pre-wake-up state, the voice interaction system is started and does not perform wake-up response on the real-time audio data. The first moment is the moment when the matching degree reaches the preset matching threshold; When the end point corresponding to the wake-up audio data is detected, control the awakened voice interaction system to perform a wake-up response on the real-time audio data collected after the end point.

2. The method according to claim 1, characterized in that, Before the frame-by-frame matching of the real-time audio data with the preset wake-up word based on the voice wake-up model, the method further includes: Labeling the time points between two adjacent characters in the voice sample as a first value, and labeling other time points as a second value to obtain a training sample; Inputting the training sample into the sub-character model to be trained for model training to obtain the end-point detection model.

3. The method according to claim 2, characterized in that, The method further includes: Obtaining at least one preset wake-up word; Determining the termination point of the last phoneme in the preset wake-up word as the end point of the preset wake-up word; Determining the end point of the preset wake-up word as the detection object of the end-point detection model.

4. The method according to claim 1, characterized in that, After the frame-by-frame matching of the real-time audio data with the preset wake-up word based on the voice wake-up model, the method further includes: When the detection time for the end point of the real-time audio data exceeds a preset time threshold, control the voice interaction system to perform a wake-up response.

5. A voice wake-up device, characterized in that, it includes: An acquisition module, configured to while collecting audio data in real time, perform the following steps based on the currently collected real-time audio data: Based on the voice wake-up model, frame-by-frame matching the real-time audio data with a preset wake-up word; When the matching degree reaches a preset matching threshold, it is determined that the preset wake-up word is detected, and based on the end-point detection model and the end point of the preset wake-up word, end-point detection is performed on the audio data after the first moment in the real-time audio data, and the voice interaction system is controlled to enter the pre-wake-up state. Wherein, in the pre-wake-up state, the voice interaction system is started and does not perform wake-up response on the real-time audio data. The first moment is the moment when the matching degree reaches the preset matching threshold; When the end point corresponding to the wake-up audio data is detected, control the awakened voice interaction system to perform a wake-up response on the real-time audio data collected after the end point.

6. The device according to claim 5, wherein, it further comprises: a marking module, configured to mark the time points between two adjacent characters in the voice sample as a first value, and mark other time points as a second value, so as to obtain a training sample; a training module, configured to input the training sample into a character segmentation model to be trained for model training, so as to obtain the end point detection model.

7. The device according to claim 6, wherein, it further comprises: a second obtaining module, configured to obtain at least one preset wake-up word; a first determining module, configured to determine the end point of the last phoneme in the preset wake-up word as the end point of the preset wake-up word; a second determining module, configured to determine the end point of the preset wake-up word as the detection object of the end point detection model.

8. The device according to claim 5, wherein, it further comprises: a second control module, configured to control the voice interaction system to perform a wake-up response when the detection time of the end point of the real-time audio data exceeds a preset time threshold.

9. An electronic device, wherein, it comprises a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the voice wake-up method according to any one of claims 1 to 4 are implemented.

10. A readable storage medium, wherein, a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the voice wake-up method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • A voice data processing method and apparatus

    CN108962262A

  • Voice processing method and equipment

    CN109994106A

  • Human-computer interaction method, human-computer interaction device, electronic equipment and storage medium

    CN110634483A